Adaptive accent speech recognition method based on feature decoupling

By adopting an adaptive accent speech recognition method based on feature decoupling in the speech recognition system, using multi-task meta-learning and a dual-branch decoder with Transformer architecture, the shortcomings of the existing system in diversity and complex accent processing are solved, and the adaptability and accuracy of the system are significantly improved.

CN119964559APending Publication Date: 2025-05-09SHENZHEN POLYTECHNIC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510029575.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing speech recognition systems face difficulties in data set coverage, limited model learning effects, and difficulty in separating accent information from language content when dealing with diversity and complex accents. Especially when there is no accent or limited data volume, the system's generalization ability and adaptability are insufficient.

Method used

Adaptive accent speech recognition method based on feature decoupling is adopted, and the model parameters are optimized through multi-task meta-learning and fine-tuning. The dual-branch decoder in the Transformer architecture is used to simultaneously perform speech recognition and accent recognition tasks, effectively separating language content and accent information.

Benefits of technology

It significantly improves the system's adaptability to diversity and complex accents, ensures that high accuracy and strong robustness can be maintained in the face of no accents or limited data volume, greatly enhances the system's generalization ability and adaptability, and provides more reliable and accurate speech recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964559A_ABST
    Figure CN119964559A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition, in particular to a feature decoupling-based adaptive accent speech recognition method, which comprises the following steps of: calling a speech recognition model obtained through multi-task element learning type adaptive training as a starting point, and deploying the fine-tuned speech recognition model in an actual application environment; inputting the preprocessed voice signal into a voice recognition model to generate a corresponding voice recognition result; in the process of generating a corresponding voice recognition result, capturing context information in a voice signal, generating recognition text output and recognition accent labels for acoustic features to be recognized, and generating a final text transcription result by using a decoding algorithm based on the context information, and generating a corresponding speech recognition result by combining the text transcription result with the recognition text output and the recognition accent tag. According to the invention, the speech recognition performance for coping with diverse and complex accent can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech recognition, and in particular to an adaptive accent speech recognition method based on feature decoupling. Background Art

[0002] As globalization accelerates, cross-language and cross-cultural communication becomes more frequent. As an important tool to promote global informatization, speech recognition technology has become increasingly important. However, when dealing with speech with accents, current speech recognition systems face two core challenges: First, the diversity and complexity of accents make it extremely difficult to build a comprehensive dataset covering all possible accents, which increases the cost of data collection and limits the learning effect of the model; second, the complexity of speech information itself and the superposition of accent features make it difficult for traditional methods to effectively separate accent information from language content, thus affecting model performance.

[0003] Especially when facing unseen accents or when the amount of data is limited, the system's generalization ability and adaptability are particularly insufficient. From the above, we can see that how to improve the speech recognition performance for diverse and complex accents remains to be solved. Summary of the invention

[0004] In order to improve the speech recognition performance for diverse and complex accents, the present application provides an adaptive accent speech recognition method based on feature decoupling.

[0005] In a first aspect, the present application provides an adaptive accent speech recognition method based on feature decoupling, which adopts the following technical solution:

[0006] A method for adaptive accent speech recognition based on feature decoupling comprises: taking a speech recognition model obtained through multi-task meta-learning adaptive training as a starting point, initializing model parameters of the speech recognition model, and fine-tuning the model parameters of the speech recognition model, and deploying the fine-tuned speech recognition model in an actual application environment; capturing a user's speech signal in the actual application environment, preprocessing the speech signal, and converting the preprocessed speech signal into a to-be-recognized spectrum graph, extracting to-be-recognized acoustic features through a feature processor based on the to-be-recognized spectrum graph, inputting the to-be-recognized acoustic features into a speech recognition model, and generating a corresponding speech recognition result; in the process of generating the corresponding speech recognition result, a Transformer architecture encoder layer in the speech recognition model processes the input to-be-recognized acoustic features to capture context information in the speech signal, a dual-branch decoder performs speech recognition tasks and accent recognition tasks on the to-be-recognized acoustic features respectively, generates a recognition text output and a recognition accent label for the to-be-recognized acoustic features, generates a final text transcription result based on the context information using a decoding algorithm, and generates a corresponding speech recognition result by combining the text transcription result with the recognition text output and the recognition accent label.

[0007] By adopting the above technical solutions, the system's ability to cope with diverse and complex accents is significantly improved through multi-task meta-learning and fine-tuning to optimize model parameters. It can not only efficiently capture and process the user's voice signals in actual application environments, but also simultaneously perform speech recognition and accent recognition tasks through the dual-branch decoder in the Transformer architecture, thereby effectively separating accent information from language content. This method ensures that the model can maintain high accuracy and strong robustness in the face of unseen accents or limited data, greatly enhancing the system's generalization ability and adaptability, thereby providing more reliable and accurate speech recognition results.

[0008] Optionally, the speech recognition model training process includes: collecting speech data with different accents, converting the speech data into a spectrogram, and extracting acoustic features from the spectrogram using a feature processor, selecting a Transformer architecture as a recognition model, initializing the parameters of the encoder and decoder in the Transformer architecture, and setting the configuration of the recognition model, wherein the multi-head attention mechanism and position encoding characteristics of the Transformer architecture can effectively capture information in the speech data; applying a meta-learning algorithm to the Transformer architecture, and performing a multi-task learning framework, introducing an accent recognition task to combine the meta-learning algorithm and the multi-task learning framework, and learning to separate accent information from language content; using the parameters of the recognition model trained by meta-learning and multi-task learning as a starting point, collecting and preprocessing speech samples with unseen accents, gradually adjusting the parameters of the recognition model based on the preprocessed speech samples with unseen accents, and gradually improving the performance of the recognition model on unseen accents through multiple iterations to determine the corresponding speech recognition model.

[0009] By adopting the above technical solutions, the multi-head attention mechanism and position encoding characteristics of the Transformer architecture are used to effectively capture the information in the speech data, and the accent recognition task is introduced to separate the accent features from the language content. The model parameters trained by meta-learning and multi-task learning are used as the starting point. Through gradual adjustment and multiple iterative optimization, especially fine-tuning for unseen accent samples, the accuracy and robustness of the model in processing unknown accents are greatly improved.

[0010] Optionally, in the process of applying the meta-learning algorithm to the Transformer architecture, the method further comprises: representing the speech recognition of the Transformer as f θ θ represents the parameter, and the data of each accent i is divided into and Will be in The first gradient descent update is performed above, updating θ to θ′, where Among them, α is the fast adaptation learning rate; in the training task, the model f(θ) optimized by the parameters obtained by training i ′) to improve the model’s performance in unseen categories The recognition performance on , the meta-objective is defined as: in Indicates that The above loss, collects the loss from a batch Use the obtained meta-validation loss to optimize the parameters in the test task: Where β is the learning step size in meta-learning, It is in accent Ai Adaptive network on ; in fine-tuning,

[0011] By adopting the above technical solution, by grouping the data of each accent and performing preliminary gradient descent updates on a specific subset of accents, the model is able to quickly adjust parameters to adapt to the characteristics of newly encountered accents. This fast adaptation mechanism utilizes a meta-objective optimization strategy to optimize the parameters in the test task by minimizing the meta-validation loss, ensuring that the model not only performs well on known accents, but also significantly improves its recognition performance on unseen accents. Ultimately, this approach enables the speech recognition system to maintain a high degree of accuracy and robustness in the face of diverse and complex accents, greatly enhancing the system's generalization ability and practical application effects.

[0012] Optionally, in the process of the multi-task learning framework, the method also includes: the multi-task learning framework includes a shared encoder and a dual-branch decoder, the shared encoder is used for both speech recognition and accent recognition, one branch decoder in the dual-branch decoder is used for speech recognition, and the other branch decoder is used for accent recognition.

[0013] By adopting the above technical solutions and designing a shared encoder and a dual-branch decoder, the performance of the model in processing diverse accents has been significantly improved. The shared encoder allows speech recognition and accent recognition tasks to share the feature extraction process, promoting information sharing and collaborative optimization between tasks. The dual-branch decoder focuses on speech recognition and accent recognition respectively, ensuring that each task can receive targeted processing, thereby effectively separating accent features from language content.

[0014] Optionally, in the process of introducing the accent recognition task, the method further comprises: in the training process, introducing an additional overall loss, the encoder and the target loss L from speech recognition asr_loss , target loss L from accent recognition accent_loss Joint optimization, overall loss L loss =βL asr_loss +λL accent_loss , where β and λ are hyperparameters, and they are usually limited by β+λ=1.

[0015] By adopting the above technical solution and introducing the accent recognition task, the model realizes the joint optimization of the encoder with the speech recognition and accent recognition tasks by adding an additional overall loss function. This joint optimization strategy ensures that the model not only focuses on improving the accuracy of speech recognition, but also effectively learns and separates accent features. The introduction of hyperparameters makes it possible to dynamically balance the importance of the two tasks during training, thereby significantly enhancing the ability to learn accent features while maximizing the performance of the main task (speech recognition).

[0016] Optionally, the method further includes a data preprocessing step, and further includes: receiving the user's original audio input, and converting it into a frequency spectrum to be identified through a feature processor, wherein the feature processor adopts a VGG model with a 6-layer convolutional neural network architecture, and the VGG model is used to extract acoustic features from the original audio signal and generate a two-dimensional frequency spectrum; passing the generated frequency spectrum as input to the Transformer model to generate a corresponding text transcription result, wherein the Transformer model includes two encoder layers and four decoder layers, and each encoder and decoder layer uses an 8-head self-attention mechanism.

[0017] By adopting the above technical solutions and introducing efficient data preprocessing steps, the performance and robustness of the speech recognition system are significantly enhanced. The feature processor uses a VGG model with a 6-layer convolutional neural network architecture, which can accurately extract acoustic features from the original audio signal and generate a two-dimensional spectrogram, providing high-quality input for subsequent processing. The generated spectrogram is then passed to a Transformer model consisting of two encoder layers and four decoder layers. An 8-head self-attention mechanism is used within each layer to ensure that the model can capture complex speech context information.

[0018] In a second aspect, the present application provides an adaptive accent speech recognition device based on feature decoupling, which adopts the following technical solution:

[0019] An adaptive accent speech recognition device based on feature decoupling, characterized by comprising:

[0020] The model initialization module uses the speech recognition model obtained through multi-task meta-learning adaptive training as a starting point to initialize the model parameters of the speech recognition model, fine-tune the model parameters of the speech recognition model, and deploy the fine-tuned speech recognition model in the actual application environment;

[0021] A speech recognition result generation module captures the user's speech signal in an actual application environment, preprocesses the speech signal, and converts the preprocessed speech signal into a to-be-recognized spectrogram, extracts the to-be-recognized acoustic features through a feature processor based on the to-be-recognized spectrogram, inputs the to-be-recognized acoustic features into a speech recognition model, and generates a corresponding speech recognition result;

[0022] The text transcription result generation module, in the process of generating the corresponding speech recognition result, the Transformer architecture encoder layer in the speech recognition model processes the input acoustic features to be recognized to capture the contextual information in the speech signal, and the dual-branch decoder performs the speech recognition task and the accent recognition task on the acoustic features to be recognized respectively, generates the recognition text output and the recognition accent label for the acoustic features to be recognized, uses the decoding algorithm based on the context information to generate the final text transcription result, and combines the text transcription result with the recognition text output and the recognition accent label to generate the corresponding speech recognition result.

[0023] In a third aspect, the present application provides an adaptive accent speech recognition method based on feature decoupling, which adopts the following technical solution:

[0024] A method for adaptive accent speech recognition based on feature decoupling comprises a processor in which a program of any one of the methods for adaptive accent speech recognition based on feature decoupling described above is run.

[0025] In a fourth aspect, the present application provides a storage medium, which adopts the following technical solution:

[0026] A storage medium stores a program of any one of the above-mentioned adaptive accent speech recognition methods based on feature decoupling.

[0027] In summary, the present application includes at least one of the following beneficial technical effects:

[0028] The system's adaptability to various accents is improved through multi-task meta-learning and fine-tuning to optimize the parameters of the speech recognition model. It can not only efficiently process the user's voice signal, but also use the dual-branch decoder in the Transformer architecture to simultaneously perform speech recognition and accent recognition tasks, thereby effectively separating language content and accent information. This method ensures that high accuracy and strong robustness can be maintained even in the face of unseen accents or limited data, greatly enhancing the system's generalization and adaptability, and providing more reliable and accurate speech recognition results.

[0029] In addition, the performance and robustness of the speech recognition system were further improved by introducing efficient preprocessing steps and a VGG model with a 6-layer convolutional neural network architecture for acoustic feature extraction, and inputting the generated spectrogram into a Transformer model with a self-attention mechanism for text transcription. The design of a shared encoder and a dual-branch decoder promotes information sharing and collaborative optimization between tasks, ensuring that each task can receive targeted processing, thereby effectively improving the results of speech recognition and accent recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1The present invention is a diagram of an architecture of an accented speech recognition based on multiple tasks in an adaptive accented speech recognition method based on feature decoupling according to an exemplary embodiment.

[0031] Figure 2 It is a diagram showing the division of data sets of two experimental devices in a method for adaptive accented speech recognition based on feature decoupling according to an exemplary embodiment.

[0032] Figure 3 It is a comparison diagram of the results of three experimental methods under a mixed region setting in an adaptive accent speech recognition method based on feature decoupling according to an exemplary embodiment.

[0033] Figure 4 The figure is a diagram showing the accent distribution ratio under a mixed area setting in an adaptive accent speech recognition method based on feature decoupling according to an exemplary embodiment.

[0034] Figure 5 It is a comparison diagram of the results of three experimental methods in a cross-region setting in an adaptive accent speech recognition method based on feature decoupling according to an exemplary embodiment.

[0035] Figure 6 The figure is a diagram showing the accent distribution ratio in a cross-region setting in an adaptive accent speech recognition method based on feature decoupling according to an exemplary embodiment.

[0036] Figure 7 It is a comparison diagram of the error rate reduction of three experimental methods under different experimental settings in an adaptive accent speech recognition method based on feature decoupling according to an exemplary embodiment.

[0037] Figure 8 The present invention is a flowchart of a method for adaptive accent speech recognition based on feature decoupling according to an exemplary embodiment.

[0038] Fig. 9 The figure is a structural block diagram of a feature decoupling-based adaptive accent speech recognition device according to an exemplary embodiment. DETAILED DESCRIPTION

[0039] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings.

[0040] In the description of this specification, the description with reference to the terms "certain embodiments", "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0041] The embodiment of the present application discloses an adaptive accent speech recognition method based on feature decoupling. In the embodiment of the present application, the method is applied to a pre-trained speech recognition model. The training process of the speech recognition model is now described in detail, specifically including:

[0042] S001, collect speech data with different accents, convert the speech data into spectrograms, and use feature processors to extract acoustic features from the spectrograms, select the Transformer architecture as the recognition model, initialize the parameters of the encoder and decoder in the Transformer architecture, and set the configuration of the recognition model.

[0043] Among them, a large speech dataset covering multiple accent types is constructed to ensure that the dataset covers a wide range of diverse accents. This includes but is not limited to regional accents, social dialects, age and gender differences, etc. The data source can be a public dataset, user contributions, or specially recorded samples; the collected raw audio data is preprocessed. It should be pointed out here that for this application, the preprocessing process is divided into three steps, including:

[0044] Data preprocessing and feature extraction: First, the system receives the user's original audio input and converts it into a spectrogram to be recognized through an efficient feature processor. The feature processor uses a VGG model with a 6-layer convolutional neural network architecture, which is specially designed to extract acoustic features from raw audio signals and generate two-dimensional spectrograms. The VGG model captures the time and frequency characteristics of the audio through multi-layer convolution operations. The generated spectrogram not only retains the key information of the original audio, but also provides high-quality input for subsequent speech recognition. This preprocessing step ensures that the system can accurately extract useful acoustic features even in complex or noisy environments, thereby improving the robustness and accuracy of speech recognition.

[0045] Transformer model structure and its functions: Next, the generated spectrogram is passed as input to the Transformer model to generate the corresponding text transcription results. The Transformer model consists of two encoder layers and four decoder layers, and each encoder and decoder layer uses an 8-head self-attention mechanism. The encoder layer is responsible for processing the input spectrogram and capturing the contextual information in the speech signal through a multi-head self-attention mechanism to ensure that the model can understand complex speech sequences. The decoder layer further processes the encoded features and combines the encoder-decoder attention mechanism to generate the final text transcription results. This structure not only simplifies the system architecture, but also improves the model's understanding of the speech content, enabling the system to maintain high accuracy and strong robustness in a variety of accent environments.

[0046] Improve speech recognition performance: Through the above-mentioned efficient data preprocessing and powerful Transformer model structure, the system significantly enhances the performance and robustness of speech recognition. The acoustic features accurately extracted by the VGG model provide a reliable foundation for the Transformer model, while the multi-head self-attention mechanism and position encoding characteristics in the Transformer architecture effectively capture the complex information in the speech data. This approach not only simplifies the system architecture, but also improves the adaptability to diverse accents, ensuring that the system provides more reliable and accurate speech recognition results in practical applications. Especially when faced with unseen accents, the system can quickly adjust parameters and optimize recognition performance, greatly enhancing the system's generalization capabilities and user experience.

[0047] S002, applies the meta-learning algorithm to the Transformer architecture and conducts a multi-task learning framework, introduces the accent recognition task to combine the meta-learning algorithm and the multi-task learning framework, and learns to separate accent information from language content.

[0048] It is important to point out here that when applying the meta-learning algorithm to the Transformer architecture, the method also includes:

[0049] Denote the speech recognition of Transformer as f θ θ represents the parameter, and the data of each accent i is divided into and Will be in The first gradient descent update is performed above, updating θ to θ′, where Among them, α is the fast adaptation learning rate;

[0050] In the training task, the model f(θ i′) to improve the model’s performance in unseen categories The recognition performance on , the meta-objective is defined as: in Indicates that The above loss, collects the loss from a batch Use the obtained meta-validation loss to optimize the parameters in the test task: Where β is the learning step size in meta-learning, It is in accent A i Adaptive network on ; in fine-tuning,

[0051] That is to say, in the improved speech recognition method, the parameters of the Transformer model are designed to be grouped according to different regions and accents. Specifically, the data set for each accent is divided into a training set and a validation set. First, for each training set with a specific accent, the system performs the first gradient descent update based on the initial model parameters to generate new parameters adapted to the accent. This process introduces a fast adaptation learning rate to adjust the pace of the update to ensure that the model can respond effectively to new accent data in a short period of time. In this way, the model can quickly and initially adapt to a variety of different accent characteristics, laying the foundation for subsequent more precise optimization.

[0052] In order to further improve the recognition performance of the model on unseen accents, the method adopts a meta-learning strategy. In this framework, the system uses the knowledge gained from multiple accent training sets to optimize a more general model. The meta-goal is to make the model not only perform well on known accents, but also quickly adapt to accents that have never been seen. To this end, a meta-validation loss is defined, which is collected based on the losses in a batch and is used to evaluate and guide the optimization direction of the model parameters. In this process, a meta-learning step size is used, which determines the extent of parameter update in each iteration. This approach allows the model to significantly improve its generalization ability without adding a lot of additional computational cost, that is, it can maintain a high recognition accuracy when encountering new, unknown accents.

[0053] Finally, in the fine-tuning phase, the model continues to be optimized based on the validation set of a specific accent to ensure that its performance is optimal. The purpose of this phase is to fine-tune the model parameters to better fit the accent characteristics in the actual application environment. Through continuous testing and adjustment on the validation set, the model can learn more subtle accent differences and incorporate this knowledge into its own structure, thereby improving the final recognition effect. This fine-tuning is not only a consolidation of previous learning results, but also the final polishing process for the adaptability and robustness of the model, ensuring that the model can provide high-quality speech recognition services in a variety of accent environments.

[0054] S003, using the parameters of the recognition model trained by meta-learning and multi-task learning as a starting point, collecting and preprocessing speech samples with unseen accents, gradually adjusting the parameters of the recognition model based on the preprocessed speech samples with unseen accents, gradually improving the performance of the recognition model on unseen accents through multiple iterations, and determining the corresponding speech recognition model.

[0055] Specifically, in the multi-task learning framework, a shared encoder is designed for speech recognition and accent recognition tasks. The shared encoder is the core part of the entire model, responsible for extracting common acoustic features from the input spectrogram and passing these features to the subsequent decoder for processing. Through the shared encoder, the model is able to share information between the two tasks, promoting feature learning and generalization capabilities. Specifically, the shared encoder can capture contextual information and temporal dependencies in the speech signal, providing a unified and rich feature representation for speech recognition and accent recognition. This approach not only reduces the number of model parameters, but also enhances the synergy between tasks, making the model more efficient and accurate when dealing with diverse accents.

[0056] In addition, the multi-task learning framework further introduces a dual-branch decoder, in which one branch focuses on the speech recognition task and the other branch is responsible for the accent recognition task. This design allows the model to process two different types of information at the same time, thereby effectively separating language content and accent features. The speech recognition decoder captures the language structure in the input features through the self-attention mechanism and generates the corresponding text transcription results; while the accent recognition decoder focuses on extracting accent-related information from the same features and generating recognized accent labels. The two decoders work independently but collaborate with each other to ensure that each task is processed in a targeted manner, improving the model's ability to understand complex speech data. In addition, the design of the dual-branch decoder also allows the model to simultaneously optimize two objective loss functions during training, namely the speech recognition loss L asr_loss and accent recognition loss L accent_loss , thereby achieving a more comprehensive task performance.

[0057] Finally, by adopting a multi-task learning framework with a shared encoder and a dual-branch decoder, the system significantly improves its adaptability and robustness to diverse accents. The shared encoder promotes information sharing between different tasks and enhances the model's overall understanding of speech content and accent characteristics; while the dual-branch decoder ensures that each task can be processed efficiently, avoiding the information confusion problem that may exist in traditional methods. This architecture not only simplifies the complexity of the system, but also improves the training efficiency and final performance of the model. Especially when faced with unseen accents, the model can better separate accent information from language content, providing more reliable and accurate speech recognition results. Overall, the design of the multi-task learning framework brings higher flexibility and stronger generalization capabilities to the speech recognition system, enabling it to perform well in practical applications.

[0058] In the training process of this application, the corresponding speech recognition loss L asr_loss and accent recognition loss L accent_loss During the training process, in order to effectively separate accent information from language content and improve the overall performance of the model, the system introduces an additional overall loss function.

[0059] L loss =βL asr_loss +λL accent_loss , which is derived from the target loss L from the speech recognition task asr_loss and the target loss L from the accent recognition task accent_loss Among them, β and λ are both hyperparameters, usually they are controlled by β+λ=1, which is used to balance the importance of the two tasks. By jointly optimizing these two objective losses, the model not only focuses on improving the accuracy of speech recognition, but also effectively learns and separates accent features, thereby enhancing the generalization ability and adaptability of the system.

[0060] The setting of hyperparameters is crucial, and they are usually adjusted according to the specific application scenario and the characteristics of the dataset. In the early stages of training, more emphasis may be placed on speech recognition tasks (for example, by setting β higher) to ensure that the model first masters the core language content; as training progresses, the weight of the accent recognition task is gradually increased (such as increasing λ) so that the model can better handle diverse accents. This dynamic adjustment strategy ensures that the model significantly enhances its ability to learn accent features while maximizing the performance of the main task (speech recognition), ultimately improving the model's recognition accuracy and robustness on unseen accents.

[0061] For the training process of the above speech recognition model, refer to Figure 1 , Figure 1The speech recognition model on the right uses the standard Transformer model framework. The task of this part of the model is to perform speech recognition, and its output is based on the learned speech information without being affected by the accent. The encoder and decoder components of this part effectively capture the information in the speech sequence through self-attention and cross-attention mechanisms. Figure 1 The accent recognition model on the left shares an encoder with the speech recognition model. This design allows the accent recognition model and the speech recognition model to jointly optimize the encoder parameters during reverse gradient backpropagation, thereby achieving information sharing between the two tasks. The task of the accent recognition model is to determine the category of the accent based on the speech features. In the accent recognition task, the output of the accent recognition model is processed by introducing a mean layer to calculate the mean and variance vectors. These mean and variance vectors are then linearly transformed through the connection layer and finally passed through an argmax layer to estimate the probability of the accent. This integrated model structure enables the speech recognition task and the accent recognition task to work together on the basis of a shared encoder, realizing the recognition of the speech content while recognizing the accent information. This architecture not only improves the overall performance of the system, but also effectively utilizes the correlation between the two tasks, making the model more generalizable.

[0062] In the training phase, corresponding experiments need to be conducted. In the data preprocessing phase, the experiments in this paper used the CommonVoice dataset for iterative training. In the fine-tuning step, each sample was iteratively tested 5 times to ensure that the model can fully learn and adapt to the speech characteristics of different accents. This iterative training method helps improve the performance and robustness of the model.

[0063] During the evaluation phase, beam search was used to generate the output of the speech recognition model. Beam search is an effective sequence generation technique that considers multiple candidate word sequences and selects the most likely sequence as the final output. Here, the beam size was set to 5 to balance the breadth and depth of the search, further improving the accuracy of the model in the speech recognition task.

[0064] Two configurations, including mixed-region and cross-region settings, are used to evaluate the effectiveness of the proposed method, specifically categorized as follows: Figure 2 As shown in Figure 2, these two different experimental settings simulate two real-life scenarios respectively.

[0065] First, the mixed-region setting is designed to simulate a large amount of speech data with accents, comprehensively covering various regional accents, and providing a relatively stable training environment, so that the model can more accurately determine the optimal parameters during the learning process. The validation set is included in the training set with accents to simulate known categories of accents, so as to more accurately determine the optimal model parameters. In this case, the dataset covers a variety of accents, and the domain gap is relatively narrow, resulting in a relatively stable training process.

[0066] In contrast, the cross-region configuration attempts to use a limited amount of speech data with accents, covering only the accents of a few regions, simulating the limited diversity of accent data in actual applications, which is closer to actual application scenarios and more challenging. The validation set is used for accents in different contexts than the training set, which is closer to the situation in real life where there are a large number of unseen accents. Therefore, the diversity of accents in the training set is limited, and the domain gap is significant, making the training more susceptible to fluctuations, but closer to the actual situation. In short, the difference between the two configurations lies in the size of the training set data, the diversity or domain gap of accents, whether the training set contains a validation set, and whether the validation set simulates accents of known or unknown categories.

[0067] The design of these two configurations not only considers the data distribution of the accent recognition task, but also fully considers the relationship between the validation set and the training set to more comprehensively evaluate the performance of the model in different scenarios. Therefore, by deeply analyzing these two configurations, we can more comprehensively understand the robustness and generalization ability of the proposed method.

[0068] During the testing of accent data, in order to ensure reliability and robustness, we adopted a strict data partitioning and multiple experiments strategy. First, the accent data involved in the test was divided into a training set and a test set, where the training set accounted for 75% of the total data and the test set accounted for 25%. This division helps to ensure that the training set contains enough samples to train the model, while the size of the test set is sufficient to fully evaluate the performance of the model. Adaptive fine-tuning training is adopted in the few-trial learning and full-sample conditions, where the training set data participates in the adaptive fine-tuning of unseen accents. The accent speech recognition model was fine-tuned under several different adaptive training sample size settings, including zero-trial learning, 5% and 25% few-trial learning, and full-sample scenarios. This setting is designed to evaluate the adaptability of the model to unseen accents and explore the performance changes under different adaptive sample sizes.

[0069] The performance evaluation standard uses the Character Error Rate (CER), which is one of the commonly used evaluation indicators in the field of accented speech recognition, and expresses its performance test results in the form of average and standard deviation. In order to ensure the accuracy and robustness of the experimental results, the experiment was repeated 10 times using different test data, and each experiment randomly sampled 100 data from the test set of the corresponding accent. This strategy of multiple experiments helps to reduce the impact of randomness on the experimental results and improve the stability of the evaluation.

[0070] The speech data is divided into a training set accounting for 75% of the total data and a test set accounting for 25% of the total data, with a ratio of 3:1 for each accent. The training set is used to perform adaptive fine-tuning for unknown accents in the case of few samples and full samples. The character error rate (CER) is used as the evaluation indicator, and its calculation formula is as follows: Where I, S and D represent the number of inserted, substituted and deleted characters respectively, and N is the number of characters in the real data. The experiment was repeated ten times, using different test data sets to ensure the reliability and accuracy of the results. Each test data set includes 100 data randomly selected from the test set of the corresponding accent. The mean and standard deviation are used to show the performance results under different sample size settings, including 0%-shot (zero sample), 5%-shot, 25%-shot and 100%-shot (full sample).

[0071] For comparative experiments and result analysis under different regional settings, refer to Figure 3 and Figure 4 , Figure 3 Comparison of the results of three experimental methods under mixed region settings. Figure 4 This is a graph of the accent distribution ratio for mixed area settings.

[0072] exist Figure 3 The performance results of three experimental methods under a mixed regional setting are compared in the figure. The horizontal axis represents the accents of Liaoning, Anhui, Zhejiang, Sichuan, and Guangxi, and the vertical axis represents different experimental methods and different data amounts under the same experimental method. CER is used as the evaluation indicator in the table. The lower the value, the better the experimental effect. Each group of data is the result of taking the average and standard deviation of ten groups of data.

[0073] In the mixed region experiment, the audio duration involved in the training was about 52 hours. By comparing the results of the three experimental methods, it was observed that as the amount of fine-tuning data increased, the overall error rate in each region showed a continuous downward trend, thus verifying the effectiveness of the three models in speech recognition tasks. Figure 4From the accent distribution ratio chart under the mixed regional setting, we can see that Zhejiang Province has the largest number of fine-tunings, and Anhui Province has the least number of fine-tunings.

[0074] By combining the above experimental tables and distribution graphs, and comparing the experimental results of Liaoning Province and Zhejiang Province, it is found that although the training set of Zhejiang Province at 100% data volume is larger than that of Liaoning Province, the fine-tuning effect is not as good as that of Liaoning Province. This has triggered thinking about the relationship between the difficulty of accent and recognition effect. The data shows that the accent information in Liaoning Province is relatively small and relatively simple, so the fine-tuning effect is significant. In the case of Zhejiang Province, due to the more complex accent information, the fine-tuning effect is not as significant, which suggests that there is a significant correlation between the difficulty of accent and recognition effect. In addition, the results of Anhui and Guangxi are relatively poor. The main reason may be that the amount of fine-tuning data in these two provinces is small, and their accents are relatively difficult. This situation increases the difficulty of recognition, resulting in relatively poor performance of the model in these two regions.

[0075] Overall, these observations not only highlight the impact of the amount of fine-tuning data on model performance, but also emphasize the importance of the difficulty of accents in speech recognition tasks. By deeply analyzing the performance of different regions, we can better understand the impact of accents on model adaptability and provide useful guidance for future accent recognition research.

[0076] Among the three experimental methods, the multi-task meta-learning method proposed in this paper shows the best results. However, it is worth noting that in the experiments in Zhejiang and Guangxi, the effect of joint training is still good. Figure 7 From the error rate drop graph in , although the effect of joint training is better than that of meta-learning (MAML), its error rate drops slowly. On the contrary, the multi-task meta-learning model (MAML+Multi-task) can significantly reduce the recognition error rate of the model while maintaining the meta-learning training speed.

[0077] From the above summary, it can be concluded that the multi-task meta-learning model has unique advantages in dealing with accent speech recognition tasks, especially in the accent scenarios of some regions, its performance is particularly outstanding. However, for the two regions of Zhejiang and Guangxi, the joint training method still has certain advantages in experimental results. This may be because the accent characteristics of these regions are significantly different from those of other regions, resulting in the slightly insufficient generalization performance of the multi-task meta-learning model in these scenarios. The joint training method can better adapt to these regions with special accent scenarios. For specific accent scenarios, it is particularly important to choose a suitable training method. In specific regional scenarios, joint training may still be an effective choice, but combined with the overall perspective, the multi-task meta-learning model performs better overall.

[0078] exist Figure 5 The performance results of three experimental methods under a cross-regional setting are compared in the table, where the horizontal axis represents the accents of Heilongjiang, Liaoning, Anhui, Zhejiang, Sichuan, and Guangxi, and the vertical axis represents different experimental methods and different amounts of data under the same experimental method. CER is used as the evaluation indicator in the table. The lower the value, the better the experimental effect. Each group of data is the result of taking the average and standard deviation of ten groups of data.

[0079] In the cross-region setting, the audio involved in the training is about 30 hours long, the amount of data is reduced compared to the mixed region, and the experimental region is adjusted. By comparing the results of the above three experimental methods, it is found that as the amount of fine-tuning data increases, the error rate in most regions continues to decline. It is worth noting that although Liaoning Province has the lowest error rate, its amount of fine-tuning data is not the least. Compared with the amount of fine-tuning data and test error rate in Zhejiang Province, it can be more clearly seen that the difficulty of accents in different regions varies.

[0080] Such a comparison reveals the impact of regional differences on model performance in accent recognition tasks. In Liaoning Province, although there is relatively more fine-tuning data, the model performs well in this region due to the relatively simple accent. In contrast, although the amount of fine-tuning data in Zhejiang Province is large, the performance of the model in this region is relatively low due to the more complex accent. This shows that the difficulty of the accent is an important factor affecting the performance of the model and needs to be fully considered during the training process. Overall, these results provide data support for a deeper understanding of the impact of regional differences in accent recognition tasks, which helps to optimize the model to better adapt to the characteristics of accents in different regions and improve the overall performance of accent recognition.

[0081] Comparing the overall data of joint training and meta-learning, it is observed that in areas with more complex accents, such as Anhui, Zhejiang, and Sichuan, when the training data is reduced to a certain extent, the advantages of meta-learning begin to emerge. Especially in situations where the accent is difficult and the amount of data is small, meta-learning shows a more obvious advantage by learning the global optimal parameters. This shows that meta-learning is more adaptable to limited training data when dealing with complex accent situations.

[0082] After introducing the meta-learning algorithm, by comparing the experimental data, it was found that when the amount of data decreased, the error rate of joint training increased significantly compared to the error rate in the mixed area. This may be because the model is prone to overfitting when faced with complex accents and limited data. On the contrary, although the two methods that added meta-learning have improved the data error rate, the overall fluctuation is relatively stable. Meta-learning better copes with the challenges brought by limited data by learning the ability to adapt to new tasks, so that the model remains relatively robust in the scenario of few-sample learning. These observations further strengthen the importance of meta-learning in accented speech recognition and provide a strong basis for choosing appropriate training methods in specific accent situations. In situations with high accent complexity and scarce data, the introduction of meta-learning not only improves the performance of the model, but also improves the model's generalization ability for limited data.

[0083] The multi-task meta-learning model based on meta-learning using a multi-task approach achieved the best results in all six experimental areas. Figure 5 As shown in the last row of , it further proves that by adding the accent recognition task, the model can better learn the accent-related information, thereby more effectively eliminating the interference of accent information in the speech recognition task and achieving better experimental results. Figure 6 The accent distribution ratio diagram under the cross-regional setting shows that the experimental results emphasize the effectiveness of the multi-task meta-learning model. By integrating the accent recognition task into the meta-learning framework, the model can more comprehensively understand and utilize accent information while learning the relevance of tasks, thereby improving overall performance. This has important implications for the complex situations in actual accent speech recognition applications and provides strong guidance for further improving model design and training strategies.

[0084] according to Figure 7 As shown in the figure, the error rate reduction diagrams of joint training (Joint), meta-learning (MAML) and multi-task meta-learning (MAML+Mulit-task) experimental methods under two different experimental settings are respectively used. The control variable is the test effect under the zero-sample data volume of Liaoning Province. In the figure, the horizontal axis represents the number of training steps (Iteration), and the vertical axis represents the word error rate (CER). Different experimental methods are distinguished by different linearity and colors.

[0085] In the mixed region, it is observed that the Transformer model using only the meta-learning (MAML) algorithm exhibits the fastest training speed, but does not achieve the optimal effect in the final fitting process. In contrast, the joint training (Joint) model has the slowest decline rate among the three methods, but performs well in the final fitting effect. This is reflected in the accented speech recognition task, where only using meta-learning has an advantage in speed, but its fitting ability is relatively weak, which reduces the training effect.

[0086] The multi-task meta-learning model (MAML+Multi-task) proposed in this paper retains the fitting effect of the Transformer model to the greatest extent, and combines the training speed of the MAML algorithm, so that it can achieve the same training effect as other methods at a faster speed. This means that in accented speech recognition, the introduction of multi-task meta-learning can balance the training speed and model performance to a certain extent, and can provide more flexible options for practical applications.

[0087] In the cross-regional setting, due to the reduction in the amount of data, different models have different degrees of increase in speech recognition error rate, which shows that the performance of the model is limited when faced with limited data, especially when there are large differences in accents across regions. With a small data set, the multi-task model under the MAML algorithm can achieve a lower speech recognition error rate while maintaining the same training speed as the MAML algorithm. This highlights the superiority of the multi-task model in the limited data scenario, and by comprehensively considering the information of different tasks, it improves the robustness of accented speech recognition.

[0088] After analyzing the above experimental results, we concluded that the multi-task meta-learning model can continue to reduce the error rate of accented speech recognition while improving the training speed. This not only provides substantial guidance for the efficient training of speech recognition models, but also maintains the accurate capture and learning of accent features.

[0089] An adaptive accent speech recognition method based on feature decoupling, referring to Figure 8 ,include:

[0090] S100, taking the speech recognition model obtained through multi-task meta-learning adaptive training as a starting point, initializing the model parameters of the speech recognition model, and fine-tuning the model parameters of the speech recognition model, and deploying the fine-tuned speech recognition model in an actual application environment.

[0091] Among them, combined with the above-mentioned training process, by collecting large-scale speech datasets containing multiple accents, we ensure that the widest possible range of accent types are covered. These data are used to train the initial Transformer architecture model, and a multi-task learning framework is adopted to enable the model to simultaneously optimize speech recognition and accent recognition tasks.

[0092] Introducing a meta-learning algorithm during the training process enhances the model's generalization ability to unseen accents by enabling it to quickly adjust model parameters between different tasks; specifically, meta-learning allows the model to quickly update its parameters on a small number of new samples to adapt to new accent characteristics. Using a model obtained through multi-task meta-learning adaptive training as a starting point to initialize its parameters, the model has been extensively trained on data from a variety of accents and has good basic performance.

[0093] Finally, the model parameters are further fine-tuned according to the needs of the actual application scenario. This step usually uses a small-scale speech dataset in a specific scenario to optimize the model's adaptability to a specific application environment. For example, in medical or customer service scenarios, the model may be fine-tuned using professional terms and common expressions in the field.

[0094] S110, capturing the user's voice signal in an actual application environment, preprocessing the voice signal, and converting the preprocessed voice signal into a frequency spectrum to be recognized, extracting the acoustic features to be recognized through a feature processor based on the frequency spectrum to be recognized, inputting the acoustic features to be recognized into a voice recognition model, and generating a corresponding voice recognition result.

[0095] In actual application environments, the system captures the user's original voice signals through a microphone or other audio acquisition devices. These signals may include factors such as background noise, different speech speeds and volume changes; the captured voice signals are preprocessed, mainly including steps such as removing background noise, normalizing the volume and adjusting the sampling rate to improve the quality of subsequent processing. Preprocessing also includes voice activity detection (VAD) to distinguish between speech segments and non-speech segments, thereby reducing unnecessary computing resource consumption.

[0096] The preprocessed speech signal is converted into a spectrogram to be recognized using a feature processor. The feature processor uses a VGG model with a 6-layer convolutional neural network architecture, which can accurately extract acoustic features from the original audio signal and generate a two-dimensional spectrogram representation. The spectrogram provides time and frequency information about the speech and is the basis for subsequent feature extraction. Based on the generated spectrogram, the acoustic features to be recognized are further extracted by the feature processor. These features include but are not limited to Mel-frequency cepstral coefficients (MFCC), filter bank energy, etc. They capture the key information in the speech signal and provide high-quality input for subsequent speech recognition.

[0097] S120, in the process of generating the corresponding speech recognition result, the Transformer architecture encoder layer in the speech recognition model processes the input acoustic features to be recognized to capture the contextual information in the speech signal, and the dual-branch decoder performs speech recognition tasks and accent recognition tasks on the acoustic features to be recognized respectively, generates recognition text output and recognition accent labels for the acoustic features to be recognized, uses a decoding algorithm based on the contextual information to generate the final text transcription result, and combines the text transcription result with the recognition text output and recognition accent label to generate the corresponding speech recognition result.

[0098] The extracted acoustic features are then fed into the Transformer model and first processed by two encoder layers. Each encoder layer uses an 8-head self-attention mechanism to effectively capture contextual information in the speech signal. The self-attention mechanism allows the model to focus on different parts of the input sequence in parallel, thereby enhancing the understanding of complex contexts.

[0099] The dual-branch decoder is responsible for speech recognition and accent recognition tasks respectively. One branch focuses on generating recognized text output, while the other branch generates recognized accent labels. This design enables the model to process language content and accent features at the same time, ensuring the separation between the two. Each decoder layer also uses an 8-head self-attention mechanism and introduces an encoder-decoder attention mechanism, which enables the decoder to refer to the encoder's output to better understand the content of the input speech.

[0100] Based on the contextual information captured by the encoder layer, a decoding algorithm is used to generate preliminary text transcription results. In this process, the model not only considers the content of the speech, but also combines accent tags to ensure that the generated text is more accurate and natural.

[0101] Finally, the text transcription results are combined with the recognition text output and the recognition accent label to generate the corresponding speech recognition results. This comprehensive processing method improves the overall performance of the system, especially when facing speech with multiple accents.

[0102] Take a case study as an example. A multinational medical technology company developed an intelligent voice assistant that aims to provide efficient and accurate medical record services to doctors and nurses around the world. Due to the increasingly obvious internationalization trend of medical services, this voice assistant needs to be able to handle English voice input with multiple accents from different countries and regions. In order to meet this demand, the company adopted an adaptive accent speech recognition method based on feature decoupling to ensure that the system can still maintain high accuracy and robustness when facing diverse accents.

[0103] 1. First, we built a large speech dataset with multiple accents, covering English speech samples with different accents from multiple countries and regions such as the United States, the United Kingdom, India, and Australia. These data were used to train the initial Transformer architecture model, and a multi-task learning framework was adopted to enable the model to optimize speech recognition and accent recognition tasks simultaneously. A meta-learning algorithm was introduced during the training process, which enhanced the model's generalization ability to unseen accents by enabling it to quickly adjust model parameters between different tasks. Specifically, meta-learning allows the model to quickly update its parameters on a small number of new samples to adapt to new accent characteristics. After multi-task meta-learning adaptive training, the model has been extensively trained on data with multiple accents and has good basic performance.

[0104] 2. Deploy the extensively trained model to the actual application environment and further fine-tune the model parameters according to the needs of the medical scenario. For example, in the medical field, voice assistants need to be able to recognize specific professional terms and common expressions, such as "electrocardiogram" (ECG), "intravenous injection" (IV), etc. To this end, the company collected a small-scale medical-specific voice data set to optimize the model in a targeted manner to better adapt it to the needs of speech recognition in the medical environment. This fine-tuning not only improves the model's recognition accuracy for professional terms, but also enhances its performance in complex medical conversations.

[0105] 3. In the actual application environment, the voice assistant captures the user's original voice signal through a microphone. These signals may include factors such as background noise, different speaking speeds and volume changes. In order to improve the quality of subsequent processing, the system preprocesses the captured voice signal, mainly including steps such as removing background noise, normalizing volume and adjusting sampling rate to improve the quality of subsequent processing. Preprocessing also includes voice activity detection (VAD) to distinguish between speech segments and non-speech segments, thereby reducing unnecessary consumption of computing resources. The preprocessed voice signal is converted into a spectrogram to be identified using a feature processor. The feature processor uses a VGG model with a 6-layer convolutional neural network architecture, which can accurately extract acoustic features from the original audio signal and generate a two-dimensional spectrogram representation. The spectrogram provides time and frequency information about the speech and is the basis for subsequent feature extraction.

[0106] 4. Next, the generated spectrogram is passed as input to the Transformer model to generate the corresponding text transcription results. The Transformer model consists of two encoder layers and four decoder layers, and each encoder and decoder layer uses an 8-head self-attention mechanism. The encoder layer is responsible for processing the input spectrogram and capturing the contextual information in the speech signal through a multi-head self-attention mechanism to ensure that the model can understand complex speech sequences. The dual-branch decoder is responsible for speech recognition tasks and accent recognition tasks respectively, with one branch focusing on generating recognition text output and the other branch generating recognition accent labels. This design enables the model to process language content and accent features at the same time, ensuring the separation between the two. Finally, based on the contextual information captured by the encoder layer, the decoding algorithm is used to generate preliminary text transcription results, and the final speech recognition results are generated by combining the recognition text output and the recognition accent label.

[0107] Through the application of the above technical solutions, the intelligent voice assistant has significantly improved its accuracy and robustness in dealing with diverse accents. In particular, when faced with unseen accents, the system can quickly adapt and provide high-quality speech recognition services. For example, at an international medical conference, doctors from different countries used the voice assistant to record the meeting content, and the system successfully recognized various English voices with local accents with an accuracy rate of more than 95%. In addition, the voice assistant performs well in the recognition of professional terms in the medical field, greatly reducing the manual recording workload of doctors and nurses and improving work efficiency and service quality.

[0108] An adaptive accent speech recognition device based on feature decoupling, referring to Fig. 9 , the devices include but are not limited to:

[0109] The model initialization module 200 uses the speech recognition model obtained through multi-task meta-learning adaptive training as a starting point to initialize the model parameters of the speech recognition model, fine-tune the model parameters of the speech recognition model, and deploy the fine-tuned speech recognition model in an actual application environment;

[0110] The speech recognition result generation module 210 captures the user's speech signal in the actual application environment, preprocesses the speech signal, and converts the preprocessed speech signal into a frequency spectrum to be recognized, extracts the acoustic features to be recognized through a feature processor based on the frequency spectrum to be recognized, inputs the acoustic features to be recognized into the speech recognition model, and generates the corresponding speech recognition result;

[0111] The text transcription result generation module 220, in the process of generating the corresponding speech recognition result, the Transformer architecture encoder layer in the speech recognition model processes the input acoustic features to be recognized to capture the contextual information in the speech signal, and the dual-branch decoder performs the speech recognition task and the accent recognition task on the acoustic features to be recognized respectively, generates the recognition text output and the recognition accent label for the acoustic features to be recognized, uses the decoding algorithm based on the context information to generate the final text transcription result, and combines the text transcription result with the recognition text output and the recognition accent label to generate the corresponding speech recognition result.

[0112] An embodiment of the present application further discloses a method for adaptive accent speech recognition based on feature decoupling, comprising a processor in which a program of any one of the above-mentioned methods for adaptive accent speech recognition based on feature decoupling is run.

[0113] The embodiment of the present application further discloses a storage medium storing a program of any one of the above-mentioned adaptive accent speech recognition methods based on feature decoupling.

[0114] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. An adaptive accent speech recognition method based on feature decoupling, characterized in that: include: A speech recognition model that has been obtained through multi-task meta-learning adaptive training is used as a starting point, model parameters of the speech recognition model are initialized, and the model parameters of the speech recognition model are fine-tuned, and the fine-tuned speech recognition model is deployed in an actual application environment; Capturing a user's voice signal in an actual application environment, preprocessing the voice signal, and converting the preprocessed voice signal into a to-be-recognized frequency spectrum, extracting to-be-recognized acoustic features through a feature processor based on the to-be-recognized frequency spectrum, inputting the to-be-recognized acoustic features into a voice recognition model, and generating a corresponding voice recognition result; In the process of generating the corresponding speech recognition results, the Transformer architecture encoder layer in the speech recognition model processes the input acoustic features to be recognized to capture the contextual information in the speech signal, and the dual-branch decoder performs speech recognition tasks and accent recognition tasks on the acoustic features to be recognized respectively, generates recognition text output and recognition accent labels for the acoustic features to be recognized, and uses a decoding algorithm based on the contextual information to generate the final text transcription result, and combines the text transcription result with the recognition text output and recognition accent label to generate the corresponding speech recognition result.

2. The method for adaptive accent speech recognition based on feature decoupling according to claim 1, characterized in that: The speech recognition model training process includes: Collect speech data with different accents, convert the speech data into spectrograms, and use feature processors to extract acoustic features from the spectrograms. Select the Transformer architecture as the recognition model, initialize the parameters of the encoder and decoder in the Transformer architecture, and set the configuration of the recognition model. The multi-head attention mechanism and position encoding characteristics of the Transformer architecture can effectively capture the information in the speech data. Apply the meta-learning algorithm to the Transformer architecture and conduct a multi-task learning framework. Introduce the accent recognition task to combine the meta-learning algorithm and the multi-task learning framework to learn to separate accent information from language content. Using the parameters of the recognition model trained by meta-learning and multi-task learning as the starting point, speech samples with unseen accents are collected and preprocessed. Based on the preprocessed speech samples with unseen accents, the parameters of the recognition model are gradually adjusted. Through multiple iterations, the performance of the recognition model on unseen accents is gradually improved, and the corresponding speech recognition model is determined.

3. The method for adaptive accent speech recognition based on feature decoupling according to claim 2, characterized in that: In the process of applying the meta-learning algorithm to the Transformer architecture, the method further includes: Denote the speech recognition of Transformer as f θ θ represents the parameter, and the data of each accent i is divided into and Will be in The first gradient descent update is performed above, updating θ to θ′, where Among them, α is the fast adaptation learning rate; In the training task, the model f(θ i ′) to improve the model’s performance in unseen categories The recognition performance on , the meta-objective is defined as: in Indicates that The above loss, collects the loss from a batch Use the obtained meta-validation loss to optimize the parameters in the test task: Where β is the learning step size in meta-learning, It is in accent A i Adaptive networks on In fine-tuning, 4. The method for adaptive accent speech recognition based on feature decoupling according to claim 1, characterized in that: In the process of the multi-task learning framework, the method also includes: The multi-task learning framework includes a shared encoder and a dual-branch decoder. The shared encoder is used for both speech recognition and accent recognition. One branch decoder in the dual-branch decoder is used for speech recognition, and the other branch decoder is used for accent recognition.

5. The method for adaptive accent speech recognition based on feature decoupling according to claim 1, characterized in that: In the process of introducing the accent recognition task, the method also includes: During the training process, an additional overall loss is introduced, the encoder and the target loss L from speech recognition asr_loss , target loss L from accent recognition accent_loss Joint optimization, overall loss L loss =βL asr_loss +λL accent_loss , where β and λ are hyperparameters, and they are usually limited by β+λ=1.

6. The method for adaptive accent speech recognition based on feature decoupling according to claim 1, characterized in that: Also comprising a data preprocessing step, the method further comprises: Receive the user's original audio input and convert it into a spectrogram to be identified through a feature processor, where the feature processor uses a VGG model with a 6-layer convolutional neural network architecture. The VGG model is used to extract acoustic features from the original audio signal and generate a two-dimensional spectrogram; The generated spectrogram is passed as input to the Transformer model to generate the corresponding text transcription result. The Transformer model consists of two encoder layers and four decoder layers. Each encoder and decoder layer uses an 8-head self-attention mechanism.

7. An adaptive accent speech recognition device based on feature decoupling, characterized in that: include: The model initialization module uses the speech recognition model obtained through multi-task meta-learning adaptive training as a starting point to initialize the model parameters of the speech recognition model, fine-tune the model parameters of the speech recognition model, and deploy the fine-tuned speech recognition model in the actual application environment; A speech recognition result generation module captures the user's speech signal in an actual application environment, preprocesses the speech signal, and converts the preprocessed speech signal into a to-be-recognized spectrogram, extracts the to-be-recognized acoustic features through a feature processor based on the to-be-recognized spectrogram, inputs the to-be-recognized acoustic features into a speech recognition model, and generates a corresponding speech recognition result; The text transcription result generation module, in the process of generating the corresponding speech recognition result, the Transformer architecture encoder layer in the speech recognition model processes the input acoustic features to be recognized to capture the contextual information in the speech signal, and the dual-branch decoder performs the speech recognition task and the accent recognition task on the acoustic features to be recognized respectively, generates the recognition text output and the recognition accent label for the acoustic features to be recognized, uses the decoding algorithm based on the context information to generate the final text transcription result, and combines the text transcription result with the recognition text output and the recognition accent label to generate the corresponding speech recognition result.

8. An adaptive accent speech recognition method based on feature decoupling, characterized in that: It comprises a processor, in which runs a program of the adaptive accent speech recognition method based on feature decoupling as described in any one of claims 1 to 6.

9. A storage medium, characterized in that: A program for the adaptive accent speech recognition method based on feature decoupling as described in any one of claims 1 to 6 is stored.

Citation Information

Cited By

  • Streaming speech recognition method for mandarin with accent in electric power major

    CN121393449A