A speech recognition method based on multi-dialect adaptive fusion
Patent Information
- Application Number
- CN202610740696.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]然而,现有的多方言语音识别方法中,直接采用并行子系统架构,并没有考虑计算资源与部署成本的优化,由此可能会导致系统参数量随方言数量线性增长,从而影响终端侧部署的可行性
本方案采用共享编码器搭配方言适配器架构,编码器可完全复用,新增方言只需训练少量适配参数,相比独立模型有效降低部署成本、数据需求与训练耗时,且支持不停机热插拔升级。通过多重优化机制提升老年方言识别准确率,依靠注意力门控实现方言与普通话无缝交替识别,避免语音切分误差。系统依托EWC策略持续优化识别精度、杜绝模型遗忘,训练效率更高,模型经压缩优化后体积小巧,可适配主流终端设备部署。
Smart Images

Figure CN122781284A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition and natural human-computer interaction technology, and in particular to a speech recognition method based on multi-dialect adaptive fusion. Background Technology
[0002] Speech recognition, as a core technology for natural human-computer interaction, is widely used in smart homes, in-vehicle terminals, and mobile devices. Among related technologies, multi-dialect speech recognition systems typically employ an architecture where multiple complete dialect recognition subsystems operate in parallel. Through the collaborative operation of speech activity detection, acoustic feature extraction, and decoders, a complete processing flow from speech acquisition to text output is constructed.
[0003] However, existing multi-dialect speech recognition methods directly employ a parallel subsystem architecture without considering the optimization of computational resources and deployment costs. This may lead to a linear increase in the number of system parameters with the number of dialects, thus affecting the feasibility of deployment on the terminal side. Furthermore, the block processing method relies on the accuracy of speech segment boundary determination, which is insufficiently adaptable to scenarios where dialects and Mandarin seamlessly switch in natural conversation. It also lacks a systematic adaptation mechanism for the speech characteristics of specific groups (such as the elderly), resulting in a significant decrease in recognition accuracy in complex application scenarios. Summary of the Invention
[0004] The main objective of this invention is to provide a speech recognition method based on multi-dialect adaptive fusion.
[0005] Another objective of this invention is to propose a speech recognition device based on multi-dialect adaptive fusion.
[0006] The third objective of this invention is to provide an electronic device.
[0007] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.
[0008] To achieve the above objectives, a first aspect of the present invention proposes a speech recognition method based on multi-dialect adaptive fusion, comprising: S1, collects user voice signals, performs activity detection on the voice signals, and extracts valid voice segments; S2 automatically detects the dialect category of valid speech segments and outputs the corresponding dialect category and confidence level. S3, load the corresponding dialect adapter according to the dialect category, input the acoustic features of the effective speech segment into the shared encoder, extract the general acoustic representation, and perform dialect-specific feature transformation on the general acoustic representation through the dialect adapter; S4 inputs the acoustic representation after feature transformation to the decoder, dynamically calculates the weights of each dialect adapter through an attention gating mechanism, and outputs the recognized text based on the weighted fused features.
[0009] Optionally, the user's voice signal is acquired, and activity detection is performed on the voice signal to extract valid voice segments, including: The effective speech segments are classified into one of the seven major dialect regions using a lightweight CNN classifier, and the first confidence level is output. When the first confidence score exceeds a preset threshold, a lightweight fine classifier is deployed in the corresponding large dialect area, and prosodic features are introduced as auxiliary features to perform second-level fine-grained dialect subclass recognition, and the corresponding dialect category and second confidence score are output.
[0010] Optionally, the dialect adapter is trained using the LoRA low-rank adaptation technique. A low-rank decomposition matrix is injected into the attention layer of the shared encoder, the original model parameters of the shared encoder are frozen, and only the low-rank adaptation parameters are trained. The number of parameters for each dialect adapter is controlled to be around 500,000.
[0011] Optionally, the acoustic features of valid speech segments are input into the shared encoder, including: 80-dimensional FBank acoustic features were extracted at a sampling rate of 16kHz. The FBank acoustic features were then input into a shared encoder containing 12 Conformer Blocks. Each Conformer Block consists of a multi-head self-attention module, a one-dimensional CNN convolution module, a feedforward network module, and layer normalization to extract a general acoustic representation.
[0012] Optionally, the weights of each dialect adapter can be dynamically calculated using an attention gating mechanism, including: A gating network is introduced into the decoder. The gating network receives the acoustic features of the current frame as input and outputs the weight vectors of each dialect adapter through a two-layer fully connected network. After Softmax normalization, the weight coefficients of each dialect adapter are obtained. The input dimension of the gating network is the feature dimension output by the shared encoder, and the output dimension is the weight vector dimension of each dialect adapter.
[0013] Optional steps for adapting elderly people's voice features: The speech rate prior parameter in the language model of the decoder is adjusted to adapt the mapping relationship between frame rate and word rate to the speech rate of 120 to 150 words per minute for the elderly, and a duration extension option is added to the CTC decoding path to allow a single phoneme to correspond to a longer frame sequence. A dynamic gain control module is added after speech activity detection and before feature extraction. It automatically performs RMS normalization gain on low-volume speech signals and adopts an adaptive compression strategy to maintain the dynamic range of speech and improve the signal-to-noise ratio in the low-volume segment. Introducing pronunciation degradation modeling into the acoustic model, a pronunciation degradation feature library is established by collecting speech data from the elderly. During the training phase, the degradation features are used as an additional input channel, enabling the model to learn the mapping relationship between degradation patterns and standard pronunciation. During the inference phase, the degree of pronunciation degradation is automatically detected and corresponding compensation strategies are applied.
[0014] Optional online adaptive learning steps: Collect the recognition results that users actively correct, and pair and store the original speech segments with the corrected text; Regularly fine-tune the dialect adapters for the corresponding users using newly collected correction data; Incremental learning employs an elastic weight solidification strategy. During fine-tuning, the importance of key parameters of the shared encoder is calculated, and strong regularization constraints are applied to important parameters to avoid catastrophic forgetting of existing knowledge. The loss function of the elastic weight solidification strategy is as follows:
[0015] in, For parameter importance, These are the old parameter values.
[0016] To achieve the above objectives, a second aspect of the present invention provides a speech recognition device based on multi-dialect adaptive fusion, comprising: The acquisition module is used to acquire user voice signals, perform activity detection on the voice signals, and extract valid voice segments; The detection module is used to automatically detect the dialect category of valid speech segments and output the corresponding dialect category and confidence level. The feature transformation module is used to load the corresponding dialect adapter according to the dialect category, input the acoustic features of the effective speech segment into the shared encoder, extract the general acoustic representation, and perform dialect-specific feature transformation on the general acoustic representation through the dialect adapter; The weighted fusion module is used to input the acoustic representation after feature transformation into the decoder, dynamically calculate the weights of each dialect adapter through an attention gating mechanism, and output the recognized text based on the weighted fusion features.
[0017] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0018] To achieve the above objectives, a third aspect of this application provides an electronic device, including a processor and a memory; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, for implementing the speech recognition method based on multi-dialect adaptive fusion as described in the first aspect embodiment.
[0019] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech recognition method based on multi-dialect adaptive fusion as described in the first aspect embodiment.
[0020] The embodiments of the present invention have the following beneficial effects: This solution employs a shared encoder paired with a dialect adapter architecture. The encoder is fully reusable, and adding a new dialect requires only training a small number of adaptation parameters. Compared to independent models, this effectively reduces deployment costs, data requirements, and training time, and supports hot-swappable upgrades without downtime. Multiple optimization mechanisms improve the accuracy of elderly dialect recognition, and attention gating enables seamless alternation between dialect and Mandarin recognition, avoiding speech segmentation errors. The system continuously optimizes recognition accuracy and prevents model forgetting using the EWC strategy, resulting in higher training efficiency. The compressed and optimized model is compact and can be deployed on mainstream terminal devices. Attached Figure Description
[0021] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a speech recognition method based on multi-dialect adaptive fusion provided in an embodiment of the present invention; Figure 2 This is a structural diagram of a speech recognition device based on multi-dialect adaptive fusion, provided in an embodiment of the present invention. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] The following description, with reference to the accompanying drawings, describes a speech recognition method and apparatus based on multi-dialect adaptive fusion according to embodiments of the present invention.
[0025] Example 1 This invention provides a speech recognition method based on multi-dialect adaptive fusion, such as... Figure 1 As shown, the method includes the following steps: S1 collects the user's voice signal, performs activity detection on the voice signal, and extracts valid voice segments.
[0026] To achieve accurate far-field speech acquisition and preprocessing, and to realize accurate dialect detection at both high and low levels, adapting to elderly and weak speech recognition scenarios, this application establishes a multi-module collaborative speech preprocessing and two-level dialect pre-detection mechanism to ensure the real-time performance and accuracy of subsequent dialect recognition.
[0027] In this embodiment, a multi-channel microphone array acquisition module is configured, specifically adopting a four- to six-channel MEMS microphone array hardware structure. This hardware structure can support far-field voice pickup within a distance range of five to eight meters. At the same time, combined with beamforming algorithm, the target sound source is directionally enhanced, which can effectively suppress environmental background noise, reverberation caused by spatial transmission, and lateral interference with human voices in the scene, significantly improving the voice acquisition quality in various complex usage environments.
[0028] In this embodiment, after the system completes the voice signal acquisition, it immediately sends the acquired voice data to the voice activity detection module for processing. This application adopts a dual-criteria detection strategy that combines energy threshold and spectral entropy. Compared with the traditional single energy detection method, this detection method has a more comprehensive judgment dimension and can accurately distinguish between effective human voices and steady-state environmental noise and instantaneous noise. It effectively avoids the system from mistakenly filtering out weak effective human voices and ensures the integrity of voice signal acquisition.
[0029] In this embodiment, considering the common characteristics of elderly users, such as weak voice volume and uneven voice energy distribution, a dynamic gain control module is added after the voice activity detection process ends and before the acoustic feature extraction process begins, to adapt to the specific voice characteristics of elderly users. In this embodiment, the dynamic gain control module can automatically identify low-volume segments in the voice signal and perform RMS normalized gain processing on the low-volume voice signal. Simultaneously, it uses an adaptive compression strategy to regulate the overall dynamic range of the voice, effectively increasing the signal amplitude of low-volume voice segments, significantly improving the signal-to-noise ratio of weak voice signals, and properly solving the practical problems of low voice volume and difficulty in device pickup for elderly users.
[0030] In this embodiment, after the system completes signal gain optimization, it truncates and filters the real-time transmitted continuous voice data stream, accurately removing silent segments and invalid noise segments while retaining all valid voice segments with recognition value, thus completing the entire set of voice preprocessing work. In this embodiment, this step also incorporates a two-level dialect pre-detection mechanism, which can preliminarily determine the dialect type of the speech in advance, laying a reliable foundation for subsequent accurate dialect recognition.
[0031] In this embodiment, the system first performs first-level coarse-grained dialect classification on the extracted valid speech segments using a lightweight CNN classifier. Speech segments of three to five seconds in length are selected, and their corresponding MFCC and FBank acoustic features are extracted as model input data. Relying on a lightweight network structure that includes multi-layer convolution, batch normalization, ReLU activation, and global average pooling, the input speech segments are determined to be any one of the seven major dialect regions: Mandarin, Wu, Cantonese, Min, Hakka, Xiang, and Gan. At the same time, the corresponding first classification confidence score is output.
[0032] In this embodiment, the overall inference latency of the coarse classification process can be controlled within 200 milliseconds, fully meeting the performance requirements of real-time voice interaction. In this embodiment, the system pre-sets a confidence threshold of 0.8. When the first confidence score obtained from the first-level classification exceeds the preset threshold, the system will automatically trigger the second-level fine-grained dialect sub-class recognition process.
[0033] In this embodiment, a multi-channel microphone array combined with beamforming is used to enhance far-field speech acquisition. Dual-criteria speech activity detection and dynamic gain control are used to optimize the pickup effect of weak speech in the elderly. A two-level coarse-fine dialect detection mechanism is used to complete the initial dialect recognition, laying the foundation for subsequent automatic dialect category detection.
[0034] S2 automatically detects the dialect category of valid speech segments and outputs the corresponding dialect category and confidence level.
[0035] To accurately determine dialect categories and adapt to the speech characteristics of the elderly, this application adopts a two-level detection architecture to train the model, detect pronunciation degradation features, and reduce device resource consumption and recognition errors.
[0036] In this embodiment, the inefficient processing method of parallel recognition of fixed dialects is abandoned. Instead, a self-developed two-level dialect detection architecture is adopted. Intelligent dialect screening is carried out by first coarsely classifying major categories and then finely classifying subcategories. The corresponding dialect type is automatically matched based on the acoustic characteristics of the speech itself. This operating mode does not require staff to manually preset dialect types, nor does it require loading multiple dialect recognition models for operation at the same time. It can effectively reduce the computing power consumption of edge terminal devices and reduce the memory resource usage of devices.
[0037] In this embodiment, a dedicated training and optimization process was implemented for the dialect detection model. During the training phase, the AISHELL-4 dialect subset, the THCHS-30 general speech dataset, and a self-built multi-scenario dialect corpus were integrated, with at least one hundred hours of annotated speech samples for each dialect type. In this embodiment, the project also separately collected at least fifty hours of elderly dialect speech data, comprehensively encompassing typical speech characteristics of the elderly, such as pronunciation degeneration, slow speech rate, and blurred tones, allowing the model to adapt well to elderly dialect speech and possess stable recognition performance.
[0038] In this embodiment, the cross-entropy loss function is selected as the parameter optimization target during model training. The internal parameters of the network are continuously adjusted iteratively, so that the recognition accuracy of coarse-grained dialect classification is stably maintained at more than 85%, effectively ensuring the authenticity and reliability of the dialect category determination results.
[0039] In this embodiment of the application, considering the common phenomenon of pronunciation degeneration among elderly users, this step simultaneously performs pronunciation degeneration feature detection and modeling. Relying on the pre-built elderly pronunciation degeneration feature library, it identifies various degeneration phenomena in speech in real time, such as weakening of dental consonants, tone shift, and drift of vowel formants.
[0040] In this embodiment, a two-level dialect detection architecture and a multi-class corpus training model are used to identify speech degradation features and quantify the degree of degradation, efficiently determine the dialect type, effectively reduce recognition errors and save equipment operating resources, and lay the foundation for subsequent dialect-specific feature transformation.
[0041] S3: Load the corresponding dialect adapter according to the dialect category, input the acoustic features of the effective speech segments into the shared encoder, extract the general acoustic representation, and perform dialect-specific feature transformation on the general acoustic representation through the dialect adapter.
[0042] To achieve adaptive transformation of acoustic features from multiple dialects and adapt to elderly speech, this application relies on a shared encoder and a lightweight adapter to extract features, balancing resource utilization and system scalability while avoiding the model learning forgetting problem.
[0043] In this embodiment, the system dynamically matches and loads the corresponding dialect adapter based on the dialect determination result and the accurate dialect category. The system adopts an on-demand loading operation mode, loading only the adapter corresponding to the current recognition scene into the device memory, while the shared encoder always remains running, thereby reasonably saving the storage space and computing power of the edge terminal device.
[0044] In this embodiment, feature extraction is performed on the processed valid speech segments. Eighty-dimensional FBank acoustic features are uniformly extracted, a standard sampling frequency of 16 kHz is set, and a uniformly formatted acoustic input sequence is generated. The regularized acoustic features are then fed into a shared encoder with a built-in twelve-layer Conformer Block structure. Each Conformer Block integrates a multi-head self-attention module, a one-dimensional CNN convolution module, a feedforward network module, and a layer normalization module, which can simultaneously capture local details and global temporal information of the speech. After being processed sequentially through multiple layers, general acoustic representation data adaptable to various dialects is output.
[0045] In this embodiment, the dialect adapters used are all trained and generated using LoRA low-rank adaptation technology. During the model training phase, all the original parameters of the shared encoder are locked, and only the low-rank decomposition matrix is embedded in the attention layer. Only a small number of adaptation parameters are trained and debugged, and the number of parameters of a single dialect adapter is maintained at around 500,000, making the overall size much smaller than that of the shared encoder. Compared with the traditional method of fine-tuning the model as a whole, this lightweight adaptation method can significantly reduce the amount of training data required, effectively reduce the time spent on model training, and support flexible expansion of dialect types. When adding a new dialect type, only the corresponding adapter needs to be trained separately, without modifying the original encoder parameters or stopping the device to run the update program, effectively enhancing the system's ability to expand applications.
[0046] In this embodiment, the loaded dialect adapter performs dialect feature conversion processing on the general acoustic representation, corrects the feature differences caused by pronunciation, intonation and accent between different dialects, and performs feature compensation by combining the elderly pronunciation degradation information obtained in the previous detection, reducing the feature distortion problem caused by pronunciation defects, and finally generating refined acoustic features that fit the current dialect and the characteristics of elderly voice.
[0047] In this embodiment, online adaptive learning-related processing logic is also reserved. The system caches user voice data and corrected text content in real time to reserve materials for subsequent incremental fine-tuning of the model. Furthermore, elastic weight solidification constraint rules are pre-set to avoid the loss of the original recognition ability during subsequent parameter adjustment.
[0048] In this embodiment, common acoustic features are extracted by dynamically loading dialect adapters, and low-rank adaptation technology is used to complete dialect feature transformation and elderly speech compensation, which saves equipment resources and improves system scalability, and also lays the foundation for subsequent calculation of the weights of each dialect adapter.
[0049] S4 inputs the acoustic representation after feature transformation to the decoder, dynamically calculates the weights of each dialect adapter through an attention gating mechanism, and outputs the recognized text based on the weighted fused features.
[0050] To achieve smooth switching between dialects and Mandarin and adapt to the speech of the elderly, this application adopts a joint decoding architecture, relies on a gating mechanism to optimize the recognition effect, and continuously improves the model's recognition performance through incremental learning.
[0051] In this embodiment, the refined acoustic representation data after dialect feature transformation and degradation compensation is input into the decoder module. The decoder used in this application combines CTC and attention joint decoding architecture, which can efficiently complete the time-series modeling and text decoding related operations.
[0052] In this embodiment, to address the recognition challenges arising from frequent switching between dialects and Mandarin in daily conversations, an attention gating mechanism is implemented within the decoder, constructing a two-layer fully connected gating network. In this embodiment, the gating network uses the acoustic features output from the shared encoder as input information, sequentially passing them through ReLU activation function operations and Softmax normalization processing to calculate the weight coefficients corresponding to each dialect adapter in real time, generating a weight vector that fits the current speech scenario.
[0053]
[0054] in, The weight vectors of each dialect adapter, after Softmax normalization, satisfy... ; , These are the weight matrices for the first and second layers of the gated network, respectively (the superscripts indicate the network layer numbers, not matrix element indices). , These are the bias vectors for the first and second layers of the gated network, respectively. This is the acoustic feature vector of the current frame; To modify the activation function of the linear unit; For the first Weighting coefficients for each dialect adapter; For the second layer of the gating network Each output value, and ; For dialect adapter index, To sum and iterate through the variables; and K represents the total number of dialect adapters.
[0055] In this embodiment, the system determines whether a language switch has occurred based on the KL divergence fluctuation of acoustic features in consecutive frames. Once the KL divergence values of three consecutive frames exceed a set threshold, the dialect and Mandarin switching recognition process is initiated. The system smoothly connects different language features by using a weighted summation method to eliminate recognition gaps and boundary errors generated during the language switching process.
[0056] In this embodiment, the decoder makes adaptation adjustments to suit the characteristics of elderly speech during operation. On the one hand, it changes the prior parameters of speech rate inside the language model to match the speech rate of 120 to 150 words per minute for the elderly, and optimizes the correspondence between frame rate and word rate.
[0057] In this embodiment, a duration extension option is added to the CTC decoding path to support matching individual phonemes with longer frame sequences, which is more in line with the speaking characteristics of elderly users who have a slower pronunciation rhythm and longer syllable durations. In this embodiment, the system performs decoding calculations based on the high-quality acoustic features after weighted fusion, ultimately generating accurate and reliable speech recognition text.
[0058] In this embodiment of the application, this step also constructs an online adaptive learning optimization closed loop. When the user manually modifies the recognized content, the system will match the original speech segment with the corrected text and store it locally using encryption.
[0059] In this embodiment, when the device is idle, the system retrieves newly added labeled data and performs incremental fine-tuning on the dialect adapter corresponding to the user. In this embodiment, the model fine-tuning stage utilizes an elastic weight-based EWC strategy throughout, statistically analyzes the importance of various parameters of the shared encoder, applies regularization constraints to core parameters, and relies on a dedicated loss function to balance new knowledge learning with the retention of existing knowledge.
[0060] in, For parameter importance (diagonal elements of the Fisher information matrix). These are the old parameter values. This processing method allows the model to continuously adapt to the user's individual pronunciation habits and prevents the model from losing its original dialect recognition ability, so that the system's recognition accuracy can gradually improve over time.
[0061] In this embodiment, a joint decoding architecture and attention gating mechanism are used to achieve smooth language recognition, adapt the output of recognized text to the speech characteristics of the elderly, and use an incremental learning strategy to iteratively optimize the model's recognition capabilities.
[0062] In specific embodiments of the present invention, the terminal device may adopt the following hardware configuration: Main control chip: SoC supporting NPU acceleration, with NPU computing power of no less than 6 TOPS.
[0063] Memory: ≥4GB LPDDR4, used to load the shared encoder and the current dialect adapter.
[0064] Storage: ≥32GB eMMC or TF card for storing the shared encoder and all dialect adapters.
[0065] Microphone array: 4-6 channel MEMS microphone array, supporting beamforming and echo cancellation.
[0066] Speaker: Used for voice feedback output.
[0067] Communication module: Wi-Fi / Bluetooth module, used for online data synchronization and OTA updates.
[0068] Step 1: Data collection and preprocessing.
[0069] We collected speech data from seven major dialect regions, referencing publicly available datasets such as AISHELL-4 dialect subsets, THCHS-30, and our self-built corpus.
[0070] At least 100 hours of annotated speech data were collected for each dialect, covering samples from different age groups, genders, and speaking styles.
[0071] Speech data of the elderly are collected separately, for at least 50 hours per dialect, to ensure full coverage of pronunciation degeneration features.
[0072] Step 2: Data Augmentation.
[0073] Add noise to the raw data: add white noise, room reverberation, background voices, etc., with a signal-to-noise ratio of 5~20dB.
[0074] Variable speed processing: speed variation from 0.8x to 1.2x to simulate different speech speed scenarios.
[0075] Volume variation: ±6dB gain variation to simulate different sound intensities.
[0076] SpecAugment: Performs time and frequency masking in the spectral dimension to enhance model robustness.
[0077] Step 3: Shared encoder pre-training.
[0078] We use data from all dialects to jointly pre-train a Conformer encoder to learn a universal acoustic representation across dialects.
[0079] Training objective: CTC loss function, AdamW optimizer, initial learning rate 1e-3, and cosine annealing learning rate scheduling.
[0080] Training epochs: approximately 100 epochs, until the validation set loss converges.
[0081] Freeze the shared encoder parameters after training is complete.
[0082] Step 4: Dialect adapter training.
[0083] For each dialect, the shared encoder parameters are fixed, and only the corresponding Adapter layer is trained.
[0084] The adapter uses LoRA low-rank adaptation technology, with rank r=8 and α=16.
[0085] Training data: Annotated data corresponding to the dialect, approximately 100 hours.
[0086] Training epochs: Approximately 20-30 epochs.
[0087] Repeat this step when adding a new dialect; there is no need to repeat step three.
[0088] Step 5: Dialect detector training.
[0089] A lightweight CNN classifier was trained using dialect speech data.
[0090] Input: FBank features of a 3-5 second speech segment.
[0091] Network structure: 3 convolutional layers (32 / 64 / 128 channels) + global average pooling + fully connected layers (7 outputs).
[0092] Training objective: Cross-entropy loss.
[0093]
[0094] in, For dialect category index, C represents the total number of dialect categories; The first one-hot encoding of the real dialect label 1 element, and ; The first output of the classifier The dialect prediction probability is output by the Softmax layer.
[0095] The accuracy on the training and validation sets reached over 85%.
[0096] Step Six: Gated Network Training.
[0097] The gating network was trained using speech data that included alternating dialect and Mandarin pronunciation.
[0098] Training objective: Minimize the CTC loss of the weighted adapter combination.
[0099] Training data can be obtained by artificially synthesizing dialect-Mandarin alternating corpora or by training on real mixed corpora.
[0100] Step 7: Online adaptive optimization.
[0101] Collect user correction data after deployment.
[0102] Regularly fine-tune the corresponding dialect adapter and adopt the EWC strategy to avoid catastrophic forgetting.
[0103] The EWC regularization coefficient λ was determined through optimization using the validation set.
[0104] When the system is running, user voice input is processed according to the following procedure: A microphone array collects speech signals, and beamforming enhances the target sound source.
[0105] The VAD module detects speech activity and extracts valid speech segments.
[0106] The dialect automatic detection module determines the dialect category of speech segments and outputs the dialect category and confidence level.
[0107] Based on the determination result, load the corresponding dialect adapter from storage into memory (skip if it has already been loaded).
[0108] Extract FBank features and input them into the shared encoder to extract a general acoustic representation.
[0109] Dialect-specific feature transformations are performed using a dialect adapter.
[0110] In the decoder, the gating network dynamically calculates the weights of each dialect adapter and achieves smooth switching through weighted summation.
[0111] CTC / attention joint decoding outputs recognized text.
[0112] If the user makes a correction, the correction data is saved for subsequent online adaptive learning.
[0113] Example 2 This invention provides a speech recognition device 10 based on multi-dialect adaptive fusion, such as... Figure 2 As shown, the device includes: The acquisition module 100 is used to acquire user voice signals, perform activity detection on the voice signals, and extract valid voice segments; The detection module 200 is used to automatically detect the dialect category of valid speech segments and output the corresponding dialect category and confidence level. The feature transformation module 300 is used to load the corresponding dialect adapter according to the dialect category, input the acoustic features of the effective speech segment into the shared encoder, extract the general acoustic representation, and perform dialect-specific feature transformation on the general acoustic representation through the dialect adapter. The weighted fusion module 400 is used to input the acoustic representation after feature transformation into the decoder, dynamically calculate the weights of each dialect adapter through an attention gating mechanism, and output the recognized text based on the weighted fusion features.
[0114] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0115] Example 3 To implement the methods of the above embodiments, the present invention also provides an electronic device, which includes a memory and a processor; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the various steps of the methods described above.
[0116] Example 4 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.
[0117] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0118] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0119] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A speech method based on multi-dialect adaptive fusion, characterized in that, include: S1, collects user voice signals, performs activity detection on the voice signals, and extracts valid voice segments; S2 automatically detects the dialect category of valid speech segments and outputs the corresponding dialect category and confidence level. S3, load the corresponding dialect adapter according to the dialect category, input the acoustic features of the effective speech segment into the shared encoder, extract the general acoustic representation, and perform dialect-specific feature transformation on the general acoustic representation through the dialect adapter; S4 inputs the acoustic representation after feature transformation to the decoder, dynamically calculates the weights of each dialect adapter through an attention gating mechanism, and outputs the recognized text based on the weighted fused features.
2. The method according to claim 1, characterized in that, The process of collecting user voice signals, performing activity detection on the voice signals, and extracting valid voice segments includes: The effective speech segments are classified into one of the seven major dialect regions using a lightweight CNN classifier, and the first confidence level is output. When the first confidence score exceeds a preset threshold, a lightweight fine classifier is deployed in the corresponding large dialect area, and prosodic features are introduced as auxiliary features to perform second-level fine-grained dialect subclass recognition, and the corresponding dialect category and second confidence score are output.
3. The method according to claim 1, characterized in that, The dialect adapter is trained using the LoRA low-rank adaptation technique. A low-rank decomposition matrix is injected into the attention layer of the shared encoder, the original model parameters of the shared encoder are frozen, and only the low-rank adaptation parameters are trained. The number of parameters for each dialect adapter is controlled to be around 500,000.
4. The method according to claim 1, characterized in that, The step of inputting the acoustic features of valid speech segments into the shared encoder includes: 80-dimensional FBank acoustic features were extracted at a sampling rate of 16kHz. The FBank acoustic features were then input into a shared encoder containing 12 Conformer Blocks. Each Conformer Block consists of a multi-head self-attention module, a one-dimensional CNN convolution module, a feedforward network module, and layer normalization to extract a general acoustic representation.
5. The method according to claim 1, characterized in that, The dynamic calculation of the weights of each dialect adapter through the attention gating mechanism includes: A gating network is introduced into the decoder. The gating network receives the acoustic features of the current frame as input and outputs the weight vectors of each dialect adapter through a two-layer fully connected network. After Softmax normalization, the weight coefficients of each dialect adapter are obtained. The input dimension of the gating network is the feature dimension output by the shared encoder, and the output dimension is the weight vector dimension of each dialect adapter.
6. The method according to claim 1, characterized in that, It also includes steps for adapting the speech features of the elderly: The speech rate prior parameter in the language model of the decoder is adjusted to adapt the mapping relationship between frame rate and word rate to the speech rate of 120 to 150 words per minute for the elderly, and a duration extension option is added to the CTC decoding path to allow a single phoneme to correspond to a longer frame sequence. A dynamic gain control module is added after speech activity detection and before feature extraction. It automatically performs RMS normalization gain on low-volume speech signals and adopts an adaptive compression strategy to maintain the dynamic range of speech and improve the signal-to-noise ratio in the low-volume segment. Introducing pronunciation degradation modeling into the acoustic model, a pronunciation degradation feature library is established by collecting speech data from the elderly. During the training phase, the degradation features are used as an additional input channel, enabling the model to learn the mapping relationship between degradation patterns and standard pronunciation. During the inference phase, the degree of pronunciation degradation is automatically detected and corresponding compensation strategies are applied.
7. The method according to claim 1, characterized in that, Online adaptive learning steps: Collect the recognition results that users actively correct, and pair and store the original speech segments with the corrected text; Regularly fine-tune the dialect adapters for the corresponding users using newly collected correction data; Incremental learning employs an elastic weight solidification strategy. During fine-tuning, the importance of key parameters of the shared encoder is calculated, and strong regularization constraints are applied to important parameters to avoid catastrophic forgetting of existing knowledge. The loss function of the elastic weight solidification strategy is as follows: in, For parameter importance, These are the old parameter values.
8. A speech device based on multi-dialect adaptive fusion, characterized in that, include: The acquisition module is used to acquire user voice signals, perform activity detection on the voice signals, and extract valid voice segments; The detection module is used to automatically detect the dialect category of valid speech segments and output the corresponding dialect category and confidence level. The feature transformation module is used to load the corresponding dialect adapter according to the dialect category, input the acoustic features of the effective speech segment into the shared encoder, extract the general acoustic representation, and perform dialect-specific feature transformation on the general acoustic representation through the dialect adapter; The weighted fusion module is used to input the acoustic representation after feature transformation into the decoder, dynamically calculate the weights of each dialect adapter through an attention gating mechanism, and output the recognized text based on the weighted fusion features.
9. An electronic device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.