A dialect recognition method and system based on deep learning

CN122821931APending Publication Date: 2026-09-25GUANGDONG UNIVERSITY OF FOREIGN STUDIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611093407.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]本发明旨在解决传统方法在潮汕方言识别中准确率低、适配成本高、难以应对内部差异等问题,提供一种基于深度学习的潮汕方言识别方法及系统,通过引入声调感知低秩适配模块和方言特性引导的训练策略,对大规模预训练语音模型进行高效微调,实现高精度、低成本的潮汕方言识别

Benefits of technology

采用LoRA微调Whisper模型,仅需更新极少量参数(约0.1%~0.5%),显著降低训练成本和存储开销,避免灾难性遗忘,支持在单GPU上快速适配。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821931A_ABST
    Figure CN122821931A_ABST
Patent Text Reader

Abstract

The application discloses a dialect recognition method and system based on deep learning, which comprises the following steps: obtaining Chaoshan dialect speech and extracting spectral features; inputting the fine-tuned Whisper model to obtain recognized text; inserting a tone perception low-rank adaptive module in the Whisper attention layer during fine-tuning; using a training set containing white reading and tone change annotation, only updating the module parameters and freezing the main body. The system comprises a speech collection module, a model fine-tuning module, a recognition inference module and a language model re-scoring module. The application captures the complex tones of Chaoshan dialect through the tone perception branch, combines the stage-by-stage course learning and the radical-stroke-phonogram language model, greatly improves the Chaoshan dialect recognition accuracy with extremely low parameter increment, supports multi-accent branches and end-side deployment, solves the problem of efficient adaptation of low-resource dialects, and has outstanding novelty, creativity and practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent speech processing and natural language understanding technology, specifically relating to a dialect recognition method and system based on deep learning. Background Technology

[0002] Chaoshan dialect retains many features of ancient Chinese, possessing eight tones, complex continuous tone sandhi rules, and significant literary and colloquial pronunciation differences. Furthermore, it contains multiple accent branches, including Chaozhou city dialect, Shantou dialect, Jieyang dialect, and Chaoyang dialect, resulting in substantial differences in pronunciation and phonology. These characteristics pose a significant challenge to automatic speech recognition of Chaoshan dialect.

[0003] Existing dialect recognition technologies are mostly based on traditional Gaussian-Hidden Markov Models (GMM-HMMs) or early end-to-end deep learning frameworks, requiring large amounts of labeled speech data and domain expert knowledge to construct pronunciation dictionaries, acoustic models, and language models. For the Chaoshan dialect, the following are its main shortcomings: The corpus resources are extremely scarce, and there is very little publicly available labeled Chaoshan dialect speech data, making it difficult to train a large-scale model from scratch.

[0004] Dialects have complex internal sound changes, and traditional solutions require separate modeling for different accents, which involves a large amount of engineering work and has poor universality.

[0005] When existing pre-trained models for Mandarin or Cantonese are directly applied to Chaoshan dialect, the recognition accuracy is low and the word error rate is often higher than 50% due to differences in phoneme system, tone system and language model.

[0006] Some studies have attempted to use multilingual pre-trained models for transfer learning, but full parameter fine-tuning can easily lead to catastrophic forgetting, and the computational cost is high, making it difficult to deploy in resource-constrained environments.

[0007] In recent years, large-scale multilingual speech pre-trained models based on Transformer, such as Whisper, have demonstrated powerful cross-language generalization capabilities. However, while Whisper natively supports approximately 99 languages, it does not cover the Chaoshan dialect. Directly using Whisper for Chaoshan dialect recognition primarily outputs Mandarin or Cantonese words with similar pronunciations, failing to generate usable text. Therefore, effectively utilizing the transfer capabilities of such large models while adapting them to the unique phonetic and linguistic characteristics of the Chaoshan dialect at a lower cost has become a pressing issue.

[0008] Low-rank adaptation (LoRA) technology achieves efficient domain or task adaptation by inserting a low-rank decomposition matrix as a bypass into the weight matrix of a pre-trained model, requiring only a small number of parameter updates. This invention reveals that, for the Chaoshan dialect, customized LoRA modules can be strategically inserted at different positions in the Whisper module. Combined with techniques such as tone perception enhancement, joint modeling of literary and colloquial pronunciations, and piecewise branch fine-tuning, this significantly improves recognition accuracy while maintaining model generalization ability and low-resource adaptability. Summary of the Invention

[0009] This invention aims to solve the problems of low accuracy, high adaptation cost, and difficulty in handling internal differences in traditional methods for Chaoshan dialect recognition. It provides a Chaoshan dialect recognition method and system based on deep learning. By introducing a tone-aware low-rank adaptation module and a training strategy guided by dialect characteristics, it can efficiently fine-tune a large-scale pre-trained speech model to achieve high-precision and low-cost Chaoshan dialect recognition.

[0010] To achieve the above objectives, the present invention provides the following technical solution: On the one hand, this invention provides a dialect recognition method based on deep learning, comprising the following steps: Step S1: Obtain the Chaoshan dialect speech signal to be identified and preprocess it to obtain the log-Mel spectrum feature sequence; Step S2: Construct a Chaoshan dialect recognition network based on the Whisper model. The network includes a pre-trained Whisper encoder and decoder, as well as a tone-aware low-rank adaptation module inserted into the attention weight matrix of some layers in the encoder and decoder. Step S3: Using the constructed Chaoshan dialect training dataset, train only the parameters of the tone perception low-rank adaptation module and the additional tone embedding layer, freeze all parameters of the pre-trained Whisper model, and obtain the fine-tuned Chaoshan dialect recognition model. Step S4: Input the preprocessed log-Mel spectrum feature sequence into the fine-tuned Chaoshan dialect recognition model. After encoder forward calculation, encoder-decoder cross attention and autoregressive decoding, the recognized text sequence is obtained.

[0011] In step S2, the tone-aware low-rank adaptation module includes a basic LoRA branch and a tone-aware branch. The basic LoRA branch performs low-rank decomposition and update on the original weight matrix, while the tone-aware branch receives the input hidden state of the current layer, extracts tone envelope features through one-dimensional convolution, and then adds them element-wise to the output of the basic LoRA branch after linear projection and low-rank decomposition to form the final increment matrix. Furthermore, sparse phoneme constraints can be introduced in the decoder embedding layer based on the structural characteristics of the initials and finals in the Chaoshan dialect.

[0012] In step S3, the construction of the training dataset includes: collecting Chaoshan dialect speech from multiple sources, covering different accents, registers, and scenarios; annotating the speech with literary and colloquial pronunciations, continuous tone sandhi, and dialect region branch labels; and expanding the phoneme sequence of the annotated text according to the pronunciation characteristics of Chaoshan dialect to establish an expanded vocabulary containing the three elements of initials, finals, and tones. The training process adopts a phased course learning approach: the first phase uses only single-character and isolated word speech to train basic phonological modeling capabilities; the second phase adds phrases and short sentences; and the third phase introduces natural dialogues and long sentences, while gradually increasing the rank of LoRA.

[0013] In the recognition process of step S4, in view of the characteristics of the Chaoshan dialect, which has many homophones and polyphonic characters, an external Chaoshan dialect semantic language model is introduced to re-score the decoded candidate sequences. This language model adopts the Transformer architecture pre-trained on Chaoshan dialect text and integrates radical embedding to enhance the constraints of similar form and similar meaning.

[0014] On the other hand, the present invention provides a dialect recognition system based on deep learning, comprising: The speech acquisition and front-end processing module is used to acquire Chaoshan dialect speech and convert it into log-Mel spectrum features; The model building and fine-tuning module is used to insert a tone-aware low-rank adaptation module into the pre-trained Whisper model and to perform efficient parameter fine-tuning using Chaoshan dialect training data. The dialect recognition and reasoning module is used to load the fine-tuned model, decode the input features, and obtain the initial text sequence. The language model rescoring module is used to rescore the initial text sequence using an external Teochew language model and output the final recognition result.

[0015] Preferably, the system may also include an accent adaptation branch selection module, which activates the corresponding branch LoRA weights during inference based on user-specified values ​​or the output of the preceding accent classifier, in order to achieve adaptive recognition of different Chaoshan dialect areas.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method. Additionally, a computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the above-described method.

[0017] The present invention also provides a flange gasket helium permeability detection system based on dynamic pressure excitation, including a flange, a gasket, a helium cylinder, a pressure reducing valve, a pressure sensor, a flow controller, and a data processing module; the data processing module is communicatively connected to the pressure sensor and the flow controller, and is used to receive detection data and perform the calculation, identification, and judgment operations in steps S2 to S5 above.

[0018] Beneficial effects Compared with the prior art, the outstanding advantages of the present invention are: Using LoRA to fine-tune the Whisper model requires updating only a very small number of parameters (approximately 0.1% to 0.5%), significantly reducing training costs and storage overhead, avoiding catastrophic forgetting, and supporting rapid adaptation on a single GPU.

[0019] A tone-aware low-rank adaptation module is proposed. By extracting fundamental frequency-related envelope features in parallel and fusing them with LoRA increments, the model can effectively capture phonemic tone information in Chaoshan dialect, thus solving the problem of insufficient modeling of tone dialects by general pre-trained models.

[0020] By constructing a phased learning training set that includes literary and colloquial pronunciations and continuous tone sandhi annotations, the model is guided to gradually master the complex phonological rules of Chaoshan dialect and improve its robustness in multi-accent and multi-register scenarios.

[0021] By introducing a semantic language model that integrates radicals and components for rescoring, the problem of homophone confusion in Chaoshan dialect can be alleviated, and the accuracy of the text can be further improved.

[0022] It supports lightweight deployment of accent branches, can be flexibly adapted to different areas within the Chaoshan region, and has high practicality and scalability. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating an embodiment of the dialect recognition method of the present invention; Figure 2 This is a structural diagram of the Whisper model with the insertion of a tone-perceiving low-rank adaptation module in this invention; Figure 3 A schematic diagram of the internal structure of the tone perception low-rank adaptation module; Figure 4 This is a diagram illustrating a phased learning and training strategy. Figure 5 This is a functional module architecture diagram of the Chaoshan dialect recognition system of the present invention; Figure 6 Example of an adaptive reasoning process for accent-based branches; Figure 7 A comparison chart of word error rates on the test set for six different implementations; Detailed Implementation The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments.

[0024] Example 1; This embodiment provides a basic LoRA-based fine-tuned Whisper method for Chaoshan dialect recognition, serving as a baseline for subsequent improvements. Figure 1 As shown, it includes: Step S101 involves collecting approximately 300 hours of Chaoshan dialect speech data, covering accents from Shantou, Chaozhou, and Jieyang. The corpus includes daily conversations, news broadcasts, excerpts from Chaoshan opera, and telephone customer service recordings. All speech data is uniformly resampled to 16kHz mono, pre-emphasized, framed (25ms frame length, 10ms frame shift), and processed using a Hamming window. 80-dimensional log-Mel filter bank features, i.e., the Mel spectrum, are then extracted.

[0025] Step S102: Construct a pre-trained Whisper model, selecting Whisper Large-v3, which contains a 32-layer encoder and a 32-layer decoder, with a hidden dimension of 1280 and 20 self-attention heads. The original Whisper vocabulary size is 51865, which does not cover the special pronunciations of Chaoshan dialect. Therefore, 200 Chaoshan dialect-specific tokens are added to the original vocabulary, including commonly used dialect characters, contractions (such as the contractions of "nao" and "wuai"), and specific modal particles, forming an expanded vocabulary.

[0026] In step S103, standard LoRA modules are inserted between each self-attention layer and feedforward layer of the encoder, and at the query, key, value, and output projection matrix of each self-attention layer and cross-attention layer of the decoder. Each LoRA module contains two low-rank matrices A and B, where A has a dimension of d×r and B has a dimension of r×d, where d is the original dimension, rank r is 16, and scaling factor α is 32. During forward propagation, the original output becomes... Only the LoRA parameter is randomly initialized, freezing the entire Whisper raw parameter.

[0027] Step S104: Training is performed using the aforementioned 300 hours of speech and corresponding text annotations. The annotated text has been converted into an extended vocabulary ID sequence. The loss function is the standard cross-entropy loss, based on the next token prediction task of autoregressive decoding. The optimizer is AdamW, with a learning rate of 5e-4 and a batch size of 16, trained for 10 epochs on 4 A100 GPUs. SpecAugment spectral augmentation (maximum temporal mask width of 30 frames, maximum frequency mask width of 20 megabands) is applied during training.

[0028] In step S105, after training is completed, the LoRA weights and the extended vocabulary embedding increment are saved. During inference, the Whisper base weights and the LoRA increment are loaded and merged into an equivalent fine-tuned model. Beam search decoding is performed on the input speech features with a beam width of 5 to obtain the recognized text.

[0029] On the test set (20 hours, not involved in training), the word error rate of this baseline model is reduced to 18.7%, while the word error rate of directly using the original Whisper Large-v3 is as high as 72.4%, which proves the effectiveness of LoRA fine-tuning. However, the error rate is still high on minimal pairs with high tone discrimination, such as "诗 / si¹", "时 / si 5 ", "四 / si 6 "), which is about 34.2%.

[0030] Example 2; This embodiment introduces a tone-aware low-rank adaptation module based on Embodiment 1 to enhance tone modeling. The structure of the tone-aware module is shown in Figure 3 . The improvement lies in that the LoRA module inserted in step S103 is not in a standard form, but has a dual-branch structure.

[0031] Taking the value projection matrix of a certain self-attention layer in the encoder as an example, the original dimension d=1280. The first branch is the basic LoRA: the matrix , r=16, calculate . The second branch is a tone-aware branch: first, the input hidden state x (sequence length T×d) is passed through a lightweight fundamental frequency extraction front-end. Since the hidden state does not directly contain acoustic fundamental frequency, a learnable proxy is adopted: x is mapped to a single-channel signal through a one-dimensional convolution (kernel width 5, stride 1), and a pseudo-fundamental frequency contour is obtained after sigmoid gating ; then two layers of dilated convolution (dilation rates 1, 2) are performed on p to obtain tone envelope features . F0 passes through the down-projection matrix and the up-projection matrix to obtain the tone increment The final LoRA output increment , where λ is a learnable adjustment factor, initialized to 0.1. This design enables the model to dynamically adjust the weight update according to the tone of each frame of signal.

[0032] During training, in addition to LoRA parameters, the newly added one-dimensional convolutional layer and parameters A2, B2 also participate in updating, while the main body of Whisper remains frozen. Tone labels are enhanced in the dataset: after each Chinese character in the text sequence, the tone category is explicitly marked with a special token (e.g., "诗<1>"). During training, the model needs to predict both characters and tone tokens, and an auxiliary tone classification loss (with a weight of 0.3) is added.

[0033] Tests show that, compared with Embodiment 1, the error rate of this embodiment on minimal pairs is reduced to 21.5%, and the overall word error rate is reduced to 15.3%, which significantly improves the ability to distinguish tones in Chaoshan dialect. This result confirms the effectiveness of tone-aware LoRA for tonal language transfer.

[0034] Example 3; This embodiment provides a systematic solution integrating radical-based language model rescoring to further reduce homophone errors. The acoustic model outputs of the previous two embodiments usually have a large number of homophone substitutions, for example, "伊人" is misrecognized as "衣人". For this purpose, a special semantic language model for Chaoshan dialect is constructed.

[0035] Specifically, about 5 million sentences of Chaoshan dialect text data are collected, from sources including Chaoshan corpus websites, online forums, Teochew opera scripts, electronic texts of local chorographies, etc. The text is segmented, and the subword granularity is a combination of Chinese characters and some high-frequency words. A small Transformer decoder language model (6 layers, hidden dimension 512, 8 heads) is pre-trained. The key improvement is the introduction of radical embedding: the embedding of each Chinese character is formed by concatenating character vectors and radical vectors. Radical vectors are obtained by splitting a Chinese character into a radical sequence (for example, "潮" is split into "氵" and "朝"), and aggregated by CNN to obtain a fixed-dimensional representation. The training task is standard next Chinese character prediction.

[0036] In the inference stage, for the beam search candidate sequences generated by the acoustic model (e.g., the top 10 candidates), the language model is used to calculate the perplexity score of each candidate, and linear interpolation rescoring is performed with the log probability of the acoustic model: score_total = log P_acoustic + β·log P_LM. β is set to 0.4.

[0037] After applying this embodiment, on the test set of the legal and news fields containing a large number of homophones, the word error rate is further reduced from the original 15.3% to 12.8%, and substitution errors are significantly reduced, for example, the accuracy of "伊人" is increased by 27%.

[0038] Example 4; This embodiment addresses the accental differences within the Chaoshan dialect by designing a segmented low-rank branch network. Different accents in the Chaoshan region exhibit systematic differences in vowels and tones. If all accent data are mixed to train a single LoRA model, the model may confuse certain opposing sound classes.

[0039] In terms of model architecture, for each Transformer layer, in addition to the shared basic LoRA module, K additional accent-branch LoRA modules are set (K is 4, corresponding to Fucheng accent, Shantou accent, Jieyang accent, and Chaoyang accent). During training, each speech is labeled with an accent. The total LoRA increment is calculated as the shared increment plus the corresponding branch increment. Branch parameters are updated only by the specific accent data, while shared parameters are learned jointly by all data. An adversarial loss for accent classification is incorporated into the training objective to encourage the shared parts to learn accent-independent features.

[0040] During inference, branches can be determined either by manual selection by the user or automatically by a pre-defined accent classifier. The pre-defined classifier is a lightweight CNN that achieves 96% accuracy in classifying 2-second speech. Once selected, only that branch is activated and shared with LoRA, resulting in inference speed consistent with a single LoRA.

[0041] The test set for testing the balance of accents from four regions showed a word error rate of 17.2% using a mixed single LoRA (no branching) model, while the word error rate of the segmented branching model in this embodiment was reduced to 14.5%, and the performance variance of each accent was reduced, thus improving practicality.

[0042] Example 5; This embodiment provides a lightweight deployment method suitable for edge devices. Whisper Large-v3 has approximately 1.55 bytes of parameters, making it difficult to run directly on mobile terminals. Therefore, Whisper Small (approximately 244M parameters) is used as the base, while simultaneously applying the tone-sensing LoRA and rescoring system compression of this invention.

[0043] First, Whisper Small was inserted into a rank-8 tone-aware LoRA algorithm, resulting in approximately 4.2M incremental parameters after training. Then, the merged model was 8-bit quantized (dynamic range quantization) and compressed into a model file of approximately 150MB. For the front-end processing, mobile-optimized Mel-spectrum extraction (using the kaldi-native-fbank library) was employed. For the language model, a small 2-layer LSTM language model with 256 hidden layers was trained, combined with radical embeddings, resulting in a final quantized file size of only 12MB.

[0044] Deployed on the Qualcomm Snapdragon 888 mobile platform and using the NCNN inference framework, the real-time rate is approximately 0.45 (0.45 seconds to process 1 second of speech), meeting real-time requirements. Under this lightweight setup, the test set error rate is 19.6%, higher than larger models but far better than the original Small model's 52.3%, demonstrating excellent practicality in scenarios such as controlling smart speakers in Chaoshan dialect and dialect input methods.

[0045] This embodiment demonstrates the end-to-end adaptability and high practicality of the method of the present invention from cloud to edge.

[0046] Example 6; This embodiment describes a method combining data augmentation and semi-supervised training to achieve usable Teochew dialect recognition with extremely low resource conditions (only 50 hours of labeled data). In many real-world scenarios, it is difficult to obtain hundreds of hours of labeled speech.

[0047] First, the tone-aware LoRA model of Example 2 was pre-trained using 50 hours of labeled data. Then, approximately 1000 hours of unlabeled Chaoshan dialect speech (such as audio recordings of Chaoshan opera, podcasts, etc.) were collected. A self-training framework was adopted: the unlabeled speech was decoded using the current model to generate pseudo-labels; approximately 400 hours of high-confidence pseudo-labeled data were obtained by using confidence filtering (only retaining sentences with an average log probability higher than a threshold of -0.8) and language model consistency filtering.

[0048] In terms of data augmentation, in addition to SpecAugment, we also designed the following enhancements for the characteristics of Chaoshan dialect: literary and colloquial pronunciation replacement enhancement (randomly replacing some literary words in the training text with corresponding colloquial words to learn the literary and colloquial mapping); and continuous tone sandhi simulation enhancement (dynamically generating tone sandhi variants in the text according to rules, using the pitch-synchronous overlapping addition method).

[0049] The pseudo-labeled data was mixed with the original labeled data, and the model was trained again. The final error rate on the test set reached 16.2%, which is close to the performance achieved using 300 hours of fully labeled data, demonstrating the high adaptability and novelty of this invention under low-resource conditions.

[0050] The above six embodiments illustrate the implementation of the present invention from the perspectives of baseline fine-tuning, tone perception enhancement, language model rescoring, accent branching, edge-side lightweighting, and low-resource semi-supervised learning. It should be understood that the technical features of the above embodiments can be combined to form a better comprehensive solution. For example, the tone perception LoRA of Embodiment 2 can be combined with the accent branching of Embodiment 4, and the tone perception structure is also used in the slice branching; the edge-side solution of Embodiment 5 can also incorporate the small language model of Embodiment 3, etc., and these combinations all fall within the protection scope of the present invention.

[0051] The system embodiments of the present invention can correspond to the method embodiments. For example... Figure 5 As shown, the system includes a speech acquisition and front-end processing module 51, a model fine-tuning configuration module 52, an acoustic recognition engine 53, and a language model re-scoring module 54. The speech acquisition module 51 can be a microphone array, responsible for beamforming and noise reduction. The configuration module 52 provides a graphical interface, allowing users to select accent branches, adjust LoRA rank, and load weights for training different datasets. The engine 53 encapsulates a Whisper inference pipeline that merges LoRA. The re-scoring module 54 embeds a Teochew dialect language model. The system can be deployed on a cloud server or an embedded terminal.

[0052] Practical application verification has shown that the Chaoshan dialect recognition technology of this invention has achieved good results in scenarios such as intelligent customer service (e.g., voice assistants for banks in the Chaoshan region), cultural preservation (transcription of Chaoshan opera oral scripts), live broadcast subtitles, and dialect navigation. Compared with existing technologies, its recognition accuracy is significantly higher, and it has low resource consumption, strong adaptability, outstanding substantive characteristics and non-obviousness, making it highly valuable for industrial applications.

Claims

1. A dialect recognition method based on deep learning, characterized in that, include: Acquire the speech of the Chaoshan dialect to be identified and extract the log-Mel spectrum feature sequence; The log-Mel spectrum feature sequence is input into the fine-tuned Whisper model. The fine-tuning process includes: constructing a Chaoshan dialect speech training set, inserting a low-rank adaptation module into the attention projection matrix of multiple transformer layers of the pre-trained Whisper model, and using the training set to update only the parameters of the low-rank adaptation module and the expanded vocabulary embedding, and freezing the original parameters of the pre-trained Whisper model. The text sequence output by the autoregressive decoding of the fine-tuned Whisper model is obtained as the recognition result.

2. The dialect recognition method based on deep learning according to claim 1, characterized in that, The construction of the Chaoshan dialect speech training set further includes: Collect Chaoshan dialect speech data with multiple accents and in multiple scenarios, and perform literary and colloquial pronunciation hierarchical annotation and continuous tone sandhi annotation on the speech; In the text annotation, a tone category label is added to each Chinese character to construct an extended vocabulary containing information on initials, finals, and tones; Based on the dialect region branch labels, assign accent branch identifiers to at least a portion of the training data.

3. The dialect recognition method based on deep learning according to claim 1, characterized in that, The low-rank adaptation module is a tone-aware low-rank adaptation module, which includes a basic low-rank decomposition branch and a tone-aware branch. The tone-aware branch receives the input hidden state of the current layer, extracts pseudo-fundamental frequency and tone envelope features through a learnable convolutional network, and then projects these features through a second low-rank decomposition matrix and fuses them with the output of the basic low-rank decomposition branch to form an incremental weight matrix.

4. The dialect recognition method based on deep learning according to claim 3, characterized in that, In the tone perception branch, the initial tone contour is first obtained from the hidden state using one-dimensional convolution and gating mechanism, then the multi-scale tone envelope is extracted by dilated convolution, and finally the increment with the same dimension as the basic branch is generated by the lower projection matrix and the upper projection matrix; during fusion, a learnable adjustment factor is introduced to control the contribution ratio of the tone branch.

5. The dialect recognition method based on deep learning according to claim 1, characterized in that, The method of updating the parameters of the low-rank adaptation module using the training set adopts a phased course learning strategy: the first phase fixes the training data as single characters and isolated words, the second phase adds phrases and short sentences, and the third phase introduces natural dialogues and long sentences; and the rank of the low-rank adaptation module is gradually increased as the phase progresses.

6. The dialect recognition method based on deep learning according to any one of claims 1 to 5, characterized in that, Following the autoregressive decoding, the following is also included: An external Chaoshan dialect semantic language model is used to re-score multiple candidate sequences obtained from decoding. The embedding layer of the semantic language model integrates the radical and component structure features of Chinese characters. The optimal sequence is selected as the final recognition result after linear interpolation of the acoustic model score and the language model score.

7. A dialect recognition system based on deep learning, characterized in that, include: The speech acquisition and front-end processing module is used to acquire Chaoshan dialect speech and convert it into a log-Mel spectrum feature sequence; The model building and fine-tuning module is used to insert a tone-aware low-rank adaptation module into the pre-trained Whisper model, update the parameters of the low-rank adaptation module using Chaoshan dialect training data, and freeze the Whisper main parameters. The dialect recognition and reasoning module loads a finely tuned model, encodes and decodes the input feature sequence, and generates an initial text recognition sequence. The language model rescoring module incorporates a Chaoshan dialect language model that integrates radicals and components. It performs semantic scoring and reordering on the initial text sequence and outputs the final recognition result.

8. The dialect recognition system based on deep learning according to claim 7, characterized in that, The model building and fine-tuning module is also used to build a training dataset containing literary and colloquial pronunciation annotations, continuous tone sandhi annotations, and dialect region branch labels; and the dialect recognition and reasoning module contains an accent branch selector, which activates the corresponding low-rank weight of the accent branch according to the user-specified or prior accent classification results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.