Multi-dialect speech recognition system and method

Through the parallel processing and dual decision-making mechanism of Mandarin links and dialect links, combined with language classification and semantic confidence evaluation, the recognition accuracy and adaptability of the multi-dial voice recognition system are optimized, and the problem of poor recognition effect of Mandarin and dialects is solved, and high flexibility and customized speech recognition functions are achieved.

CN120260567APending Publication Date: 2025-07-04AISPEECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510401726.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing multi-dial speech recognition system is difficult to achieve the best recognition effect between Mandarin and dialect, and a single model can easily lead to an increase in the misrecognition rate.

Method used

The parallel processing of Mandarin links and dialect links is adopted, combined with language classification models and semantic confidence evaluation, and the identification accuracy and adaptability are optimized through a dual decision mechanism, including model enhancement modules to support specific vocabulary optimization and data updates.

Benefits of technology

It improves the accuracy and stability of multi-dial voice recognition, meets the customized needs of different users, and improves the performance of Mandarin and dialect recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260567A_ABST
    Figure CN120260567A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition, in particular to a multi-dialect speech recognition system and method.The method comprises the steps that input audio is received, audio features are extracted, the audio features are input into a mandarin link recognition model and a dialect link recognition model at the same time, and the dialect link recognition model comprises a multi-dialect recognition model and a language classification model; and an identification result is output, a first decision is made based on an output result of the language classification model, and if the output result is a dialect, the identification result of the dialect link is directly adopted as a final result. And if the output result is mandarin, entering a second decision judgment, in the second decision judgment, calling a semantic model to respectively perform semantic confidence calculation on a mandarin link identification result and a dialect link identification result, and adopting an identification result with high semantic confidence as a final result. According to the invention, through parallel processing of the mandarin link and the dialect link, the accuracy and adaptability of speech recognition can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular, to a multi-dialect speech recognition system and method. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and deep learning technologies, speech recognition systems have been widely used in various fields such as in-vehicle navigation, smart home, and mobile communication. Users' demands for the naturalness, real-time performance, and personalization of speech interaction have been continuously increasing. The multi-dialect free speech recognition system has emerged, providing new technical support to meet the communication needs of users of different languages and dialects.

[0003] In the prior art, most multi-dialect speech recognition systems mainly adopt a single speech recognition model. Its basic principle is to train a model to simultaneously recognize Mandarin and various dialects under a unified framework, and finally use the model output as the recognition result. At the same time, industrial-level customized applications usually rely on a unified hotword interface to implement, and some customized optimizations need to be completed through operations such as interpolation with a separate model to adapt to the recognition requirements of specific fields. There are also some technical solutions that attempt to use a single model to simultaneously solve the recognition tasks of multi-dialects and Mandarin, and all use a unified architecture to achieve recognition output.

[0004] However, due to the extremely similar pronunciations of some words between Mandarin and some dialects, the existing solutions that use a single model to simultaneously solve the recognition tasks of multi-dialects and Mandarin are difficult to take both into account, resulting in the recognition effects of Mandarin and dialects not reaching the best level ultimately. Summary of the Invention

[0005] This application provides a multi-dialect speech recognition system and method, which can optimize the accuracy and adaptability of speech recognition through parallel processing of the Mandarin link and the dialect link, combined with a language classification model and semantic confidence evaluation. This application provides the following technical solutions:

[0006] In a first aspect, this application provides a multi-dialect speech recognition system, and the system includes:

[0007] An audio input module, configured to receive the input audio and extract the audio features of the input audio;

[0008] A model recognition module, configured to simultaneously input the extracted audio features into a Mandarin link recognition model and a dialect link recognition model, where the dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs a recognition result;

[0009] The first-level decision-making module is used to make the first-level decision based on the output result of the language classification model. If the output result is a dialect, directly adopt the recognition result of the dialect link as the final result. If the output result is Mandarin, enter the second-level decision-making judgment.

[0010] The second-level decision-making module is used to calculate the semantic confidence of the recognition results of the Mandarin link and the dialect link respectively by invoking the semantic model in the second-level decision-making judgment, and adopt the recognition result with a higher semantic confidence as the final result.

[0011] In a specific feasible implementation, in the model recognition module, a model enhancement module is embedded, and the model enhancement module is used to optimize speech recognition and adapt to different user needs.

[0012] In a specific feasible implementation, the model enhancement module includes a hot word enhancement sub-module, a data hot update sub-module, and a data return sub-module, which are used to optimize the recognition of specific vocabulary, update domain data in real time, and return user data respectively.

[0013] In a specific feasible implementation, the hot word enhancement sub-module includes request-level hot words, device-level hot words, and product-level hot words.

[0014] In a specific feasible implementation, the data hot update sub-module relies on the data automatic pull framework to perform T+1-level version list updates and T+7-level general domain data updates in each field.

[0015] In a specific feasible implementation, the data return sub-module is used to automatically collect online acoustic data and customer-defined corpora and return them to the training platform to optimize the model.

[0016] In a second aspect, the present application provides a multi-dialect speech recognition method, which is applied to the multi-dialect speech recognition system as described in the first aspect, and adopts the following technical solutions:

[0017] A multi-dialect speech recognition method includes:

[0018] Receiving the input audio and extracting the audio features of the input audio;

[0019] Simultaneously inputting the extracted audio features into the Mandarin link recognition model and the dialect link recognition model, where the dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs recognition results;

[0020] Making the first-level decision based on the output result of the language classification model. If the output result is a dialect, directly adopt the recognition result of the dialect link as the final result. If the output result is Mandarin, enter the second-level decision-making judgment;

[0021] In the second-level decision-making judgment, the semantic model is called to calculate the semantic confidence of the Mandarin link recognition result and the dialect link recognition result respectively, and the recognition result with a higher semantic confidence is used as the final result.

[0022] In a specific feasible implementation, the extracted audio features are input into the Mandarin link recognition model and the dialect link recognition model simultaneously, where the dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs recognition results as follows:

[0023] The Mandarin link recognition model is trained using a vertical model for a specific domain;

[0024] The dialect link recognition model is jointly trained based on dialect data and Mandarin data.

[0025] In a third aspect, the present application provides an electronic device, which includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a multi-dialect speech recognition method as described in the second aspect.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored, and the program is used to implement a multi-dialect speech recognition method as described in the second aspect when executed by a processor.

[0027] In summary, the beneficial effects of the present application at least include:

[0028] (1) Traditional multi-dialect speech recognition methods often use a single model to simultaneously recognize Mandarin and dialects, resulting in difficulty in balancing between Mandarin and dialects for the model and affecting the overall recognition effect. The present application adopts a parallel architecture of a Mandarin link and a dialect link, enabling the Mandarin link to be optimized independently of the dialect link, thereby ensuring that the Mandarin recognition performance is not affected and maintaining the original recognition accuracy and customization function. At the same time, the dialect recognition model does not need to take into account the Mandarin recognition task and can focus on the optimization training of dialect speech data, thus achieving an improvement in dialect recognition ability. This architecture ensures that the recognition performance of both links reaches the optimal level and avoids the performance degradation caused by a single model taking into account multiple tasks.

[0029] (2) Since the pronunciations of Mandarin and some dialects are extremely similar, traditional single models are prone to confusion during the recognition process, leading to an increase in the misrecognition rate. This application introduces a dual decision-making mechanism. Through the first-level judgment of the language classification model, it first determines whether the audio is a dialect. If it is a dialect, the recognition result of the dialect link is directly adopted to avoid the misprocessing of dialect audio by the Mandarin link. If it is Mandarin, it enters the second-level decision-making, calculates the semantic confidence based on the semantic model, and selects the recognition result that better conforms to the context. Through this dual judgment, the system can effectively avoid misrecognition caused by the confusion between Mandarin and dialects, and improve the accuracy and stability of the final recognition result.

[0030] (3) The enhancement module and customization function provided by this application make the speech recognition system have higher flexibility and scalability. For example, it supports hotword enhancement at different levels (request level, device level, product level) to meet the needs of different users for optimizing specific vocabulary; it supports hotword optimization based on geographical location information to improve the recognition effect in specific application scenarios such as navigation; it supports data hot update and feedback mechanism, enabling the system to adjust the recognition model at any time and maintain high adaptability to new field data. The caller can freely combine and apply these modules according to their own needs to achieve highly customized speech recognition functions, greatly enhancing the adaptability and user experience of the system.

[0031] This application proposes a multi-dialect speech recognition solution based on a dual decision-making mechanism. Through the parallel processing of the Mandarin link and the dialect link, combined with the language classification model and semantic confidence evaluation, it optimizes the accuracy and adaptability of speech recognition.

[0032] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of this application and describes them in detail in conjunction with the drawings as follows. Brief Description of the Drawings

[0033] Figure 1 It is a structural block diagram of the multi-dialect speech recognition system in the embodiment of this application.

[0034] Figure 2 It is a schematic flow chart of the multi-dialect speech recognition method in the embodiment of this application.

[0035] Figure 3 It is a schematic overall flow chart of the multi-dialect speech recognition method in the embodiment of this application.

[0036] Figure 4 It is a block diagram of the electronic device for multi-dialect speech recognition in the embodiment of this application. Detailed Description of the Embodiment

[0037] The following will further describe in detail the specific implementation manners of the present application in combination with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0038] Referring to Figure 1 , which is a structural block diagram of a multi-dialect speech recognition system provided by an embodiment of the present application. The system at least includes the following modules:

[0039] An audio input module, configured to receive the input audio and extract the audio features of the input audio.

[0040] A model recognition module, configured to input the extracted audio features into a Mandarin link recognition model and a dialect link recognition model at the same time. The dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs recognition results.

[0041] A first-level decision-making module, configured to perform a first-level decision based on the output result of the language classification model. If the output result is a dialect, directly use the recognition result of the dialect link as the final result. If the output result is Mandarin, enter the second-level decision-making judgment.

[0042] A second-level decision-making module, configured to, in the second-level decision-making judgment, call a semantic model to calculate the semantic confidence degrees of the Mandarin link recognition result and the dialect link recognition result respectively, and use the recognition result with a higher semantic confidence degree as the final result.

[0043] In addition, preferably, in the model recognition module, a model enhancement module is further embedded, configured to optimize the accuracy and customization ability of speech recognition, enable the system to better adapt to different user requirements, and improve the overall performance of Mandarin and dialect recognition.

[0044] The model enhancement module includes three sub-modules: a hot word enhancement sub-module, a data hot update sub-module, and a data feedback sub-module, which are respectively configured to optimize the recognition of specific vocabulary, update domain data in real time, and feedback user data to continuously optimize the model.

[0045] Among them, the hot word enhancement sub-module is configured to improve the recognition accuracy of specific vocabulary, including request-level hot words (supporting the "visible and speakable" function), device-level hot words (supporting the "intelligent address book" function), and product-level hot words (supporting the "product self-customization" function). In addition, the recognition of hot spot place names can be optimized based on the geographical location information of the device (such as Beidou, GPS) to ensure accurate recognition of local vocabulary in different regions and improve the recognition effect of applications such as navigation.

[0046] The data hot update sub-module is used to ensure that the speech recognition system can quickly adapt to the dynamic changes in the industry, achieving T+1 updates (daily iteration of data in specific fields) and T+7 updates (weekly iteration of data in general fields). Relying on a perfect data automatic pulling framework, within the fields concerned by various industries, the version list data at the T+1 level can be quickly updated, and the general field data at the T+7 level can also be iterated regularly, thereby continuously optimizing the speech recognition performance.

[0047] The data return sub-module is used to automatically collect online acoustic data and customer-defined corpora, and return them to the training platform to optimize the model. The acoustic data return mechanism built by the middle platform can regularly update the Mandarin and dialect models. At the same time, the customer-defined corpora in different fields will also be synchronously returned, enriching the capabilities of the speech recognition system. In addition, combined with the internal self-training platform, the customer-defined corpora and regular acoustic supervised data are regularly returned and updated to continuously optimize the recognition model.

[0048] Figure 2 It is a schematic flowchart of a multi-dialect speech recognition method provided by an embodiment of the present application. This method is applied to the above multi-dialect speech recognition system and includes at least the following steps:

[0049] Step S201: Receive the input audio and extract the audio features of the input audio.

[0050] In step S201, after the system receives the input audio signal, it directly uses the MFCC feature extraction technology. By preprocessing the input audio, obtaining the spectrum using the fast Fourier transform, then calculating the energy distribution of the spectrum on the Mel scale through the Mel filter bank, and performing logarithmic operations and discrete cosine transforms on these energy values, the Mel frequency cepstral coefficient audio features (MFCC audio features) are extracted. The MFCC audio features can effectively capture the acoustic characteristics of the speech signal and provide a highly distinguishable input for the subsequent recognition module.

[0051] Step S202: Input the extracted audio features into the Mandarin link recognition model and the dialect link recognition model simultaneously. The dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs the recognition result.

[0052] In step S202, the extracted audio features are input into both the Mandarin link recognition model and the dialect link recognition model simultaneously. The dialect link recognition model consists of a multi-dialect recognition model and a language classification model, and outputs the recognition result of the dialect link and the language type corresponding to the audio respectively. The above models are jointly trained based on dialect data and Mandarin data, and are mainly used to identify and distinguish dialects in more than a dozen different regions such as Sichuan dialect, Cantonese, Shandong dialect, Shanghai dialect, Chongqing dialect, Shaanxi dialect, Henan dialect, Hebei dialect, Tianjin dialect, Anhui dialect, etc., and output the corresponding dialect types at the same time. The Mandarin link recognition model is trained using vertical models for specific fields such as in-vehicle, home, wearable, and conference. Through this method of simultaneously inputting audio features and running in parallel, the system can output the recognition results of the two links respectively, thus ensuring the recognition performance of multiple dialects while retaining the advantages and related functions of Mandarin recognition, and supporting the implementation of various customized solutions for users.

[0053] Step S203: Make a first-level decision based on the output result of the language classification model. If the output result is a dialect, directly use the recognition result of the dialect link as the final result. If the output result is Mandarin, enter the second-level decision judgment.

[0054] In step S203, the system makes a first-level decision based on the output of the language classification model. That is to say, when the language classification model judges the input audio, if the output result is [Dialect], the system directly uses the recognition result of the dialect link as the final result; if the output is [Mandarin], it means that there may be some dialect words in the audio whose pronunciations are similar to Mandarin. At this time, the system will not directly select the result of the Mandarin link, but enter the subsequent second-level decision judgment.

[0055] It should be noted that in this application, the multi-dialect recognition model and the language classification model are grouped together to form the dialect link because they work together to complete the recognition and judgment of dialect audio. Specifically, the multi-dialect recognition model focuses on generating dialect recognition results based on dialect data, while the language classification model is responsible for judging the language attribute of the input audio, that is, outputting [Dialect] or [Mandarin]. This combined design enables the system to directly select the appropriate recognition result based on the output of the language classification model in the first-level decision: when the output is [Dialect], directly use the result of the dialect recognition model to optimize the accuracy of dialect recognition; when the output is [Mandarin], enter the second-level decision for further judgment. Grouping these two models together helps to form a complete processing link for multi-dialect recognition and achieve specialized customized processing and optimization.

[0056] Traditional multi-dialect speech recognition methods often use a single model to simultaneously recognize Mandarin and dialects, making it difficult for the model to balance between Mandarin and dialects and affecting the overall recognition effect. This application adopts a parallel architecture of a Mandarin link and a dialect link, enabling the Mandarin link to be optimized independently of the dialect link, thus ensuring that the Mandarin recognition performance is not affected and maintaining the original recognition accuracy and customization functions. At the same time, the dialect recognition model does not need to take into account the Mandarin recognition task and can focus on the optimization training of dialect speech data, thereby achieving an improvement in dialect recognition ability. This architecture ensures that the recognition performance of both links reaches the optimal level and avoids the performance degradation caused by a single model's need to balance multiple tasks.

[0057] Step S204: In the second-level decision-making judgment, call the semantic model to calculate the semantic confidence of the recognition results of the Mandarin link and the dialect link respectively, and use the recognition result with a higher semantic confidence as the final result.

[0058] In step S204, the system calculates the semantic confidence of the recognition results of the Mandarin link and the dialect link respectively based on the semantic model, and uses the recognition result with a higher confidence as the final output. Semantic confidence measures the rationality of the recognition result in terms of grammatical structure and semantic logic, and is usually calculated through language models, semantic matching algorithms, or deep learning networks. For example, use a pre-trained language model to evaluate the rationality of the recognition result, or calculate the semantic similarity by matching common sentence templates. If the semantic confidence of the dialect link is higher than that of the Mandarin link, it means that the recognition result of the dialect link is more natural and reasonable in the current context, so this result is adopted; otherwise, the recognition result of the Mandarin link is adopted. This dual decision-making mechanism effectively solves the problem that the output result of the Mandarin model is inaccurate after some dialect audio is misclassified as Mandarin, and improves the overall recognition accuracy of the system.

[0059] Since the pronunciations of Mandarin and some dialects are extremely similar, traditional single models are prone to confusion during the recognition process, resulting in an increase in the misrecognition rate. This application introduces a dual decision-making mechanism. Through the first-level judgment of the language classification model, first determine whether the audio is a dialect. If it is a dialect, directly adopt the recognition result of the dialect link to avoid misprocessing of dialect audio by the Mandarin link; if it is Mandarin, enter the second-level decision-making, calculate the semantic confidence based on the semantic model, and select the recognition result that better conforms to the context. Through this dual judgment, the system can effectively avoid misrecognition caused by the confusion between Mandarin and dialects, and improve the accuracy and stability of the final recognition result.

[0060] In summary, combined with Figure 3, this application proposes a multi-dialect speech recognition solution based on a dual decision-making mechanism. Through the parallel processing of the Mandarin link and the dialect link, combined with a language classification model and semantic confidence evaluation, the accuracy and adaptability of speech recognition are optimized.

[0061] This solution first receives the input audio and extracts its audio features. Subsequently, the extracted audio features are simultaneously input into the Mandarin link recognition model and the dialect link recognition model. The dialect link includes a multi-dialect recognition model and a language classification model. The multi-dialect recognition model is responsible for identifying and distinguishing more than a dozen different regional dialects, and simultaneously outputs the recognition result and the dialect type. The language classification model, on the other hand, determines the language attribute of the audio, that is, outputs a [Mandarin] or [Dialect] label. The Mandarin link recognition model generates a Mandarin recognition result based on training data in specific domains (such as in-vehicle, home, wearable, conference, etc.). The system makes a first decision based on the output of the language classification model: if the language classification model determines that the input audio is [Dialect], then directly uses the recognition result of the dialect link as the final result; if it is determined to be [Mandarin], then enters the second decision-making stage. In the second decision-making stage, the system calls the semantic model to calculate the semantic confidence of the recognition results of the Mandarin link and the dialect link respectively, and selects the result with a higher confidence as the final output. Semantic confidence measures the rationality of the recognition result in terms of grammatical structure and semantic logic, and can be calculated through a language model, a semantic matching algorithm, or a deep learning network. For example, use a pre-trained language model to evaluate the rationality of the recognition result, or match common sentence templates to calculate the semantic similarity. If the semantic confidence of the dialect link is higher than that of the Mandarin link, then use the dialect recognition result; otherwise, use the Mandarin recognition result.

[0062] The technical solution of this application solves the accuracy problem existing in the existing multi-dialect and Mandarin recognition using a single model. Due to the similar pronunciation of Mandarin and some dialects in traditional methods, it is easy to cause misrecognition, and it is difficult for a single model to balance the recognition effects of Mandarin and dialects. This solution processes the speech signal through independent Mandarin and dialect links respectively, and introduces language classification and semantic confidence decision-making, optimizing the recognition strategy in different language situations, thereby improving the overall recognition accuracy. In addition, this solution also supports various customization requirements. For example, optimize the recognition effect in a specific domain through the vertical model of the Mandarin link, and combine the dialect link to ensure the dialect recognition ability to meet the diverse needs in practical applications.

[0063] Optionally, this application takes the multi-dialect speech recognition method provided in each embodiment as an example for illustration in an electronic device. The electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.

[0064] Figure 4It is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.

[0065] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0066] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the multi-dialect speech recognition method provided by the method embodiment in the present application.

[0067] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.

[0068] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.

[0069] Optionally, the present application further provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the multi-dialect speech recognition method in the above method embodiments.

[0070] Optionally, the present application further provides a computer product including a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the multi-dialect speech recognition method in the above method embodiments.

[0071] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0072] The above embodiments only represent several implementation manners of the present application, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A multi-dialect speech recognition system, characterized in that The system includes: An audio input module for receiving the input audio and extracting the audio features of the input audio; A model recognition module for simultaneously inputting the extracted audio features into a Mandarin link recognition model and a dialect link recognition model, where the dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs recognition results; A first-level decision-making module for making a first-level decision based on the output result of the language classification model. If the output result is a dialect, directly use the recognition result of the dialect link as the final result. If the output result is Mandarin, enter the second-level decision-making judgment; A second-level decision-making module for, in the second-level decision-making judgment, calling a semantic model to calculate the semantic confidence of the recognition results of the Mandarin link and the dialect link respectively, and using the recognition result with a higher semantic confidence as the final result.

2. The multi-dialect speech recognition system according to claim 1, wherein In the model recognition module, a model enhancement module is embedded, and the model enhancement module is used to optimize speech recognition and adapt to different user needs.

3. The multi-dialect speech recognition system according to claim 2, characterized in that, The model enhancement module includes a hot word enhancement sub-module, a data hot update sub-module, and a data feedback sub-module, which are respectively used to optimize the recognition of specific vocabulary, update domain data in real time, and feedback user data.

4. The multi-dialect speech recognition system according to claim 3, characterized in that, The hot word enhancement sub-module includes request-level hot words, device-level hot words, and product-level hot words.

5. The multi-dialect speech recognition system according to claim 3, characterized in that, The data hot update sub-module relies on a data automatic pulling framework to perform T+1-level version list updates and T+7-level general domain data updates in each field.

6. The multi-dialect speech recognition system according to claim 3, characterized in that, The data feedback sub-module is used to automatically collect online acoustic data and customer-defined corpora and feedback them to the training platform to optimize the model.

7. A multi-dialect speech recognition method, applied to the multi-dialect speech recognition system according to any one of claims 1 to 6, characterized in that, including: Receiving the input audio and extracting the audio features of the input audio; Simultaneously inputting the extracted audio features into a Mandarin link recognition model and a dialect link recognition model, where the dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs recognition results; Making a first-level decision based on the output result of the language classification model. If the output result is a dialect, directly use the recognition result of the dialect link as the final result. If the output result is Mandarin, enter the second-level decision-making judgment; In the second-level decision-making judgment, calling a semantic model to calculate the semantic confidence of the recognition results of the Mandarin link and the dialect link respectively, and using the recognition result with a higher semantic confidence as the final result.

8. The multi-dialect speech recognition method according to claim 1, wherein The step of simultaneously inputting the extracted audio features into a Mandarin link recognition model and a dialect link recognition model, where the dialect link recognition model includes a multi-dialect recognition model and a language classification model, and the model outputs recognition results is expanded: The Mandarin link recognition model is trained using a vertical model for a specific field; The dialect link recognition model is jointly trained based on dialect data and Mandarin data.

9. An electronic device, characterized in that, The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a multi-dialect speech recognition method according to any one of claims 7 to 8.

10. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, and when the program is executed by a processor, it is used to implement a multi-dialect speech recognition method according to any one of claims 7 to 8.

Citation Information

Cited By

  • Identification method, device and equipment based on multilingual speech recognition large model

    CN121438838A