Speech recognition model training method and speech processing method of intelligent customer service

By training and fusing speech recognition models, the problem of insufficient accuracy in dialect speech recognition in intelligent customer service systems has been solved, achieving accurate recognition of both Mandarin and dialects, thus improving communication efficiency and user experience.

CN119626209BActive Publication Date: 2026-05-12CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
Filing Date
2024-11-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing intelligent customer service systems lack accuracy in recognizing dialects, which affects the user's question-and-answer experience.

Method used

By acquiring speech sample data and generating labels, a pre-trained model is trained, followed by dialect-adaptive training. The pre-trained model and the dialect understanding model are then fused to generate a language understanding model that can simultaneously recognize Mandarin and dialects.

Benefits of technology

It improves the accuracy of speech recognition, enabling more precise matching of answers to user questions, thereby enhancing communication efficiency and service experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119626209B_ABST
    Figure CN119626209B_ABST
Patent Text Reader

Abstract

The application provides a speech recognition model training method, a speech processing method and device of an intelligent customer service, electronic equipment and a computer readable storage medium, comprising: obtaining speech sample data, generating a label for each speech sample data, training a pre-training model using the speech sample data to obtain a trained pre-training model, performing dialect adaptability training on the pre-training model using speech sample data with the dialect label to obtain a dialect understanding model; and performing model fusion on the pre-training model and the dialect understanding model to obtain a language understanding model. The application generates a label for speech sample data, trains a pre-training model using a dialect, and enables the dialect understanding model to accurately recognize a dialect. The model after fusing the pre-training model can accurately recognize Mandarin and dialects, can more accurately match answers when a user asks a question, improves communication efficiency, and provides a better service experience for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition, specifically to a speech recognition model training method, a speech processing method for intelligent customer service, an apparatus, an electronic device, and a computer-readable storage medium. Background Technology

[0002] In recent years, intelligent customer service systems have been widely developed due to their wide coverage and rapid response. Many human customer service representatives have integrated with intelligent customer service to serve more users, and have received a lot of positive feedback.

[0003] Nowadays, most of the design and optimization of intelligent customer service focuses on Mandarin, serving the majority of users who speak Mandarin. Intelligent customer service uses voice recognition to identify the user's question and provide a response. Voice recognition is quite accurate in recognizing Mandarin.

[0004] In addition to a large number of Mandarin speakers, there are also many users who speak dialects. The speech recognition effect is poor for these users after voice input, which affects the question-and-answer effect. Summary of the Invention

[0005] This invention provides a speech recognition model training method to address the problem of insufficient accuracy in dialect recognition in prior art.

[0006] Accordingly, embodiments of the present invention also provide a voice processing method, apparatus, electronic device, and computer-readable storage medium for intelligent customer service, to ensure the implementation and application of the above methods.

[0007] In a first aspect, embodiments of the present invention provide a speech model training method, the method comprising:

[0008] Acquire speech sample data and generate a label for each speech sample data; the speech sample data includes Mandarin and various local dialects, and the labels include Mandarin labels and dialect labels, the dialect labels being used to indicate the type of dialect;

[0009] The pre-trained model is trained using the speech sample data to obtain the trained pre-trained model.

[0010] Using speech sample data with the dialect labels, the pre-trained model is subjected to dialect adaptive training to obtain a dialect understanding model; the dialect understanding model is used to recognize the text recognition results of dialect speech data, and the pre-trained model is used to recognize the text recognition results of Mandarin speech data.

[0011] The pre-trained model and the dialect understanding model are fused to obtain a language understanding model; the language understanding model is used to identify the text recognition results of dialect speech data and Mandarin speech data.

[0012] Secondly, embodiments of the present invention provide a speech recognition method, the method comprising:

[0013] Obtain the audio data to be processed from the user input;

[0014] The audio data to be processed is input into the trained language understanding model to obtain the text recognition result corresponding to the audio data to be processed.

[0015] Based on the text recognition results and the preset knowledge base, determine the response statement that matches the text recognition results;

[0016] The response statement is sent to the user's device so that the response statement serves as a reply to the audio data to be processed.

[0017] Thirdly, embodiments of the present invention provide a speech recognition model training apparatus, the apparatus comprising:

[0018] The first acquisition module is used to acquire voice sample data and generate a label for each voice sample data.

[0019] The pre-training module is used to train a pre-trained model using the speech sample data to obtain a trained pre-trained model.

[0020] The dialect training module is used to perform dialect adaptive training on the pre-trained model using speech sample data with the dialect labels to obtain a dialect understanding model.

[0021] The fusion module fuses the pre-trained model and the dialect understanding model to obtain a language understanding model.

[0022] Fourthly, embodiments of the present invention provide a voice processing device for intelligent customer service, the device comprising:

[0023] The second acquisition module acquires the audio data to be processed input by the user;

[0024] The recognition module is used to input the audio data to be processed into the trained language understanding model to obtain the text recognition result corresponding to the audio data to be processed.

[0025] The question-and-answer module is used to determine the response statement that matches the text recognition result based on the text recognition result and the preset knowledge base;

[0026] The response module is used to send the response statement to the user's device so that the response statement serves as a reply to the audio data to be processed.

[0027] Fifthly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of one or more methods as described in embodiments of the present invention.

[0028] In a sixth aspect, embodiments of the present invention provide a readable storage medium that, when instructions in the readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform one or more methods as described in embodiments of the present invention.

[0029] Compared with related technologies, the embodiments of the present invention have the following advantages:

[0030] In this invention, speech sample data is acquired, and labels are generated for each speech sample data. A pre-trained model is trained using the speech sample data to obtain a trained pre-trained model. Speech sample data with dialect labels is then used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Finally, the pre-trained model and the dialect understanding model are fused to obtain a language understanding model. This invention generates labels for speech sample data, classifying the speech sample data to be used for training. After generating labels for the speech sample data, the pre-trained model is trained using the speech sample data, giving the trained pre-trained model a foundation for speech recognition. Since the speech sample data contains a large amount of audio data with Mandarin labels, the accuracy of Mandarin recognition continuously increases during training. Therefore, the pre-trained model has a good understanding ability of Mandarin. After the pre-trained model is trained, all speech samples with dialect labels are used to perform dialect-adaptive training on the pre-trained model. Since the speech samples used are all dialect-labeled, the training results closely approximate the dialect training results, giving the pre-trained model dialect recognition capabilities, resulting in a dialect understanding model. Finally, the dialect understanding model and the pre-trained model are fused, enabling the fused language understanding model to accurately recognize both Mandarin and dialects. By accurately recognizing both dialects and Mandarin, answers to user questions can be matched more precisely, improving communication efficiency and providing a better service experience for users.

[0031] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a flowchart illustrating the steps of a speech recognition model training method provided in an embodiment of the present invention;

[0034] Figure 2 This is a flowchart of another speech recognition model training method provided in an embodiment of the present invention;

[0035] Figure 3 This is a flowchart of the steps of a voice processing method for intelligent customer service provided in an embodiment of the present invention;

[0036] Figure 4 This is a schematic diagram illustrating the overall process of an intelligent customer service system interacting with users, as provided in an embodiment of the present invention.

[0037] Figure 5 This is a structural diagram of a speech recognition model training device provided in an embodiment of the present invention;

[0038] Figure 6 This is a structural diagram of a voice processing device for intelligent customer service provided in an embodiment of the present invention;

[0039] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] When users interact with computers, their input typically needs to be converted into computer-readable information, such as keystrokes, binary codes, or character sequences. Voice communication is a more convenient and efficient method for users. Applying this method to communication between users and computers requires recognizing the user's voice input and converting the audio information into computer-readable input.

[0042] Voice recognition is used in all aspects of our lives. For example, when it is inconvenient to manually input information, it is very convenient and quick to input voice information through voice recognition. In recent years, in particular, intelligent customer service systems have developed rapidly and have become an important means of communication with users. When users encounter problems, they need to contact customer service, and voice communication is the most convenient and quick way for users to communicate. Users can explain their problems by voice input, and then the intelligent customer service will perform voice recognition, understand the user's problem, search for answers and provide feedback to complete the customer service work.

[0043] The speech recognition model training method of the present invention will be further described below with reference to the relevant accompanying drawings and embodiments:

[0044] Figure 1 This is a flowchart illustrating the steps of a speech recognition model training method provided in an embodiment of the present invention, as follows: Figure 1 As shown, the method may include:

[0045] Step 101: Obtain speech sample data and generate a label for each speech sample data.

[0046] In this embodiment of the invention, voice sample data refers to audio data used for communication. The audio data contains information about the user's language communication. When the user communicates, they may not necessarily use Mandarin, but may use local dialects. The specific language used is influenced by the user's lifestyle.

[0047] Dialects are variations of a language in different regions, and their pronunciation usually differs from Standard Mandarin. In some areas, specific words with concrete meanings have emerged and are only used locally. Sometimes, the same word can have different usage scenarios and specific meanings in different regions. The speech sample data obtained in this embodiment of the invention includes Standard Mandarin speech sample data and dialect speech sample data.

[0048] Among them, the label of the voice sample data is an identification symbol carried by the voice sample data. The label stores the dialect information of the voice sample, which is used to indicate whether the voice sample data is Mandarin or a dialect, and which dialect the voice sample data belongs to.

[0049] For example, if the audio content of a voice sample is "The weather is really nice today", and the user who recorded this voice sample speaks Sichuan dialect, the identifier information carried by the tag is {dialect; Sichuan dialect}; if the user who recorded this voice sample speaks Mandarin, the identifier information carried by the tag is {Mandarin; null}.

[0050] After acquiring the speech sample data, a label is generated for each speech sample data, indicating the dialect type of the speech sample data. This distinguishes different dialect speech sample data, making the sample data have clear branches within the dialect category, avoiding confusion between dialects, and ensuring that the model can fully learn the characteristics of various dialects.

[0051] Step 102: Use the speech sample data to train a pre-trained model to obtain a trained pre-trained model.

[0052] In this embodiment of the invention, the pre-trained model is trained by supervised classification, while the speech sample data is used to provide the model with sample data for supervised classification. By training the pre-trained model with the speech sample data, the pre-trained model is able to recognize speech data.

[0053] Specifically, a pre-trained model is a model pre-trained on a large dataset for a specific original task. This model can then be used on a target task, fine-tuned to suit the characteristics of that task, thereby improving its performance. Essentially, it utilizes the theory of transfer learning. By inputting speech sample data into the pre-trained model and outputting the text recognition results of those speech samples, the difference between the recognition results and the samples is calculated to determine the loss function, thus completing the training of the pre-trained model and giving it speech recognition capabilities. These models typically have millions to billions or even more parameters and can handle extremely complex tasks such as natural language processing, computer vision, and speech recognition.

[0054] Furthermore, the pre-trained model only has the function of recognizing Mandarin speech, and its performance in recognizing dialect speech is not ideal. However, the number of dialect speech sample data is relatively small compared to the number of Mandarin speech sample data. It is difficult to achieve a good training effect for the pre-trained model with a small number of sample data so that it can recognize dialect speech data. Therefore, in this embodiment of the invention, a pre-trained model for recognizing Mandarin speech data is first trained, and then the pre-trained model is further adjusted using dialect speech data to make up for the difficulty caused by the scarcity of dialect speech sample data in dialect speech recognition.

[0055] Step 103: Using the speech sample data with the dialect labels, perform dialect adaptation training on the pre-trained model to obtain a dialect understanding model.

[0056] In this embodiment of the invention, dialect adaptive training is an improved training method that uses limited dialect speech sample data and adjusts specific layers in the pre-trained model to improve the training effect of the pre-trained model on dialects.

[0057] Pre-trained models are typically chosen from those trained on large-scale general datasets. Models used for speech recognition usually include the following layers:

[0058] Input layer: Used to receive raw audio signals or other features.

[0059] Convolutional layers: used to extract local features, such as edges and textures in a spectrogram, and are often used to capture the time and frequency features of audio signals.

[0060] Recurrent layers: used to capture long-term dependencies in time-series data, which is particularly important for speech recognition because speech signals are continuous in time.

[0061] Attention mechanism: used to focus on important parts of the input sequence, improving the model's interpretability and performance.

[0062] Fully connected layer: Used to map extracted features to the final output category, typically used in hierarchical tasks.

[0063] Output layer: Generates the final recognition result, such as a text sequence.

[0064] Among these methods, adding more feature dimensions to the input layer can better capture the unique acoustic features of dialects; increasing the number of convolutional layers or adjusting the kernel size can better capture the unique frequency and temporal features of dialects; and optimizing the parameters of the attention mechanism can also make the model pay better attention to the unique features of dialects.

[0065] In this embodiment of the invention, a pre-trained model is obtained, and specific layers of the pre-trained model are adjusted. Speech sample data with dialect labels is input to obtain the output value of one training cycle. The output value and the speech sample data with dialect labels are used to calculate the loss value for this training cycle. Based on the loss value, a loss function is determined, and the parameters of the pre-trained model can be trained. After multiple rounds of training operations, the training objective is met, resulting in a dialect understanding model. This embodiment of the invention does not limit the selection of the loss function. By training the pre-trained model, a model with basic speech recognition functions can be improved to recognize dialect speech data. This allows for the training of a dialect understanding model within a limited amount of speech sample data with dialect labels, compensating for the scarcity of such data.

[0066] In addition to adjusting specific layers in the pre-trained model, dialect recognition and understanding capabilities can be optimized through joint training. Multi-task learning allows the model to learn general features by sharing parameters in the first four layers. A multi-task loss function is designed to combine loss terms from multiple tasks, creating a multi-task dataset to receive the multi-task learning dataset. The weights of the task losses are dynamically adjusted based on the difficulty of the tasks, and the losses for different tasks are alternately optimized during training, enabling the model to achieve good performance on various tasks. Step 104 involves fusing the pre-trained model and the dialect understanding model to obtain a language understanding model.

[0067] In related technologies, the imbalance between dialect speech data and standard Mandarin speech data, with dialect speech data being less abundant than standard Mandarin speech data, limits the model's performance in recognizing and understanding dialect speech data.

[0068] In this embodiment of the invention, a pre-trained model is trained using speech sample data, enabling it to recognize certain general features and thus recognize Mandarin speech data. Although there are some differences between dialects and Mandarin, some features are still quite similar. Therefore, dialect speech sample data with a smaller amount of data is used to perform dialect-adaptive training on the pre-trained model. This adjusts the pre-trained model so that, in addition to basic Mandarin speech understanding, it pays more attention to and captures dialect features. This allows the trained dialect understanding model to recognize dialect speech data, solving the problem that insufficient dialect speech sample data makes it difficult to train a dialect understanding model for understanding dialect speech, and improving the utilization efficiency of dialect speech sample data.

[0069] In this embodiment of the invention, the pre-trained model is trained using speech sample data with dialect labels and speech sample data with Mandarin labels, enabling the pre-trained model to have a good effect on recognizing the general features of speech data. Therefore, the pre-trained model can have accurate recognition results when recognizing Mandarin speech data. The dialect understanding model is trained on the basis of the pre-trained model using speech sample data with dialect labels, enabling the dialect understanding model to have a good effect on recognizing dialect features. Therefore, the dialect understanding model has accurate recognition results for dialect speech data. In order to enable the model to accurately recognize both Mandarin speech data and dialect speech data, the pre-trained model and the dialect understanding model are fused together, integrating the recognition effect of general features and the recognition effect of dialect features into one model, generating a language understanding model that has the function of accurately recognizing both Mandarin speech data and dialect speech data.

[0070] For example, model A has a good capture effect on feature 'a', and can accurately identify it when it receives data 'a'; model B has a good capture effect on feature 'b', and can accurately identify it when it receives data 'b'. In this case, a model C is needed that can simultaneously recognize both feature 'a' and feature 'b'. By using model fusion, model C can capture both feature 'a' and feature 'b' effectively, thus achieving accurate identification of data 'a' and data 'b'.

[0071] In addition, model fusion can integrate the ability to understand Mandarin speech data and dialect speech data, which previously required pre-trained models and dialect understanding models, into a single language understanding model. This eliminates the need to additionally determine whether the speech data is dialect and input it into the corresponding model. Furthermore, it reduces the computational cost and storage requirements of two models to the computational cost and storage requirements of a single model, thereby reducing operational steps, improving efficiency, and saving costs.

[0072] In this invention, speech sample data is acquired, and labels are generated for each speech sample data. A pre-trained model is trained using the speech sample data to obtain a trained pre-trained model. Speech sample data with dialect labels is then used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Finally, the pre-trained model and the dialect understanding model are fused to obtain a language understanding model. This invention generates labels for speech sample data, classifying the speech sample data to be used for training. After generating labels for the speech sample data, the pre-trained model is trained using the speech sample data, giving the trained pre-trained model a foundation for speech recognition. Since the speech sample data contains a large amount of audio data with Mandarin labels, the accuracy of Mandarin recognition continuously increases during training. Therefore, the pre-trained model has a good understanding ability of Mandarin. After the pre-trained model is trained, all speech samples with dialect labels are used to perform dialect-adaptive training on the pre-trained model. Since the speech samples used are all dialect-labeled, the training results closely approximate the dialect training results, giving the pre-trained model dialect recognition capabilities, resulting in a dialect understanding model. Finally, the dialect understanding model and the pre-trained model are fused, enabling the fused language understanding model to accurately recognize both Mandarin and dialects. By accurately recognizing both dialects and Mandarin, answers to user questions can be matched more precisely, improving communication efficiency and providing a better service experience for users.

[0073] Reference Figure 2 The training method for a speech recognition model may include the following steps:

[0074] Step 201: Obtain speech sample data.

[0075] For details, please refer to step 101 above; it will not be repeated here.

[0076] Optionally, step 201 may specifically include:

[0077] Sub-step 2011: Transcribe the obtained speech sample data into text to obtain the text transcription result.

[0078] In this embodiment of the invention, a large amount of speech sample data can be collected through recording devices or online platforms. Professional personnel can then transcribe the speech sample data into text, obtaining the corresponding text information. Through text transcription, the information of the speech sample data is concretely realized in text form, providing a clearer representation of specific vocabulary and expressions in dialects. This adds a textual dimension to the speech sample data, resulting in richer and more informative sample data. It adds more features to the limited dialect speech sample data, providing more dialect-informed sample data for subsequent model training, thus improving the utilization rate of the dialect speech sample data.

[0079] For example, when the pronunciation of a speech sample is "jīn tiān de tiān qì zhēn hǎo yā", the text content "The weather is really nice today" can be obtained through text transcription. The text content is the text content information corresponding to the speech information contained in the speech sample data.

[0080] Furthermore, in many everyday scenarios, we often encounter situations where pronunciations correspond to multiple texts. Transcription of speech sample data can more directly determine the text information that the speech sample data should correspond to in the current context. This text information can be determined by professionals based on the context. Correspondingly, training a model with speech sample data containing text transcription information can link the speech sample data with the context in which it appears. When the model recognizes the context, it can make a clear judgment on the meaning carried by the speech data, allowing the model to clearly understand the content corresponding to the current speech data through the context, thus forming a more accurate recognition.

[0081] For example, when the pronunciation of a voice sample data is "bào fù", it is difficult to determine the corresponding text information based solely on the pronunciation. For such homophonic words with different meanings, it is necessary to consider the context information before and after to determine the specific meaning it wants to express and match the corresponding text result. If the context information is about a person's ambition, then the text information corresponding to this voice is likely to be "ambition". If the context information is about money, then the corresponding text information is likely to be "sudden wealth". If the context information is about some unlucky experiences, the corresponding text information may be "retaliation".

[0082] Sub-step 2012: Combine the text transcription result with the initials and finals of the text to obtain text sample data.

[0083] In the embodiments of the present invention, the initials and finals are generated in combination with the voice sample data. The initials and finals are also the literal representation of the voice. If the voice sample data is in Mandarin, the corresponding initials and finals are the standard initials and finals corresponding to Mandarin. If the voice sample data is in a dialect, then the initials and finals are adjusted approximately according to the specific pronunciation. For example, in Cantonese, the pronunciation of "一" is relatively similar to "yā", so "yā" is used as the initial of "一", while in Mandarin, the pronunciation of "一" is "yī", and both the initial and final have changed. At this time, the initials and finals should be recorded according to the voice sample data.

[0084] Among them, for the convenience of recording the initials and finals, the initials of "你好" (nǐ hǎo) can be marked as "ni3hao3", and the finals of "你好" (nǐ hǎo) can be marked as "i" and "ao".

[0085] In addition, recording the initials and finals of the voice sample data is also to better improve the utilization rate of the dialect voice sample data, so that the dialect voice sample data has more features. By marking the initials, the model can understand the features at the initial level. Especially when using the dialect voice sample data to train the pre-trained model, because there are some similarities between the initials of Mandarin and the dialect, it can help the pre-trained model better master the function of recognizing dialect voice data. And by marking the finals, the model can better understand the features at the phoneme level. By learning the finals, the model can better master the pronunciation rules in the dialect, help the model extract more refined voice features, improve the recognition accuracy, and distinguish the subtle differences in different dialects, so as to avoid the confusion of dialects.

[0086] Sub-step 2013: Generate labels for the text sample data and the voice sample data respectively.

[0087] The labels generated from text sample data are similar to those generated from speech sample data. The labels store dialect information of the speech sample, indicating whether the speech sample is in Mandarin or a dialect, and which dialect it belongs to.

[0088] In this invention, the text sample data includes the text content, pinyin and vowel information transcribed from the corresponding speech sample data. Both the text sample data and the speech sample data have information to distinguish between Mandarin and various local dialects, forming a unified format that is convenient for inputting into the model for training.

[0089] Step 202: Use the speech sample data to train a pre-trained model to obtain a trained pre-trained model.

[0090] For details, please refer to step 102 above; it will not be repeated here.

[0091] Optionally, prior to step 202, the following steps are also included:

[0092] Sub-step 2021: Noise filtering is performed on the speech sample data to obtain noise-filtered speech sample data.

[0093] The noise in the voice sample data can include, for example, background noise, faint human voices in the background, and electrical noise generated by the recording device when it is interfered with.

[0094] In this embodiment of the invention, noise filtering is performed on the speech sample data to minimize the impact of noise on the speech sample data. This ensures that the speech features extracted by the model are not affected by noise when the speech sample data is used to train the model, thereby improving the accuracy of the model for speech recognition.

[0095] Noise can be filtered using linear filtering, Wiener filtering, subspace algorithms, or machine learning methods. This embodiment of the invention does not limit the selection of noise filtering methods.

[0096] Sub-step 2022 involves performing audio segmentation on each character of the noise-filtered speech sample data to obtain single-character speech data.

[0097] In this embodiment of the invention, the speech sample data after noise filtering is segmented word by word to form the smallest language understanding unit, resulting in single-word speech data. After training a pre-trained model, the pre-trained model can recognize the smallest language understanding unit, i.e., a single character. After recognition, and understanding the context, it can better recognize the content corresponding to the speech data. If the audio data is not segmented into the smallest language understanding unit, some sentences with the same meaning but different sentence structures will require the pre-trained model to input richer speech sample data for training in order to recognize the content.

[0098] For example, in everyday spoken language, we often greet others with "Have you eaten?". In some specific regions, inversion is common, expressing it as "Have you eaten?". This inversion doesn't affect understanding or usage in daily communication. If audio is segmented into the smallest unit of understanding, the resulting audio data would be "you," "eat," "food," "eat," and "ma." This segmentation can still recognize inverted sentences like "Have you eaten?" However, if the audio data isn't segmented or is segmented by grammatical structure to obtain similar structures like "you," "eat," and "ma," the recognition of inverted sentences will be less effective.

[0099] Sub-step 2022: After the audio segmentation is completed, the pre-trained model is trained based on the single-word speech data and the labels of the speech sample data to which the single-word speech data belongs, to obtain the trained pre-trained model.

[0100] In this embodiment of the invention, by automatically detecting the end of the audio segmentation process in a speech sample data, single-word speech data arranged according to the audio content of the speech sample data is obtained. The single-word speech data is automatically input into the pre-training model in sequence, so that the input of speech sample data into the pre-training model does not require manual intervention, thereby improving the efficiency of training the model and reducing the manual cost of training the model.

[0101] For example, the audio information of a speech sample is “jīn tiān de tiān qì zhēn hǎo yā”. After audio segmentation of the speech sample, we get “jīn”, “tiān”, “de”, “tiān”, “qì”, “zhēn”, “hǎo”, and “yā”. After the segmentation process is completed, the eight single-character speech data are sequentially input into the pre-trained model.

[0102] Step 203: Using the speech sample data with the dialect labels, perform dialect adaptation training on the pre-trained model to obtain a dialect understanding model.

[0103] For details, please refer to step 103 above; it will not be repeated here.

[0104] Optionally, step 203 may specifically include:

[0105] Sub-step 2031: Combine the text sample data with the dialect tag with the speech sample data with the dialect tag to obtain multimodal dialect data containing the speech sample data and the corresponding text sample data.

[0106] In this embodiment of the invention, multimodal dialect data refers to text sample data with dialect labels and speech sample data with dialect labels. The labels of the speech sample data contain dialect type information, and the text transcribed text carrying the speech sample data in the text sample data is close to the pinyin and vowels of the dialect pronunciation. The labels of the text sample data contain dialect type information.

[0107] In addition, generating corresponding text samples with dialect labels from speech sample data is beneficial because dialect pronunciation differs from standardized pronunciation. Text sample information can establish a connection between the differentiated pronunciation and the standardized pronunciation. Furthermore, some words and phrases in dialects have regional characteristics and are not found in other regions or have very different meanings. Text samples can form a concept unique to that region, avoiding contamination of the meaning of expressions in other regions.

[0108] Sub-step 2032: Using the multimodal dialect data, perform dialect adaptive training on the pre-trained model to obtain a dialect understanding model.

[0109] Dialect-adaptive training involves training a pre-trained model with sample data containing dialect features, while adjusting the layer structure of the pre-trained model to improve its ability to capture dialect features.

[0110] In this embodiment of the invention, a pre-trained model is obtained by training a pre-trained model using speech sample data. Then, multimodal data is used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Since multimodal data contains as many dialect features as possible, including dialect pronunciation, corresponding characters, pinyin, vowels, and dialect usage areas, the pre-trained model with basic speech data understanding capabilities evolves towards understanding dialect speech data. Dialect pronunciation, characters, pinyin, vowels, and regional information all provide the pre-trained model with connections to Mandarin speech sample information, enabling the training of a dialect understanding model from limited dialect speech sample data. Simultaneously, speech and text sample data allow the model to better understand phoneme-level features, master pronunciation rules in dialects, help the model extract more refined speech features, improve recognition accuracy, distinguish subtle differences between different dialects, and thus avoid dialect confusion.

[0111] Optionally, step 203 may specifically include:

[0112] Sub-step 2033: When an update to the speech sample data with the dialect label is detected, the pre-trained model is trained to adapt to the dialect using the speech sample data with the dialect label to obtain a dialect understanding model.

[0113] Currently, the existing speech sample data is unbalanced, with abundant Mandarin speech sample data and scarce dialect speech sample data. Therefore, when new dialect speech sample data is received, the pre-trained model should be trained in a timely manner to obtain a dialect understanding model and improve the dialect recognition accuracy of the dialect understanding model.

[0114] Furthermore, the increasingly rapid speed of information dissemination has a significant impact on language. Timely training of dialect understanding models with new dialect speech sample data is also to ensure that dialect understanding models can keep up with the evolution and development of dialects.

[0115] Sub-step 2034: When the number of concurrent tasks processed by the dialect understanding model is less than a preset threshold, the pre-trained model is trained on dialect-adaptive training using speech sample data with the dialect label to obtain the dialect understanding model.

[0116] In this embodiment of the invention, the processing performance of the dialect understanding model is also monitored.

[0117] For example, in the pre-trained model before dialect adaptation training, the speech data recognized under concurrent tasks is 100 records. For the dialect understanding model, if the speech data processed by the dialect understanding model under concurrent tasks is less than 100 records, it is considered that the processing performance of the dialect understanding model does not meet the requirements and training needs to be continued to improve the response speed of speech data recognition.

[0118] Sub-step 2035: Upon receiving a fine-tuning instruction, in accordance with the fine-tuning instruction, a pre-trained model is trained using the speech sample data, and a dialect-adaptive training is performed on the pre-trained model using speech sample data with the dialect label to obtain a dialect understanding model.

[0119] In a specific embodiment of the present invention, the accuracy of the dialect recognition model can be evaluated based on user feedback to determine whether the user approves of it. If the user is dissatisfied, a fine-tuning instruction is issued to continue training the dialect understanding model through dialect adaptation.

[0120] Optionally, step 203 may specifically include:

[0121] Sub-step 2036: Increase the number of convolutional layers in the convolutional layer of the pre-trained model to obtain a convolutional layer that extracts dialect features.

[0122] Convolutional layers are used to extract local features, typically the time and frequency features of audio signals. By increasing the number of convolutional layers, the unique time and frequency features of dialects can be captured more effectively.

[0123] In this embodiment of the invention, by freezing some convolutional layers, such as the convolutional layers used to extract features of Mandarin speech sample data in the early pre-trained model, while keeping the convolutional layers that extract features of dialect speech sample data working normally, the convolutional layers can increase the recognition accuracy of dialect features while retaining general features.

[0124] Sub-step 2037: Increase the number of recurrent layers in the recurrent layer of the pre-trained model to obtain a bidirectional recurrent layer.

[0125] The recurrent layer extracts long-term dependencies in time-series data, while speech data has temporal continuity.

[0126] In this embodiment of the invention, by increasing the number of recurrent layers and using bidirectional recurrent layers, the dialect recognition model's ability to extract long-term dependencies is enhanced, while also taking into account the contextual information of the speech sample data.

[0127] Sub-step 2038: Increase the number of fully connected layers in the fully connected layer of the pre-trained model to obtain a fully connected layer with dialect features.

[0128] The fully connected layer is used to map the extracted features to the final output category.

[0129] In this embodiment of the invention, the number of fully connected layers is increased so that the additional fully connected layers are used specifically for mapping dialect features, thus ensuring the classification task of dialect features.

[0130] Step 204: Input the speech sample data into the pre-trained model and output the first predicted text recognition result of the speech sample data.

[0131] The process involves inputting speech sample data into a pre-trained model to obtain the model's recognition result, which is the first predicted text recognition result. Re-recognizing the speech sample data using the pre-trained model is to provide a concrete recognition result for the pre-trained model's speech recognition performance. Furthermore, the recognition result after being recognized by the pre-trained model makes the feature extraction and recognition results of the pre-trained model more consistent with the model's training habits. When the target model is trained using the first predicted text recognition result and its associated speech sample data, it helps the target model learn richer features.

[0132] Step 205: Input the speech sample data into the dialect understanding model and output the second predicted text recognition result.

[0133] In this process, the input speech sample data is fed into the dialect understanding model, and the recognition result of the dialect understanding model on the speech sample data is obtained, which is the second predicted text recognition result.

[0134] Step 206: The first predicted text recognition result and the second predicted text recognition result are fused to form the target predicted text recognition result of the speech sample data.

[0135] The first and second predicted text recognition results are fused together to form the target predicted text recognition result for the speech sample data. This target predicted text recognition result includes recognition data from a pre-trained model that recognizes Mandarin and recognition data from a dialect understanding model that recognizes dialects.

[0136] Step 207: Train the initial model based on the speech sample data with the Mandarin label, the speech sample data with the dialect label, and the speech sample data associated with the target predicted text recognition result to obtain the language understanding model.

[0137] Among them, the speech sample data associated with the target predicted text recognition result refers to the speech sample data corresponding to the predicted text recognition result. Since it is the result of recognition by two models, one speech sample data corresponds to two target predicted text recognition results.

[0138] The initial model refers to the initial model used to train the language understanding model using speech sample data with Mandarin labels, speech sample data with dialect labels, and speech sample data associated with the target predicted text recognition results. Specifically, a pre-trained model can be used as the initial model.

[0139] Optionally, step 207 may specifically include:

[0140] Sub-step 2071 involves inputting the speech sample data with the Mandarin label, the speech sample data with the dialect label, and the speech sample data associated with the target predicted text recognition result into the initial model to obtain the third predicted text recognition result output by the initial model.

[0141] In this embodiment of the application, an initial model is used to identify speech sample data of Mandarin labels, speech sample data of dialect labels, and speech sample data associated with the target predicted text recognition results, so that the initial model has an initial recognition result for the above three sample data.

[0142] Sub-step 2072: Determine the first loss value based on the difference between the third predicted text recognition result and the target predicted text recognition result.

[0143] The difference represented by the first loss value indicates the difference between the speech recognition function of the initial model and the pre-trained model and dialect understanding model. The smaller the first loss value, the more the function of the initial model is biased towards the pre-trained model and dialect understanding model.

[0144] Sub-step 2073: Determine the second loss value based on the difference between the third predicted text recognition result, the Mandarin information of the speech samples stored in the Mandarin label, and the dialect information of the speech samples stored in the dialect label.

[0145] The difference represented by the second loss value indicates the accuracy of the initial model in recognizing real speech sample data. The smaller the second loss value, the higher the accuracy of the initial model in recognizing real speech sample data.

[0146] Sub-step 2074 trains the initial model based on the first loss value, the second loss value, and a preset loss function to obtain the language understanding model.

[0147] Specifically, a loss function is calculated using the first and second loss values ​​to balance the recognition accuracy of the initial model on real speech sample data with the recognition similarity between the pre-trained model and the dialect understanding model after training, thus obtaining the language understanding model.

[0148] Furthermore, by using a loss function, the impact of the two trained models and real speech sample data on the language understanding model is balanced. This ensures that the language understanding model does not distort the recognition of real speech sample data while learning the speech features of the two trained models, thus guaranteeing the accuracy of the language understanding model in speech recognition. This enables the language understanding model to recognize both Mandarin and dialect speech data. Moreover, by integrating the functions of the two models into the language understanding model, the computational cost and storage requirements of the two models are reduced to the computational cost and storage requirements of a single model, thereby saving costs.

[0149] For example, the pre-trained model has a computational cost of 100 Ops and a storage requirement of 10GB, used to recognize Mandarin speech data. The dialect understanding model also has a computational cost of 100 Ops and a storage requirement of 10GB, used to recognize dialect speech data. The two together complete the recognition of Mandarin and dialect speech data. The fused language understanding model still uses 100 Ops and still requires 10GB of storage to recognize Mandarin and dialect speech data, but it can retain the functions of the two models very well.

[0150] In this invention, speech sample data is acquired, and labels are generated for each speech sample data. A pre-trained model is trained using the speech sample data to obtain a trained pre-trained model. Speech sample data with dialect labels is then used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Finally, the pre-trained model and the dialect understanding model are fused to obtain a language understanding model. This invention classifies the speech sample data to be used for training by generating labels. Besides the basic categories of dialect and Mandarin, the dialect labels further subdivide the specific dialect system to which the speech sample data belongs, making the label classification of the speech sample data more detailed. This provides a foundation for training a dialect understanding model capable of recognizing multiple dialects. Simultaneously, text sample data is generated from the speech sample data. After generating labels for the speech and text sample data, the pre-trained model is trained, giving it a foundation for speech recognition. Since the speech sample data contains a large amount of audio data with Mandarin labels, the accuracy of Mandarin recognition continuously increases during training. Therefore, the pre-trained model has a good understanding ability of Mandarin. After the pre-trained model is trained, multimodal data consisting of all dialect-labeled speech and text samples is used to perform dialect-adaptive training on the pre-trained model. Because the text samples better connect dialects and Standard Mandarin, the dialect understanding model trained through dialect adaptation achieves higher accuracy in dialect recognition. Finally, the dialect understanding model and the pre-trained model are fused, giving the fused language understanding model the ability to accurately recognize both Standard Mandarin and dialects. By accurately recognizing both dialects and Standard Mandarin, the model can more precisely match answers to user questions, improving communication efficiency and providing a better service experience for users.

[0151] Figure 3 This is a flowchart illustrating the steps of a voice processing method for intelligent customer service provided in an embodiment of the present invention, as follows: Figure 3 As shown, the method may include:

[0152] Step 301: Obtain the audio data to be processed input by the user.

[0153] One way to obtain user input is by having the user record a recording device or by recording a phone call between the user and the intelligent customer service.

[0154] Step 302: Input the audio data to be processed into the trained language understanding model to obtain the text recognition result corresponding to the audio data to be processed.

[0155] In this embodiment of the invention, the user inputs the audio data to be processed, which may be Mandarin or a dialect. The trained language understanding model recognizes the dialect or Mandarin as text and obtains the recognition result.

[0156] Step 303: Based on the text recognition results and the preset knowledge base, determine the response statement that matches the text recognition results.

[0157] The text recognition results are matched with a pre-set knowledge base to identify questions. The knowledge base then returns a response based on the answers to the stored questions, which is then relayed to the user by the intelligent customer service.

[0158] Step 304: Send the response statement to the user's device so that the response statement serves as a reply to the audio data to be processed.

[0159] For example, if a user asks about tomorrow's weather, the intelligent customer service will use a language understanding model to recognize the user's voice input, obtain computer-readable input information, and match it with a preset knowledge base. The knowledge base will then return the information that tomorrow will be sunny to the intelligent customer service. The intelligent customer service will then convert the information returned by the knowledge base into audio data that the user can understand and reply to the user, saying, "Hello, tomorrow will be sunny, it's a good day to go out and exercise."

[0160] Figure 4 For the entire process of interaction between the intelligent customer service system and the user:

[0161] When a user calls the intelligent customer service, the user inputs voice data, which is received by the intelligent customer service. On one hand, the intelligent customer service recognizes the user's dialect, retrieves the language understanding model for recognition, and obtains computer-readable input information. Based on the computer-readable input information after speech recognition, it matches relevant knowledge in a pre-set knowledge base to obtain the answer to the user's question. The answer text is then converted into audio data using the corresponding speech and given to the user. On the other hand, the user's speech triggers the update of dialect speech sample data. After generating labels for the dialect speech sample data, the pre-trained model is trained to obtain a dialect understanding model. Through model fusion, a language understanding model is formed. During communication with users, voice sample data with dialect labels is continuously accumulated, so that the final language understanding model can accurately recognize and understand dialect speech data, so as to provide better voice communication results for users in subsequent communication.

[0162] In this invention, speech sample data is acquired, and labels are generated for each speech sample data. A pre-trained model is trained using the speech sample data to obtain a trained pre-trained model. Speech sample data with dialect labels is then used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Finally, the pre-trained model and the dialect understanding model are fused to obtain a language understanding model. This invention generates labels for speech sample data, classifying the speech sample data to be used for training. After generating labels for the speech sample data, the pre-trained model is trained using the speech sample data, giving the trained pre-trained model a foundation for speech recognition. Since the speech sample data contains a large amount of audio data with Mandarin labels, the accuracy of Mandarin recognition continuously increases during training. Therefore, the pre-trained model has a good understanding ability of Mandarin. After the pre-trained model is trained, all speech samples with dialect labels are used to perform dialect-adaptive training on the pre-trained model. Since the speech samples used are all dialect-labeled, the training results closely approximate the dialect training results, giving the pre-trained model dialect recognition capabilities, resulting in a dialect understanding model. Finally, the dialect understanding model and the pre-trained model are fused, enabling the fused language understanding model to accurately recognize both Mandarin and dialects. By accurately recognizing both dialects and Mandarin, answers to user questions can be matched more precisely, improving communication efficiency and providing a better service experience for users.

[0163] Figure 5 This invention provides a speech recognition model training device, such as... Figure 5 As shown, the device may include:

[0164] The first acquisition module 401 is used to acquire speech sample data and generate a label for each speech sample data.

[0165] The pre-training module 402 is used to train a pre-trained model using the speech sample data to obtain a trained pre-trained model.

[0166] The dialect training module 403 is used to perform dialect adaptive training on the pre-trained model using speech sample data with the dialect labels to obtain a dialect understanding model.

[0167] The fusion module 404 performs model fusion on the pre-trained model and the dialect understanding model to obtain a language understanding model.

[0168] Optionally, the first acquisition module 401 may specifically include:

[0169] The text transcription submodule performs text transcription based on the acquired speech sample data to obtain the text transcription result.

[0170] The text sample submodule combines the text transcription results with the vowels and pinyin to obtain text sample data.

[0171] The tagging submodule generates tags for the text sample data and the voice sample data respectively.

[0172] Optionally, the pre-training module 402 may specifically include:

[0173] The noise filtering submodule performs noise filtering on the speech sample data to obtain noise-filtered speech sample data.

[0174] The audio segmentation submodule performs audio segmentation on the noise-filtered speech sample data character by character, and obtains single-character speech data after segmentation.

[0175] The sample processing submodule, after detecting that the audio segmentation has ended, trains the pre-trained model based on the single-word speech data and the labels of the speech sample data to which the single-word speech data belongs, and obtains the trained pre-trained model.

[0176] Optionally, the dialect training module 403 may specifically include:

[0177] The update detection submodule, when it detects that the speech sample data with the dialect label has been updated, uses the speech sample data with the dialect label to perform dialect adaptive training on the pre-trained model to obtain a dialect understanding model.

[0178] The performance detection submodule, when it detects that the concurrent processing task of the dialect understanding model is less than a preset threshold, uses speech sample data with the dialect label to perform dialect adaptive training on the pre-trained model to obtain the dialect understanding model.

[0179] When the instruction detection submodule receives a fine-tuning instruction, it trains a pre-trained model using the speech sample data in accordance with the fine-tuning instruction, and performs dialect adaptation training on the pre-trained model using speech sample data with the dialect label to obtain a dialect understanding model.

[0180] The convolutional layer submodule adds the number of convolutional layers to the convolutional layers of the pre-trained model to obtain a convolutional layer that extracts dialect features.

[0181] The recurrent layer submodule adds the number of recurrent layers to the recurrent layer of the pre-trained model to obtain a bidirectional recurrent layer.

[0182] The fully connected layer submodule adds the number of fully connected layers to the fully connected layer of the pre-trained model to obtain a fully connected layer with dialect features.

[0183] The multimodal data submodule combines text sample data with dialect tags with speech sample data with dialect tags to obtain multimodal dialect data containing the speech sample data and the corresponding text sample data.

[0184] The multimodal training submodule uses the multimodal dialect data to perform dialect adaptive training on the pre-trained model to obtain a dialect understanding model.

[0185] Optionally, the fusion module 404 may specifically include:

[0186] The first predicted text submodule takes the speech sample data into the pre-trained model and outputs the first predicted text recognition result of the speech sample data.

[0187] The second predicted text submodule inputs the speech sample data into the dialect understanding model and outputs the second predicted text recognition result.

[0188] The target predicted text submodule fuses the first predicted text recognition result and the second predicted text recognition result to form the target predicted text recognition result of the speech sample data.

[0189] The third predicted text submodule inputs the speech sample data with the Mandarin label, the speech sample data with the dialect label, and the speech sample data associated with the target predicted text recognition result into the initial model to obtain the third predicted text recognition result output by the initial model.

[0190] The first loss value submodule determines the first loss value based on the difference between the third predicted text recognition result and the target predicted text recognition result.

[0191] The second loss value submodule determines the second loss value based on the difference between the third predicted text recognition result, the Mandarin information of the speech samples stored in the Mandarin tag, and the dialect information of the speech samples stored in the dialect tag.

[0192] The loss function submodule trains the initial model based on the first loss value, the second loss value, and a preset loss function to obtain the language understanding model.

[0193] In this invention, speech sample data is acquired, and labels are generated for each speech sample data. A pre-trained model is trained using the speech sample data to obtain a trained pre-trained model. Speech sample data with dialect labels is then used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Finally, the pre-trained model and the dialect understanding model are fused to obtain a language understanding model. This invention generates labels for speech sample data, classifying the speech sample data to be used for training. After generating labels for the speech sample data, the pre-trained model is trained using the speech sample data, giving the trained pre-trained model a foundation for speech recognition. Since the speech sample data contains a large amount of audio data with Mandarin labels, the accuracy of Mandarin recognition continuously increases during training. Therefore, the pre-trained model has a good understanding ability of Mandarin. After the pre-trained model is trained, all speech samples with dialect labels are used to perform dialect-adaptive training on the pre-trained model. Since the speech samples used are all dialect-labeled, the training results closely approximate the dialect training results, giving the pre-trained model dialect recognition capabilities, resulting in a dialect understanding model. Finally, the dialect understanding model and the pre-trained model are fused, enabling the fused language understanding model to accurately recognize both Mandarin and dialects. By accurately recognizing both dialects and Mandarin, answers to user questions can be matched more precisely, improving communication efficiency and providing a better service experience for users.

[0194] Figure 6 This invention provides a voice processing device for intelligent customer service, such as... Figure 6 As shown, the device may include:

[0195] The second acquisition module 501 acquires the audio data to be processed input by the user;

[0196] The recognition module 502 is used to input the audio data to be processed into the trained language understanding model to obtain the text recognition result corresponding to the audio data to be processed.

[0197] The question-and-answer module 503 is used to determine the response statement that matches the text recognition result based on the text recognition result and the preset knowledge base.

[0198] The response module 504 is used to send the response statement to the user's device so that the response statement serves as a reply to the audio data to be processed.

[0199] In this invention, speech sample data is acquired, and labels are generated for each speech sample data. A pre-trained model is trained using the speech sample data to obtain a trained pre-trained model. Speech sample data with dialect labels is then used to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model. Finally, the pre-trained model and the dialect understanding model are fused to obtain a language understanding model. This invention generates labels for speech sample data, classifying the speech sample data to be used for training. After generating labels for the speech sample data, the pre-trained model is trained using the speech sample data, giving the trained pre-trained model a foundation for speech recognition. Since the speech sample data contains a large amount of audio data with Mandarin labels, the accuracy of Mandarin recognition continuously increases during training. Therefore, the pre-trained model has a good understanding ability of Mandarin. After the pre-trained model is trained, all speech samples with dialect labels are used to perform dialect-adaptive training on the pre-trained model. Since the speech samples used are all dialect-labeled, the training results closely approximate the dialect training results, giving the pre-trained model dialect recognition capabilities, resulting in a dialect understanding model. Finally, the dialect understanding model and the pre-trained model are fused, enabling the fused language understanding model to accurately recognize both Mandarin and dialects. By accurately recognizing both dialects and Mandarin, answers to user questions can be matched more precisely, improving communication efficiency and providing a better service experience for users.

[0200] The present invention also provides an electronic device, see [link to relevant documentation]. Figure 7 It includes: a processor 601, a memory 602, and a computer program 6021 stored in the memory and executable on the processor. When the processor executes the program, it implements the speech recognition model training method of the aforementioned embodiments.

[0201] The present invention also provides a readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to execute the speech recognition model training method of the foregoing embodiments.

[0202] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0203] It should be noted that all information and data obtained in the embodiments of the present invention were obtained with the authorization of the information / data holder.

[0204] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0205] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0206] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.

[0207] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0208] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present invention. The present invention can also be implemented as a device or apparatus program for performing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0209] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0210] The user information (including but not limited to user device information, user personal information, etc.) and related data involved in this invention are all information authorized by the user or authorized by all parties.

[0211] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0212] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0213] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for training a speech recognition model, the method comprising: The method comprises: acquiring voice sample data and generating labels for each of the voice sample data; the voice sample data comprises Mandarin and various local dialects, and the labels comprise a Mandarin label and a dialect label, the dialect label being used to indicate the category of the dialect; training a pre-trained model using the voice sample data to obtain a trained pre-trained model; performing dialect adaptability training on the pre-trained model using voice sample data with the dialect label to obtain a dialect understanding model; the dialect understanding model is used to identify the text recognition result of dialect voice data, and the pre-trained model is used to identify the text recognition result of Mandarin voice data; performing model fusion on the pre-trained model and the dialect understanding model to obtain a language understanding model; the language understanding model is used to identify the text recognition result of dialect voice data and Mandarin voice data; the model fusion of the pre-trained model and the dialect understanding model to obtain a language understanding model comprises: inputting the voice sample data into the pre-trained model to output a first predicted text recognition result of the voice sample data; inputting the voice sample data into the dialect understanding model to output a second predicted text recognition result; fusing the first predicted text recognition result and the second predicted text recognition result to form a target predicted text recognition result of the voice sample data; inputting voice sample data with the Mandarin label, voice sample data with the dialect label, and voice sample data associated with the target predicted text recognition result into an initial model to obtain a third predicted text recognition result output by the initial model; determining a first loss value according to the difference between the third predicted text recognition result and the target predicted text recognition result; determining a second loss value according to the difference between the third predicted text recognition result, the Mandarin information of the voice sample stored in the Mandarin label, and the dialect information of the voice sample stored in the dialect label; training the initial model according to the first loss value, the second loss value, and a preset loss function to obtain the language understanding model; the preset loss function is a loss function corresponding to the recognition accuracy of the real voice sample data.

2. The method of claim 1, wherein, Before the training of the pre-trained model using the voice sample data to obtain the trained pre-trained model, the method further comprises: performing noise filtering on the voice sample data to obtain noise-filtered voice sample data; performing audio segmentation on the noise-filtered voice sample data word by word to obtain single-character voice data after segmentation; the training of the pre-trained model using the voice sample data to obtain the trained pre-trained model comprises: after detecting the end of the audio segmentation, training the pre-trained model according to the single-character voice data and the label of the voice sample data to which the single-character voice data belongs to obtain the trained pre-trained model.

3. The method of claim 1, wherein, the dialect adaptability training of the pre-trained model using voice sample data with the dialect label to obtain a dialect understanding model comprises: When an update to the speech sample data with the dialect label is detected, the pre-trained model is trained to adapt to the dialect using the speech sample data with the dialect label to obtain a dialect understanding model. When the number of concurrent tasks processed by the dialect understanding model is less than a preset threshold, the pre-trained model is trained using speech sample data with the dialect label to obtain the dialect understanding model. Upon receiving a fine-tuning instruction, the system trains a pre-trained model using the speech sample data and performs dialect-adaptive training on the pre-trained model using speech sample data with the dialect label to obtain a dialect understanding model.

4. The method of claim 1, wherein, The pre-trained model includes convolutional layers, recurrent layers, and fully connected layers. During the dialect-adaptive training of the pre-trained model, the method further includes: The number of convolutional layers is increased in the convolutional layer of the pre-trained model to obtain a convolutional layer that extracts dialect features; when extracting features from speech sample data with dialect labels, the convolutional layers other than the convolutional layer with dialect features are frozen to retain the general features in the extracted features. The number of recurrent layers is increased in the recurrent layer of the pre-trained model to obtain a bidirectional recurrent layer; The number of fully connected layers is increased in the pre-trained model to obtain a fully connected layer that maps dialect features; the fully connected layer is used to map the dialect category of the speech sample data with the dialect label.

5. The method of claim 1, wherein, The step of acquiring speech sample data and generating a label for each speech sample data includes: The obtained speech sample data is used to perform text transcription to obtain the text transcription result; the text transcription result includes the text data corresponding to the speech sample data; The text transcription results are combined with the vowels and pinyin to obtain text sample data; Tags are generated for the text sample data and the speech sample data respectively; the text sample data includes Mandarin and various local dialects, and the tags include Mandarin tags and dialect tags, with the dialect tags indicating the type of dialect; The step of using speech sample data with the dialect labels to perform dialect-adaptive training on the pre-trained model to obtain a dialect understanding model includes: By combining the text sample data with the dialect label with the speech sample data with the dialect label, multimodal dialect data containing the speech sample data and the corresponding text sample data is obtained. Using the multimodal dialect data, the pre-trained model is subjected to dialect adaptive training to obtain a dialect understanding model; the dialect understanding model is used to recognize dialect speech data to obtain text recognition results. 6.A voice processing method of intelligent customer service, characterized in that, include: Obtain the audio data to be processed from the user input; The audio data to be processed is input into the trained language understanding model according to any one of claims 1 to 5 to obtain the text recognition result corresponding to the audio data to be processed. Based on the text recognition results and the preset knowledge base, determine the response statement that matches the text recognition results; The response statement is sent to the user's device so that the response statement serves as a reply to the audio data to be processed. 7.A voice recognition model training apparatus, characterized by comprising: include: The first acquisition module is used to acquire voice sample data and generate a label for each voice sample data. The voice sample data includes Mandarin and various local dialects, and the tags include Mandarin tags and dialect tags, with the dialect tags used to indicate the type of dialect; The pre-training module is used to train a pre-trained model using the speech sample data to obtain a trained pre-trained model. The dialect training module is used to perform dialect adaptive training on the pre-trained model using speech sample data with the dialect label to obtain a dialect understanding model; the dialect understanding model is used to recognize the text recognition results of dialect speech data, and the pre-trained model is used to recognize the text recognition results of Mandarin speech data. The fusion module is used to fuse the pre-trained model and the dialect understanding model to obtain a language understanding model; the language understanding model is used to recognize the text recognition results of dialect speech data and Mandarin speech data. The fusion module may specifically include: The first predicted text submodule is used to input the speech sample data into the pre-trained model and output the first predicted text recognition result of the speech sample data. The second predicted text submodule is used to input the speech sample data into the dialect understanding model and output the second predicted text recognition result. The target predicted text submodule is used to fuse the first predicted text recognition result and the second predicted text recognition result to form the target predicted text recognition result of the speech sample data; The third predicted text submodule is used to input the speech sample data with the Mandarin label, the speech sample data with the dialect label, and the speech sample data associated with the target predicted text recognition result into the initial model to obtain the third predicted text recognition result output by the initial model. The first loss value submodule is used to determine a first loss value based on the difference between the third predicted text recognition result and the target predicted text recognition result. The second loss value submodule is used to determine a second loss value based on the difference between the third predicted text recognition result, the Mandarin information of the speech samples stored in the Mandarin tag, and the dialect information of the speech samples stored in the dialect tag. The loss function submodule is used to train the initial model based on the first loss value, the second loss value, and a preset loss function to obtain the language understanding model; the preset loss function is the loss function corresponding to the recognition accuracy of real speech sample data.

8. A voice processing apparatus of an intelligent customer service, characterized by, include: The second acquisition module acquires the audio data to be processed input by the user; The recognition module is used to input the audio data to be processed into the trained language understanding model according to any one of claims 1 to 5 to obtain the text recognition result corresponding to the audio data to be processed. The question-and-answer module is used to determine the response statement that matches the text recognition result based on the text recognition result and the preset knowledge base; The response module is used to send the response statement to the user's device so that the response statement serves as a reply to the audio data to be processed.

9. An electronic device, comprising: It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1 to 6.

10. A readable storage medium, characterized by, When the instructions in the readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1 to 6.