Speech recognition method, device, equipment and product
By identifying and optimizing the dialect types of minority languages, the problem of low speech recognition accuracy is solved, and higher speech recognition accuracy and speech signal quality are achieved.
Patent Information
- Application Number
- CN202510652891.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
There is a problem of low accuracy in the application of existing speech recognition technologies in ethnic minority languages, especially when dealing with dialects and new vocabulary.
By identifying the dialect type of the speech signal of the target language and generating corresponding speech optimization strategies based on the dialect type, the acoustic features are optimized to improve the accuracy of speech recognition.
The accuracy of speech recognition in ethnic minority languages is improved, especially when dealing with dialects and new vocabulary, and the quality of speech signals and the reliability of recognition results are enhanced.
Smart Images

Figure CN120183383A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and particularly to a speech recognition method, apparatus, device and product. Background Art
[0002] With the continuous progress and wide application of intelligent speech recognition technology, the demand for its support for diverse language environments is also increasing day by day. Especially for those ethnic minority languages with a small number of speakers but high cultural value, their speech recognition technology is gradually becoming a research hotspot.
[0003] Although the current speech recognition technology has made remarkable progress in the application of mainstream languages, for languages with a small number of speakers, their corresponding speech recognition systems have the problem of low speech recognition accuracy. Therefore, it is urgent to improve the speech recognition effect for those ethnic minority languages with a small number of speakers but high cultural value. Summary of the Invention
[0004] Based on the above technical status quo, this application provides a speech recognition method, apparatus, device and product, which can improve the speech recognition accuracy of languages with a small number of speakers.
[0005] To achieve the above technical objectives, this application specifically proposes the following technical solutions: According to the first aspect of the embodiments of this application, a speech recognition method is provided, including: when determining that the language type of the speech signal of the target language is a dialect of the target language, identifying the dialect type of the speech signal; according to the speech optimization strategy corresponding to the dialect type, optimizing the acoustic features of the speech signal to obtain optimized acoustic features, where the speech optimization strategy is generated according to the accent characteristics of the dialect type and is a strategy for improving the speech recognition effect; performing speech recognition based on the optimized acoustic features to obtain the speech recognition result of the speech signal.
[0006] According to the second aspect of the embodiments of this application, a speech recognition apparatus is provided, including: a dialect type identification unit, configured to identify the dialect type of the speech signal when determining that the language type of the speech signal of the target language is a dialect of the target language; an optimization processing unit, configured to optimize the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, where the speech optimization strategy is generated according to the accent characteristics of the dialect type and is a strategy for improving the speech recognition effect; a speech recognition unit, configured to perform speech recognition based on the optimized acoustic features to obtain the speech recognition result of the speech signal.
[0007] According to a third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor; the memory is connected to the processor and is used for storing programs; the processor is used for implementing the speech recognition method according to any one of the first aspect and the implementation manners of the first aspect by running the programs stored in the memory.
[0008] According to a fourth aspect of the embodiments of the present application, a computer program product is provided, including computer program instructions, and when the computer program instructions are run by a processor, the processor is enabled to implement the speech recognition method according to any one of the first aspect and the implementation manners of the first aspect.
[0009] A speech recognition method, device, equipment and product provided by the embodiments of the present application. When it is determined that the language type corresponding to the speech signal of the target language is a dialect of the target language, the method recognizes the dialect type of the speech signal, and then optimizes the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features. Finally, speech recognition is performed based on the optimized acoustic features to obtain the final speech recognition result. Among them, the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and is used to improve the speech recognition effect. Since the dialect type is recognized and the dialect is subjected to speech optimization processing before speech recognition is performed according to the acoustic features, the speech quality of the acoustic features is enhanced, thereby improving the accuracy of subsequent speech recognition. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.
[0011] Figure 1 It is a flowchart of a speech recognition method provided by the embodiments of the present application.
[0012] Figure 2 It is a flowchart of performing speech recognition based on optimized acoustic features provided by the embodiments of the present application.
[0013] Figure 3 It is a flowchart of performing speech recognition when the speech signal is the standard language of the target language provided by the embodiments of the present application.
[0014] Figure 4 It is a flowchart of the construction process of the corpus provided by the embodiments of the present application.
[0015] Figure 5Flowchart of the confidence update model based on the speech recognition result provided by the embodiment of the present application.
[0016] Figure 6 Flowchart of updating the model based on the corrected speech recognition result provided by the embodiment of the present application.
[0017] Figure 7 Schematic structural diagram of a speech recognition device provided by the embodiment of the present application.
[0018] Figure 8 Schematic structural diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners
[0019] The technical solution provided by the embodiment of the present application can be applied to application scenarios that require speech recognition of the target language in fields such as medical treatment, education, or justice, so as to improve the speech recognition accuracy of the target language in these scenarios. For example, in the health consultation service scenario, users using the target language can describe their health problems by phone, and the speech recognition system automatically recognizes and gives corresponding suggestions or transfers to the appropriate medical staff. In the distance education platform, the online education platform can integrate the speech recognition function of the target language, allowing teachers to teach in the language of the target language and generate subtitles or notes in real time, facilitating students to review the course content. In the judicial process, by introducing the speech recognition technology of the target language, real-time recording during the court trial can be achieved, ensuring that the statements of all participants can be accurately recorded.
[0020] The technical solution provided by the embodiment of the present application can be exemplarily applied to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged as a software program to be run. When the hardware device executes the processing process of the technical solution of the embodiment of the present application, or when the above software program is run, the automatic splitting of the target task and the automatic invocation of the application program interfaces required for the task can be realized, achieving the purpose of the target task. The embodiment of the present application only gives an exemplary introduction to the specific processing process of the technical solution of the present application, and does not limit the specific implementation form of the technical solution of the present application. Any technical implementation form that can execute the processing process of the technical solution of the present application can be adopted by the embodiment of the present application.
[0021] Next, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0022] Before introducing the solution of the present application, the related technologies will be introduced first: The application of speech recognition technology for ethnic minority languages is becoming increasingly in-depth in the medical, educational, and judicial fields, and has been deeply applied in many key fields, demonstrating a powerful enabling effect. Currently, in order to improve the recognition effect on ethnic minority languages, self-training acoustic models and self-training language models are usually used to optimize the speech recognition system.
[0023] When constructing a self-training acoustic model, in order to improve the robustness and data security of the acoustic model, a common method is to integrate general synthetic data into the customer's real-scene data to improve the recognition effect in general scenarios. This method enables customers to avoid saving historical data for iterative training, thereby reducing the training time. In addition, in a multilingual application scenario, a high-resource language pre-trained acoustic model can be used and then migrated to a low-resource language environment with less data to ensure the recognition accuracy of the low-resource language, thus effectively solving the problem of poor recognition effect caused by insufficient data volume of low-resource languages.
[0024] For self-training language models, a common practice is to collect relevant scenario data and real data at the conversation level, and input them into a semantic generalization system to generate a corpus containing the speaker's intention. Subsequently, the generated corpus is mixed with the real general corpus to optimize the language model, thereby improving the accuracy of speech recognition.
[0025] Although the above strategies can improve the speech recognition effect to a certain extent, the recognition effect for ethnic minority languages (such as Uyghur) is still not ideal. The specific reasons are as follows: (1) Traditional speech recognition systems are mainly used to recognize single-standard Uyghur. However, in actual applications, users often mix Chinese words or phrases into the conversation, that is, Chinese loanwords frequently appear. Since the existing systems are only optimized for standard Uyghur, they will have problems with incorrect recognition when processing Chinese content. In addition, Uyghur belongs to the Altaic language family while Chinese is a tonal language (tone changes can change the meaning of words). There are significant differences in grammar structure, pronunciation methods, etc. between the two languages. Even if the speech recognition system tries to add Chinese support, due to the fundamental differences between the two languages, especially the unique tonal system of Chinese, the recognition accuracy of the Chinese part is still relatively low.
[0026] (2) The growth of network neologisms in Uyghur is rapid. For example, policy terms such as infrastructure workers or social media terms such as live broadcast rooms. These emerging words such as network buzzwords, continuously emerging network neologisms, and newly launched APP names are difficult for the speech recognition system with a relatively lagging update speed to capture in real time due to their rapid iteration and diverse contexts, resulting in a relatively low speech recognition accuracy.
[0027] (3) Uyghur has a variety of dialects, and there are differences in aspects such as pronunciation, intonation, and pronunciation methods among different dialects, which makes it difficult for speech recognition systems to accurately recognize with a unified standard. For example, there may be special phonemes or pronunciation habits in some dialects that are different from standard Uyghur. Currently, most speech recognition engines only support dialects in a single region or neighboring regions and cannot meet the dialect scenario requirements of large-scale applications.
[0028] In summary, the current speech recognition technology cannot effectively meet the speech recognition needs of some languages, thus restricting the popularization and application of speech recognition technology among a wider user group.
[0029] In view of this, the embodiments of the present application are committed to providing a speech recognition method, device, equipment, and product. Before performing speech recognition based on the acoustic features of a speech signal, the dialect type of the speech signal is recognized, and speech optimization processing is performed on the dialect to enhance the quality of its acoustic features, thereby improving the accuracy of the subsequent speech recognition process and ensuring more accurate and reliable speech recognition results. This will be described in detail one by one in the following embodiments.
[0030] Exemplary Method Figure 1 This is a flowchart of a speech recognition method provided by an embodiment of the present application. As Figure 1 shown, the speech recognition method provided in this embodiment includes steps S101 - S103: S101. When it is determined that the language type corresponding to the speech signal of the target language is a dialect of the target language, recognize the dialect type of the speech signal.
[0031] In this embodiment, the target language refers to the language to be subjected to speech recognition. The target language can be a minority language with a relatively small number of speakers. In some examples, the target language can be Uyghur.
[0032] In some embodiments, step S101 includes: obtaining the speech signal of the target language; extracting acoustic features from the speech signal; determining the language type corresponding to the speech signal according to the acoustic features, and the language type includes a standard language or a dialect; if the language type corresponding to the speech signal is a dialect of the target language, then recognize the dialect type corresponding to the speech signal.
[0033] Among them, the voice signal of a user using the target language can be collected by a voice collection device. The voice collection device can select a microphone with high sensitivity and low noise or other devices suitable for specific scenarios. For example, for voice collection in a mobile environment, a portable recording device or an intelligent terminal integrated with a high-quality microphone (such as a smart phone and a tablet computer) can be adopted. In a fixed environment, such as an office or a meeting room, a desktop microphone or a suspended microphone array system can be used to effectively capture the voice signal of the user.
[0034] After the voice signal is obtained, an acoustic feature extraction method can be used to extract acoustic features from the voice signal. Among them, the acoustic feature extraction method can include: Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), Perceptual Linear Prediction (PLP), Zero-Crossing Rate (ZCR), etc. Taking Mel-Frequency Cepstral Coefficients as an example, it divides the voice signal into frames and adds windows, and calculates the Fast Fourier Transform (FFT) of each frame to obtain the spectrum; then, maps the spectrum to the Mel scale, calculates the energy through a set of Mel filter banks, and finally takes the logarithm of the output of the filter banks and performs a Discrete Cosine Transform (DCT) to obtain the MFCC features.
[0035] In order to improve the accuracy of speech recognition, after the acoustic features of the voice signal are extracted, it is necessary to determine which language type the voice signal belongs to according to the acoustic features, such as whether it belongs to a standard language or a dialect, and then determine whether further preprocessing is required according to the specific language type, such as voice signal enhancement processing, to improve the accent problem of the dialect in the voice signal, thereby improving the quality of the voice signal, and then performing speech recognition.
[0036] In some examples, a first language recognition model can be used to determine the language type corresponding to the voice signal according to the acoustic features. The first language recognition model can be pre-trained according to voice samples (including standard languages and their corresponding dialects) and their corresponding true language types. Among them, the first language recognition model can be a (Spoken Language Identification, LID) model.
[0037] When the language type corresponding to the voice signal is a standard language, speech recognition can be directly performed based on the acoustic features to obtain the speech recognition result corresponding to the voice signal.
[0038] When the language type corresponding to the speech signal is a dialect, due to regional differences, there are also different types of dialects. Therefore, it is necessary to further determine the dialect type corresponding to the speech signal.
[0039] In some examples, a second language identification model can be used to determine the dialect type corresponding to the speech signal according to the acoustic features. The second language identification model can be pre-trained according to speech samples (including different types of dialects) and their corresponding true dialect types. Among them, the second language identification model can be a (Spoken Language Identification, LID) model.
[0040] It should be noted that the first language identification model and the second language identification model can be used separately or in combination. When used separately, first, the language type of the speech signal is identified by the first language identification model. When the language type is a dialect, the specific dialect type is further identified by the second language identification model. When used in combination, the first language identification model and the second language identification model can be combined into a comprehensive language identification model. The comprehensive language identification model can be pre-trained according to speech samples (including standard languages and different types of dialects) and their corresponding true language types and dialect types.
[0041] S102. According to the speech optimization strategy corresponding to the dialect type, perform speech optimization processing on the acoustic features of the speech signal to obtain optimized acoustic features.
[0042] In order to improve the speech recognition effect of dialects, different speech optimization strategies can be set for different dialect types to improve the accent problem of dialects in the speech signal, thereby improving the speech signal quality and the subsequent speech recognition accuracy. Among them, the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and used to improve the speech recognition effect.
[0043] In some embodiments, the dialect type includes the first regional dialect type, the second regional dialect type, or the third regional dialect type. The first regional dialect type corresponds to the first speech optimization strategy, and the first speech optimization strategy is used to perform voiceless consonant enhancement processing on the acoustic features; the second regional dialect type corresponds to the second speech optimization strategy, and the second speech optimization strategy is used to perform vowel elongation processing on the acoustic features; the third regional dialect type corresponds to the third speech optimization strategy, and the third speech optimization strategy is used to perform correction processing on syllable truncation of the acoustic features.
[0044] Taking Uyghur as an example, according to geographical distribution, the regions where Uyghur is spoken can be divided into the first region, the second region, and the third region in order from north to south. Among them, the Uyghur dialect in the first region has the characteristic of relatively weak voiced consonants; the Uyghur dialect in the second region shows that the vowels are relatively short; while in the third region, the Uyghur dialect often has phenomena such as syllable weakening, omission, or liaison, which may lead to inaccurate syllable truncation.
[0045] In view of the differences in the pronunciation characteristics of different regional dialect types, different voice optimization strategies can be set for them to optimize the voice of different types of dialects, thereby improving the quality and clarity of the voice signal.
[0046] Among them, when selecting appropriate voice optimization strategies to optimize different types of dialects, it specifically includes: if the dialect type is the first regional dialect type, perform enhanced processing on the voiced consonants of the acoustic features to obtain optimized acoustic features; if the dialect type is the second regional dialect type, perform vowel elongation processing on the acoustic features to obtain optimized acoustic features; if the dialect type is the third regional dialect type, perform correction processing on the syllable truncation of the acoustic features to obtain optimized acoustic features.
[0047] Specifically, if it is detected that the dialect type is the first regional dialect, the first voice optimization strategy is activated. This strategy aims to enhance the voiced consonants in the acoustic features. For example, adjust the energy distribution of the voiced consonants to improve their significance in the frequency spectrum, so that the voiced consonants in the processed voice signal are more clearly distinguishable, thereby obtaining optimized acoustic features.
[0048] If it is detected that the dialect type is the second regional dialect, the second voice optimization strategy is applied. This strategy aims to elongate the vowels in the acoustic features. For example, increase the sound duration of specific vowel segments and adjust the formant parameters to ensure that the vowel pronunciation is full and easy to identify, thereby obtaining optimized acoustic features.
[0049] If the detected dialect type is the third regional dialect, the third voice optimization strategy is adopted. This strategy aims to correct the syllable truncation problem in the acoustic features. By precisely analyzing and adjusting the syllable boundaries in the continuous speech stream, solve the problem of inaccurate syllable truncation caused by weakening, omission, or liaison, ensure that each syllable can be correctly parsed and expressed, and finally form high-quality optimized acoustic features.
[0050] Through the above targeted voice optimization processing method, the naturalness and clarity of the voice signal can be improved, which helps to improve the subsequent voice recognition accuracy. Next, voice recognition can be performed based on the optimized acoustic features to obtain the voice recognition result. For specific details, refer to the detailed introduction of step S103 as follows: S103. Perform speech recognition based on optimized acoustic features to obtain the speech recognition result of the speech signal.
[0051] Since the speech signal may contain new words, resulting in incorrect recognition. Therefore, in this embodiment, it is also possible to first perform decoding based on the optimized acoustic features to obtain the initial recognition text, and then, for each word segment included in the initial recognition text, determine whether it belongs to a new word one by one, and further select different processing strategies based on the judgment results to improve the speech recognition accuracy of new words. The following is a detailed introduction to this implementation process with reference to the accompanying drawings: Figure 2 This is a flowchart of speech recognition based on optimized acoustic features provided by an embodiment of the present application. As Figure 2 shown, speech recognition based on optimized acoustic features includes the following steps S201 - S203: S201. Based on the optimized acoustic features, determine the word segmentation result corresponding to the speech signal.
[0052] In some embodiments, step S201 includes: encoding based on the optimized acoustic features through an encoder to obtain a hidden layer feature representation; decoding based on the hidden layer feature representation through a decoder to obtain the initial speech recognition result corresponding to the speech signal; and performing word segmentation processing on the initial speech recognition result to obtain the word segmentation result corresponding to the speech signal.
[0053] Among them, the encoder can be a cross - language Transformer encoder; the decoder can be a cross - language Transformer decoder.
[0054] The initial speech recognition result is the initial recognition text corresponding to the speech signal. By performing word segmentation processing on the initial recognition text, the initial recognition text can be segmented into each word segment to obtain the word segmentation result corresponding to the speech signal.
[0055] After obtaining the word segmentation result, it is necessary to determine for each word segment in the word segmentation result whether it is a new word or a non - new word one by one. Among them, non - new words can be understood as regular words, that is, words that have been included in a pre - constructed corpus; correspondingly, new words can include network new words, that is, words that have not been included in the pre - constructed corpus.
[0056] In some implementation manners, the GPT - 4V (GPT - 4 with Vision) model can be used to determine whether each word segment belongs to a new word. For example, in the model, an instruction format of "Please classify the following words as new words or non - new words" is input to obtain the judgment result of whether each word segment belongs to a new word or a non - new word.
[0057] Among them, GPT-4V can be continuously trained based on the latest network data, and it can capture the changing trends of language, including the emergence and popularity of new words. Since the model can not only recognize new words but also understand the cultural and semantic evolution behind them, the model can analyze their meanings and generate reasonable transliteration expressions.
[0058] In some implementation manners, for each word segment in the word segmentation result, it can also be determined one by one whether it is a new word or a non-new word through a pre-constructed corpus.
[0059] After that, suitable decoding strategies are selected for decoding according to different judgment results to obtain the speech recognition result. For details, please refer to the introduction of the following steps S202 and S203.
[0060] S202: For each word segment among all word segments, when it is determined that the word segment is a non-new word, based on the correspondence between the non-new word in the corpus and the transliteration result in the target language, determine the transliteration result of the word segment in the target language, and decode according to the transliteration result in the target language to obtain the speech recognition result.
[0061] Specifically, when it is determined that a certain word segment is a non-new word, it means that it has been included in the pre-constructed corpus, and then the transliteration result of the word segment in the target language can be directly determined based on the correspondence between the non-new word in the corpus and the transliteration result in the target language, and the speech recognition result can be obtained by re-decoding according to it.
[0062] S203: For each word segment among all word segments, when it is determined that the word segment is a new word, determine the transliteration result of the word segment in the target language based on the determined semantics of the word segment, and decode according to the transliteration result of the word segment in the target language to obtain the speech recognition result.
[0063] When it is determined that a certain word segment is a new word, since the transliteration result of the word segment in the target language cannot be found in the pre-constructed corpus. Therefore, it is necessary to determine the source and semantics of the word segment through the GPT-4V model, and then determine the transliteration result of the word segment in the target language according to the source and semantics, and decode according to it to obtain the speech recognition result.
[0064] In some examples, a language model (such as an N-Gram model) can be used to re-decode according to the transliteration result in the target language to obtain the speech recognition result.
[0065] In this embodiment, it is first determined whether the specific language type corresponding to the speech signal of the target language is a certain specific dialect under this language; if it is determined to be a dialect, the specific dialect type of this speech signal is further identified, and based on the identified dialect type, a speech optimization strategy customized in advance according to the accent characteristics of this dialect is applied to perform targeted optimization processing on the acoustic features of the speech signal, so as to obtain better-quality acoustic features. Finally, the speech recognition process is performed using these optimized acoustic features to generate the final speech recognition result, thereby improving the accuracy of the speech recognition result.
[0066] Please refer to Figure 3 , in another embodiment of the present application, there is also provided an implementation process of how to perform speech recognition when the language type corresponding to the speech signal of the target language is the standard language of the target language, which may specifically include the following steps S301 and S302: S301. When it is determined that the language type corresponding to the speech signal of the target language is the standard language of the target language, based on the acoustic features of the speech signal, determine the word segmentation result corresponding to the speech signal.
[0067] Among them, each word segmentation is included in the word segmentation result.
[0068] Before step S301, it is necessary to obtain the speech signal of the target language; extract the acoustic features of the speech signal; and determine the language type corresponding to the speech signal according to the acoustic features, and this language type includes the standard language or the dialect.
[0069] It should be noted that the specific implementation methods of how to obtain the speech signal of the target language, extract its acoustic features, and determine the language type corresponding to the speech signal according to the acoustic features are similar to the specific implementation methods in step S101. For the detailed introduction of the specific implementation methods in step S101, reference can be specifically made, and details are not described here again.
[0070] In some embodiments, step S301 includes: encoding based on the acoustic features through an encoder to obtain a hidden layer feature representation; decoding based on the hidden layer feature representation through a decoder to obtain an initial speech recognition result corresponding to the speech signal; and performing word segmentation processing on the initial speech recognition result to obtain the word segmentation result corresponding to the speech signal.
[0071] Among them, the encoder can be a cross-lingual Transformer encoder; the decoder can be a cross-lingual Transformer decoder.
[0072] The initial speech recognition result is the initial recognition text corresponding to the speech signal. By performing word segmentation processing on the initial recognition text, the initial recognition text can be segmented into each word segmentation to obtain the word segmentation result corresponding to the speech signal.
[0073] After obtaining the word segmentation results, it is necessary to determine whether each word in the word segmentation results is a new word or a non-new word one by one. Among them, non-new words can be understood as regular words; correspondingly, new words can include new Internet words.
[0074] In some examples, the GPT-4V (GPT-4 with Vision) model can be used to determine whether each word belongs to a new word according to the Internet corpus. For example, by inputting the instruction format of "Please classify the following words as new words or non-new words" into the model, the judgment result of whether each word belongs to a new word or a non-new word can be obtained.
[0075] Among them, GPT-4V can be trained based on the latest Internet data and can capture the changing trends of language, including the emergence and popularity of new words. Since the model can not only identify new words but also understand the cultural and semantic evolution behind them, the model can analyze its meaning and generate a reasonable transliteration expression.
[0076] S302. Perform speech recognition based on the word segmentation results corresponding to the speech signal to obtain the speech recognition result of the speech signal.
[0077] In some embodiments, step S302 includes: for each word in each word segmentation, when it is determined that the word is a non-new word, based on the correspondence between the non-new word in the corpus and the transliteration result in the target language, determine the transliteration result in the target language corresponding to the word, and decode according to the transliteration result in the target language to obtain the speech recognition result; for each word in each word segmentation, when it is determined that the word is a new word, based on the determined semantics of the word, determine the transliteration result in the target language corresponding to the word, and decode according to the transliteration result in the target language corresponding to the word to obtain the speech recognition result.
[0078] Specifically, when it is determined that a certain word is a non-new word, it means that it has been included in the pre-constructed corpus, and then the transliteration result in the target language corresponding to the word can be directly determined based on the correspondence between the non-new word in the corpus and the transliteration result in the target language, and the speech recognition result can be obtained by re-decoding it.
[0079] When it is determined that a certain word is a new word, since the transliteration result in the target language corresponding to the word cannot be found in the pre-constructed corpus. Therefore, it is necessary to determine the source and semantics of the word through the GPT-4V model, and then determine the transliteration result in the target language corresponding to the word according to the source and semantics, and decode according to it to obtain the speech recognition result.
[0080] In some embodiments, if it is determined that a word segment is a non-emerging word and the corresponding phonetic transcription result in the target language cannot be determined according to the correspondence between the non-emerging words in the corpus and the phonetic transcription results in the target language, then the word segment can be skipped and the corresponding recognition result can be output in blank format.
[0081] In some examples, the speech recognition result can be re-decoded according to the phonetic transcription result in the target language through a language model (such as an N-Gram model).
[0082] In this embodiment, the initial speech recognition result is segmented, and it is determined whether each word segment is an emerging word, so as to adopt different processing strategies for different judgment results to improve the recognition accuracy of each word segment. On this basis, by re-decoding each processed word segment, the accuracy of the overall speech recognition result is further improved.
[0083] Please refer to Figure 4 , in another embodiment of the present application, a process for constructing a corpus is further provided, which may specifically include the following steps S401 to S403: S401. Establish the correspondence between the speech data and the recognition text of the standard language in the target language, and the correspondence between the dialect speech data in the target language and its corresponding recognition text, to obtain an initial corpus.
[0084] In some embodiments, the specific implementation process of establishing the correspondence between the speech data and the recognition text of the standard language of the standard language can be as follows: In some examples, by collecting the speech data corresponding to the standard language in the target language in scenarios such as the medical field, the judicial field, the education field, and daily communication, and through a double-blind annotation method, different annotators independently annotate it to obtain the correspondence between the speech data and the recognition text of the standard language of the standard language.
[0085] Among them, the sampling rate of the speech data corresponding to the standard language is 16 kHz, and the signal-to-noise ratio ≥ 40 dB. To increase the diversity of the data, speech data in different fields can be collected. For example, the respective proportions of the speech data in the medical field, the judicial field, the education field, and daily communication are 30%, 25%, 20%, and 25% in turn. The duration of each speech data is controlled within 5 - 15 seconds / sentence, and the time interval between sentences is greater than or equal to 1 second.
[0086] Furthermore, the collection objects of the speech data can also be users who are between 18 and 65 years old and whose mother tongue is the target language. The ratio of the number of users of different genders among these users is 1:1.
[0087] In addition, the voice data can also be from different regions, such as the first region, the second region, and the third region.
[0088] To ensure the accuracy of the annotation results, the annotation consistency of different annotators needs to be ≥95%.
[0089] It should be noted that the voice data of the standard language can include standard voice data and Chinese loanwords. For Chinese loanwords, they can be transliterated, and the corresponding relationship between the Chinese loanwords and the transliteration results in the target language can be constructed. For the standard voice data, the text data of the target language can be directly annotated for it, so as to construct the corresponding relationship between the voice data and the text data of the standard language.
[0090] In some embodiments, the specific implementation process of constructing the corresponding relationship between the dialect voice data and the recognized text can be as follows: To ensure the diversity of the dialect voice data, the dialect voice data from different regions (such as the first region, the second region, and the third region) can be collected at the network end. The cumulative duration of these dialect voice data is 5000 hours, and it includes recordings under environmental noises such as hospitals, markets, and schools.
[0091] At the same time, the user can mark the sentences and words with recognition errors and upload them to the cloud server in real time. The cloud server corrects the text data corresponding to the dialect voice data according to the sentences and words with recognition errors marked by the user, and constructs the corresponding relationship between the dialect voice data and the corrected text data according to the corrected text data.
[0092] In addition, a dialect feature extraction model can also be constructed based on the dialect voice data. This model can judge which type of dialect a piece of voice belongs to according to the acoustic features of the voice data. This model can use XGBoost.
[0093] For example, there is a piece of dialect voice. After processing, 25 MFCC feature values are extracted. These feature values are input into the trained XGBoost model, and the model can output results in the following form: Dialect of the first region: 80%; Dialect of the second region: 15%; Dialect of the third region: 5%. This indicates that the model believes that this piece of voice is most likely to be the dialect of the first region.
[0094] S402. If the network vocabulary of the target language obtained is not included in the initial corpus, query the transliteration result of the network vocabulary in the target language through online networking, and add the corresponding relationship between the network vocabulary and its transliteration result in the target language to the initial corpus to obtain the corpus.
[0095] In step S402, through the social media crawler system, monitor the content on social platforms such as official media websites, TikTok, and Instagram with a burst index ≥ 7 (calculated based on the number of forwards, likes, or comments) and a usage rate ≥ 80%. This process aims to capture emerging popular words and phrases on the network, ensuring that such highly popular and relevant content can be promptly discovered and incorporated into the corpus to update and enrich the initial corpus.
[0096] For online words, if it is found that they are not included in the initial corpus, it is necessary to initiate an online query for the transliteration results of the online words and generate the corresponding relationship between the online words and the transliteration results in the target language, and add it to the initial corpus.
[0097] S403. If the corresponding relationship between the online word and the transliteration result in the target language cannot be queried through online connection, then generate the corresponding relationship between the online word and its corresponding transliteration result in the target language based on the determined semantics of the online word, and add the corresponding relationship to the initial corpus to obtain the corpus.
[0098] If the corresponding relationship between the online word and the transliteration result in the target language cannot be queried through online connection, then the semantics of the word can be further determined through the GPT-4V model, generate the corresponding relationship between the online word and its corresponding transliteration result in the target language, and add it to the initial corpus to obtain the corpus.
[0099] For example, for online words, when no corresponding relationship is queried online, the GPT-4V model can be further used to determine whether the online word belongs to an online buzzword. For example, if it is found through the model that a new online word is derived from the homophone of an existing word, then add a semantic label to it: {"source": "online homophone meme", "mapping suggestion": "new online word → homophone of existing word"}. And write this mapping relationship into the initial corpus.
[0100] It should be noted that for new online words, regardless of whether they appear in voice or text form, the corresponding relationship between them and the transliteration results in the target language can be established. Therefore, the corpus in this embodiment can also be understood as a Chinese-Uyghur transliteration rule library.
[0101] In this embodiment, a corpus is constructed. The corpus not only includes the correspondence between the speech data of standard languages and the recognized texts, but also includes the correspondence between the speech data of Chinese loanwords and the transliteration results in the corresponding target languages, as well as the correspondence between the speech data or text data of network neologisms and the transliteration results in the corresponding target languages. Therefore, when performing speech recognition, this corpus can be applied to obtain the transliteration results in the corresponding target languages for Chinese loanwords, dialects, and network neologisms respectively, thereby helping the language model to understand the speech data during decoding and further improving the accuracy of speech recognition.
[0102] Please refer to Figure 5 , in another embodiment of the present application, the speech recognition method of the embodiments of the present application can be executed by a speech recognition model, and the speech recognition model can be obtained by training with training samples and corresponding labels. The training samples include various speech data collected during the construction of the corpus (such as including standard Uyghur, standard Uyghur containing Chinese loanwords, Uyghur dialects, standard Uyghur containing network neologisms, and Uyghur dialects containing network neologisms, etc. in various forms). The label of each training sample is the recognized text or transliteration result of the corresponding speech data generated according to the corpus. The output of the speech recognition model also includes the confidence level of the speech recognition result, and the model can be dynamically updated based on this confidence level. Specifically, it can include the following steps S501 to S503: S501. When it is determined that the confidence level of the speech recognition result is less than or equal to the preset confidence threshold, obtain the corrected speech recognition result.
[0103] Among them, the confidence level represents the degree of credibility of the speech recognition model for its output result, reflects the evaluation of the certainty of the recognition result by the model, and it is a probability estimate used to represent the possibility that the result is correct. If the confidence level of the speech recognition result is less than or equal to the preset confidence threshold, it means that the accuracy of the speech recognition result is relatively low and the model has insufficient confidence in this result; on the contrary, if the confidence level of the speech recognition result is greater than the preset confidence threshold, it means that the speech recognition result has a relatively high accuracy and the model has a high confidence in its output.
[0104] When the confidence level of the speech recognition result is less than or equal to the preset confidence threshold, it is necessary to obtain the corrected speech recognition result and update the model according to it.
[0105] Among them, the confidence threshold can be set according to empirical values, and the value range is [0.9, 0.97]. For example, it can be set to 0.94, 0.95, 0.96, etc.
[0106] In this embodiment, the execution subject of the speech recognition method may be the edge side, such as a terminal device. Then, obtaining the corrected speech recognition result includes: uploading the speech recognition result to the cloud server so that the cloud server corrects the speech recognition result according to multimodal data, and obtains the corrected speech recognition result. The multimodal data includes the corrected text data and the user's lip movement video; receiving the corrected speech recognition result returned by the cloud server.
[0107] The edge side initiates a cloud API request to the cloud server to upload the speech recognition result with a low confidence level to the cloud server for processing, so that the cloud server combines the corrected text and multiple modal data including the image of the user's lip movement to improve the accuracy of the speech recognition result.
[0108] During the interaction process, the user can manually correct the speech recognition result output by the model to obtain the corrected text data. The image data includes the lip movement video of the user speaking corresponding to the speech signal. The lip movement recognition technology can capture the parts of the speech signal that may be interfered or blurred by noise, so as to provide supplementary information for speech recognition.
[0109] After comprehensively analyzing and correcting the speech recognition result based on the above multimodal data, the cloud server generates a more accurate result and returns it to the edge side. The corrected result not only considers the speech signal itself, but also integrates text, image and other relevant information, thus significantly improving the reliability of the recognition result.
[0110] Further, the speech recognition model can also be updated by using the corrected speech recognition result to improve the performance of the model. For the detailed introduction, please refer to step S502 as follows: S502. Update the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model.
[0111] Please refer to Figure 6 , in another embodiment of the present application, after obtaining the corrected speech recognition result, updating the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model may specifically include the following steps S601 to S604: S601. Determine the influencing factors with a confidence level less than or equal to the preset confidence threshold according to the corrected speech recognition result.
[0112] Among them, the influencing factors include the first influencing factor or the second influencing factor. The first influencing factor includes the existence of words not included in the corpus. The second influencing factor includes dialects.
[0113] In some embodiments, step S601 includes the following steps a1 to a4: Step a1: Determine the perplexity of the corrected speech recognition result through a language model, and determine the matching degree between the acoustic features of the speech signal and the dialect feature library.
[0114] Step a2: If the perplexity is greater than or equal to a preset perplexity threshold and the matching degree is less than a preset matching degree threshold, then determine that the influencing factors for the confidence level being less than or equal to the preset confidence level threshold include the first influencing factor.
[0115] If the perplexity of a certain sentence by the model is greater than or equal to the preset perplexity threshold, it means that the sentence may contain words not included in the corpus, resulting in a low confidence level of the recognition result.
[0116] If the matching degree between the acoustic features of a certain sentence and the dialect feature library is less than the preset matching degree threshold, it means that the sentence does not belong to a dialect, so the reason for the low confidence level of the recognition result is not caused by dialect characteristics.
[0117] Step a3: If the perplexity is less than the preset perplexity threshold and the matching degree is greater than or equal to the preset matching degree threshold, then determine that the influencing factors for the confidence level being less than or equal to the preset confidence level threshold include the second influencing factor.
[0118] If the perplexity of a certain sentence by the model is less than the preset perplexity threshold, it means that there may be no words not included in the corpus in the sentence, so the reason for the low confidence level of the recognition result is not the problem of new words.
[0119] If the matching degree between the acoustic features of a certain sentence and the dialect feature library is greater than or equal to the preset matching degree threshold, it means that the language type of the sentence is a dialect, so the reason for the low confidence level of the recognition result may be related to dialect characteristics.
[0120] Step a4: If the perplexity is greater than or equal to the preset perplexity threshold and the matching degree is greater than or equal to the preset matching degree threshold, then the influencing factors for the confidence level being less than or equal to the preset confidence level threshold include the first influencing factor and the second influencing factor.
[0121] If the perplexity of a certain sentence by the model is greater than or equal to the preset perplexity threshold, it means that it may be the influencing factor that there are words not included in the corpus in the sentence, resulting in a low confidence level of the recognition result.
[0122] And, if the matching degree between the acoustic features of a certain sentence and the dialect feature library is greater than or equal to the preset matching degree threshold, it means that the language type of the sentence is a dialect, so the reason for the low confidence level of the recognition result may involve both the problem of new words and dialect characteristics.
[0123] S602. If the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor, generate the semantics of the vocabulary and update the speech recognition model based on the semantics of the vocabulary.
[0124] Specifically, when the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor, it indicates that there are obvious new words in the sentence, but it is unlikely to be a dialect. At this time, zero-shot classification can be performed through the GPT-4V model to parse each word segment in the sentence to identify the new words and generate the semantics of the words, so as to update the speech recognition model.
[0125] S603. If the influencing factors with a confidence level less than or equal to the preset confidence threshold include the second influencing factor, perform speech enhancement processing on the dialect and update the speech recognition model based on the dialect after speech enhancement processing.
[0126] Specifically, when the influencing factors with a confidence level less than or equal to the preset confidence threshold include the second influencing factor, it means that there are no obvious new words in the sentence, but it is very likely to belong to a certain dialect. Therefore, it is first necessary to identify the corresponding dialect type, and then optimize the acoustic features according to the voice optimization strategy corresponding to the dialect type. Subsequently, obtain the speech recognition result based on the optimized acoustic features and update the speech recognition model accordingly.
[0127] S604. If the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor, generate the semantics of the vocabulary, perform speech enhancement processing on the speech signal, and update the speech recognition model based on the semantics of the vocabulary and the dialect after speech enhancement processing.
[0128] Specifically, when the influencing factors with a confidence level less than or equal to the preset confidence threshold include both the first influencing factor and the second influencing factor, it indicates that there may be new words in the sentence and it also has dialect characteristics. In this case, the above two processing methods can be adopted simultaneously, and the processing results of the two methods can be fused (combined) to obtain the final speech recognition result, and the speech recognition model is updated according to this result.
[0129] In some embodiments, there may also be a situation where the influencing factors with a confidence level less than or equal to the preset confidence threshold do not involve the first influencing factor nor the second influencing factor. This indicates that there are neither new words nor dialect characteristics in the sentence. In this case, no special processing is required.
[0130] After obtaining the speech signal with enhanced semantics and / or speech of the vocabulary, the corrected speech recognition result can be updated according to the semantics of the vocabulary to obtain an optimized speech recognition result, and the corrected speech recognition result can be updated according to the recognition result regenerated from the speech signal with enhanced speech to obtain an optimized speech recognition result.
[0131] After that, the speech signal can be used as a training sample, and the optimized speech recognition result can be used as the corresponding label, and the speech recognition model can be retrained using these data. Through this process, the model can learn a more accurate mapping relationship, thereby obtaining an updated speech recognition model.
[0132] In some embodiments, the vocabulary semantics generated by the GPT-4V model can be used to adjust the transliteration rules in the corpus. For example, for a new word originating from a network homophone, its semantic label can be set as {"source": "network homophone meme", "mapping suggestion": "network new word → homophone of existing word"}.
[0133] In addition, these semantic labels can also be used to optimize the word segmentation model. For example, for a new word, its semantic category is a network abbreviation, then during the word segmentation process, this word will be regarded as a whole instead of being split into individual Chinese characters, thereby improving the accuracy and rationality of word segmentation.
[0134] Continue to refer to Figure 5 , after updating the speech recognition model based on the corrected speech recognition result through step S502 to obtain an updated speech recognition model, step S503 can also be included.
[0135] S503. Optimize the speech recognition result based on the updated speech recognition model to obtain an optimized speech recognition result.
[0136] After completing the model update, use the optimized speech recognition model to process new speech signals, further improve the accuracy of speech recognition, and generate an optimized speech recognition result. Finally, return the optimized result to the edge side for subsequent applications or interactions.
[0137] In this embodiment, when the confidence level of the speech recognition result output by the speech recognition model is less than or equal to the preset confidence threshold, the speech recognition result output by the model is corrected, and based on the corrected speech recognition result, it is judged whether the reason for the low confidence level is due to the influence of more new words or dialect accents, so as to take targeted processing measures and use them to update the model to continuously optimize the speech recognition performance of the model.
[0138] In some embodiments, to continuously optimize the performance of the speech recognition model, the speech recognition model can also be dynamically updated through a user feedback mechanism, which specifically includes the following steps: when it is determined that the confidence level is greater than the preset confidence threshold, obtain the feedback data of the user on the speech recognition result; update the speech recognition model according to the feedback data to obtain an updated speech recognition model; deploy the updated speech recognition model on the edge side for speech recognition.
[0139] When the confidence level of the speech recognition result is higher than the preset threshold (e.g., 90%), the result can be presented to the user and the user is requested to confirm or correct it. The user can provide feedback through simple interaction methods, such as clicking the "correct" or "wrong" button, or directly entering the correct text content.
[0140] After collecting the user's feedback data, the data is screened, cleaned, and labeled. If the user confirms that the recognition result is correct, the speech sample and its corresponding recognition result are marked as "positive samples" to enhance the learning effect of the model; if the user provides a corrected result, the original speech and the corrected content are combined into "negative samples" to correct the errors of the model. Subsequently, the newly added training data will be uploaded to the cloud server for incremental training of the speech recognition model to generate an updated speech recognition model.
[0141] After the model is updated, the new speech recognition model will be deployed on edge devices (such as smartphones, smart home devices, or in-vehicle systems, etc.) to achieve lower latency and higher operating efficiency. At the same time, the performance of the deployed model can also be monitored. If it is found that the model behaves abnormally (e.g., the recognition accuracy drops significantly), the rollback mechanism is automatically triggered to restore to the previous version; if the model behaves normally, it is continuously optimized and more feedback data is accumulated to further improve the recognition ability of the model.
[0142] In this embodiment, the model is updated by combining user feedback when the confidence level is higher than the preset confidence threshold, thereby further optimizing the speech recognition performance of the model. While making full use of the reliability of the high-confidence results, the model is continuously improved with the help of user feedback, which can effectively improve the accuracy and robustness of the model for speech recognition.
[0143] Exemplary device Corresponding to the above speech recognition method, an embodiment of the present application also provides a speech recognition device. Figure 7 It is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application. As Figure 7As shown in the figure, the speech recognition device provided by the embodiment of the present application includes: a dialect type recognition unit 701, an optimization processing unit 702, and a speech recognition unit 703; wherein, the dialect type recognition unit 701 is configured to recognize the dialect type of the speech signal when determining that the language type corresponding to the speech signal of the target language is the dialect of the target language; the optimization processing unit 702 is configured to optimize the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type, and obtain optimized acoustic features, where the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and is used to improve the speech recognition effect; the speech recognition unit 703 is configured to perform speech recognition based on the optimized acoustic features to obtain the speech recognition result of the speech signal.
[0144] In some embodiments, the dialect type includes a first regional dialect type, a second regional dialect type, or a third regional dialect type; the first regional dialect type corresponds to a first speech optimization strategy, and the first speech optimization strategy is used to perform voiceless consonant enhancement processing on the acoustic features; the second regional dialect type corresponds to a second speech optimization strategy, and the second speech optimization strategy is used to perform vowel elongation processing on the acoustic features; the third regional dialect type corresponds to a third speech optimization strategy, and the third speech optimization strategy is used to perform correction processing on syllable truncation of the acoustic features.
[0145] In some embodiments, the optimization processing unit 702 optimizes the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, including: if the dialect type is the first regional dialect type, then perform voiceless consonant enhancement processing on the acoustic features to obtain the optimized acoustic features; if the dialect type is the second regional dialect type, then perform vowel elongation processing on the acoustic features to obtain the optimized acoustic features; if the dialect type is the third regional dialect type, then perform correction processing on syllable truncation of the acoustic features to obtain the optimized acoustic features.
[0146] In some embodiments, the speech recognition unit 703 performs speech recognition based on the optimized acoustic features to obtain the speech recognition result of the speech signal, including: based on the optimized acoustic features, determining the word segmentation result corresponding to the speech signal, where the word segmentation result includes each segmented word; for each segmented word among the segmented words, when it is determined that the segmented word is not a newly emerging word, based on the correspondence between the segmented word in the corpus and the transliteration result of the target language, determining the transliteration result of the target language corresponding to the segmented word, and decoding according to the transliteration result of the target language to obtain the speech recognition result; for each segmented word among the segmented words, when it is determined that the segmented word is a newly emerging word, based on the determined semantics of the segmented word, determining the transliteration result of the target language corresponding to the segmented word, and decoding according to the transliteration result of the target language corresponding to the segmented word to obtain the speech recognition result.
[0147] In some embodiments, the method is executed by a speech recognition model, and the speech recognition result includes the confidence level of the speech recognition result. The apparatus further includes: a model update unit 704, configured to perform the following steps: in the case where it is determined that the confidence level of the speech recognition result is less than or equal to a preset confidence threshold, obtaining a corrected speech recognition result; updating the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model; and optimizing the speech recognition result based on the updated speech recognition model to obtain an optimized speech recognition result.
[0148] In some embodiments, the model update unit 704 obtains the corrected speech recognition result, including: uploading the speech recognition result to a cloud server, so that the cloud server corrects the speech recognition result according to multi-modal data to obtain the corrected speech recognition result, where the multi-modal data includes corrected text data and the user's lip movement video; and receiving the corrected speech recognition result returned by the cloud server.
[0149] In some embodiments, the model update unit 704 updates the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model, including: determining, according to the corrected speech recognition result, influencing factors with a confidence level less than or equal to a preset confidence threshold, where the influencing factors include a first influencing factor or a second influencing factor; if the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor, generating the semantics of new words in the corrected speech recognition result, and updating the speech recognition model based on the semantics of the new words; if the influencing factors with a confidence level less than or equal to the preset confidence threshold include the second influencing factor, performing speech enhancement processing on the dialect, and updating the speech recognition model based on the dialect after the speech enhancement processing; if the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor, generating the semantics of the words, performing speech enhancement processing on the dialect, and updating the speech recognition model based on the semantics of the words and the dialect after the speech enhancement processing.
[0150] In some embodiments, the model update unit 704 determines, according to the corrected speech recognition result, the influencing factors with a confidence level less than or equal to the preset confidence threshold, including: determining the perplexity of the corrected speech recognition result and the matching degree between the acoustic features of the speech signal and the dialect feature library; if the perplexity is greater than or equal to a preset perplexity threshold and the matching degree is less than a preset matching degree threshold, determining that the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor; if the perplexity is less than the preset perplexity threshold and the matching degree is greater than or equal to the preset matching degree threshold, determining that the influencing factors with a confidence level less than or equal to the preset confidence threshold include the second influencing factor; if the perplexity is greater than or equal to the preset perplexity threshold and the matching degree is greater than or equal to the preset matching degree threshold, determining that the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor.
[0151] In some embodiments, the method is executed by a speech recognition model, and the speech recognition result includes the confidence level of the speech recognition result. The model update unit 704 is further configured to perform the following steps: when it is determined that the confidence level is greater than the preset confidence threshold, obtaining feedback data of the speech recognition result; updating the speech recognition model according to the feedback data to obtain an updated speech recognition model; and deploying the updated speech recognition model on the edge side for speech recognition. In some embodiments, the corpus is constructed through the following steps: establishing the correspondence between the speech data of the standard language of the target language and the recognized text, as well as the correspondence between the dialect speech data of the target language and its corresponding recognized text, to obtain an initial corpus; if the network vocabulary of the target language obtained is not included in the initial corpus, querying online for the transliteration result of the network vocabulary in the target language, and adding the correspondence between the network vocabulary and its corresponding transliteration result in the target language to the initial corpus to obtain the corpus; if the correspondence between the network vocabulary and the transliteration result in the target language cannot be queried online, generating the correspondence between the network vocabulary and its corresponding transliteration result in the target language based on the determined semantics of the network vocabulary, and adding the correspondence to the initial corpus to obtain the corpus.
[0152] In some embodiments, the speech recognition unit 703 is further configured to perform the following steps: when determining that the language type corresponding to the speech signal of the target language is the standard language of the target language, determining the word segmentation result corresponding to the speech signal based on the acoustic features of the speech signal, where the word segmentation result includes each segmented word; performing speech recognition based on the word segmentation result corresponding to the speech signal to obtain the speech recognition result of the speech signal.
[0153] The speech recognition device provided in this embodiment belongs to the same inventive concept as the speech recognition method provided in the above embodiments of the present application, can execute the speech recognition method provided in any of the above embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the speech recognition method. For the technical details not described in detail in this embodiment, reference can be made to the specific processing content of the speech recognition method provided in the above embodiments of the present application, which will not be elaborated here.
[0154] It should be noted that the functions implemented by the above dialect type determination unit 701, optimization processing unit 702, speech recognition unit 703, and model update unit 704 can be implemented by the same or different processors respectively, and the embodiments of the present application do not make any limitations.
[0155] It should be understood that the units in the above device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, instructions are stored in the memory, and the processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of a hardware circuit, and the functions of some or all of the units can be implemented through the design of the hardware circuit. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented through the design of the logical relationship of the components in the circuit. Again, for example, in another implementation, the hardware circuit can be implemented through a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file, so as to implement the functions of some or all of the above units. All units of the above device can be all implemented in the form of a processor calling software, or all implemented in the form of a hardware circuit, or some implemented in the form of a processor calling software, and the remaining part implemented in the form of a hardware circuit.
[0156] In the embodiments of the present application, the processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit, and the logical relationship of the hardware circuit is fixed or can be reconstructed. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA, etc. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as an NPU, a TPU, a DPU, etc.
[0157] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method. For example: a CPU, a GPU, an NPU, a TPU, a DPU, a microprocessor, a DSP, an ASIC, an FPGA, or a combination of at least two of these processor forms.
[0158] In addition, all or part of the units in the above device can be integrated together or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the units of the device. The types of the at least one processor may be different. For example, it includes a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0159] An embodiment of the present application also provides a vehicle equipped with the above voice recognition system for performing the steps of the above voice recognition method.
[0160] Exemplary electronic device An embodiment of the present application provides an electronic device. Refer to Figure 8 As shown, the electronic device includes: A memory 200 and a processor 210; Wherein, the memory 200 is connected to the processor 210 for storing programs; The processor 210 is configured to implement the voice recognition method disclosed in any of the above embodiments by running the programs stored in the memory 200.
[0161] Specifically, the above electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0162] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them: The bus may include a path for transmitting information between various components of the computer system.
[0163] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0164] The processor 210 may include a main processor and may also include a baseband chip, a modem, etc.
[0165] The program for implementing the technical solution of the present invention is stored in the memory 200, and the operating system and other key services may also be stored. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.
[0166] The input device 230 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0167] The output device 240 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0168] The communication interface 220 may include devices of any transceiver type for communicating with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0169] The processor 210 executes the program stored in the memory 200 and calls other devices, and can be used to implement each step of any one of the voice recognition methods provided in the above embodiments of the present application.
[0170] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs the program stored on the memory through the data interface to execute the voice recognition method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be referred to the embodiments of the voice recognition method described above.
[0171] Exemplary computer program product and storage medium In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the voice recognition method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0172] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0173] In addition, an embodiment of the present application may also be a storage medium having a computer program stored thereon, and the computer program is executed by a processor to perform the steps in the speech recognition method according to various embodiments of the present application described in any of the above embodiments of the present specification, and specifically may implement the steps in the above speech recognition method embodiments.
[0174] For the foregoing method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described order of actions, because according to the present application, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0175] It should be noted that the various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments may be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiments.
[0176] The steps in the methods of the various embodiments of the present application may be adjusted, combined, and deleted according to actual needs, and the technical features recorded in the various embodiments may be replaced or combined.
[0177] The modules and sub-modules in the devices and terminals in the various embodiments of the present application may be combined, divided, and deleted according to actual needs.
[0178] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.
[0179] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.
[0181] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0182] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0183] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0184] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A speech recognition method, characterized in that: include: When it is determined that the language type corresponding to the speech signal of the target language is a dialect of the target language, identifying the dialect type of the speech signal; According to the speech optimization strategy corresponding to the dialect type, the acoustic features of the speech signal are optimized to obtain optimized acoustic features, wherein the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and used to improve the speech recognition effect; Speech recognition is performed based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
2. The method according to claim 1, characterized in that The dialect type includes a first regional dialect type, a second regional dialect type, or a third regional dialect type; The first regional dialect type corresponds to a first speech optimization strategy, and the first speech optimization strategy is used to perform voiced consonant enhancement processing on the acoustic feature; The second regional dialect type corresponds to a second speech optimization strategy, and the second speech optimization strategy is used to perform vowel extension processing on the acoustic features; The third regional dialect type corresponds to a third voice optimization strategy, and the third voice optimization strategy is used to perform a correction process of syllable truncation on the acoustic features.
3. The method according to claim 2, characterized in that The step of performing speech optimization processing on the acoustic features according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features includes: If the dialect type is the first regional dialect type, performing voiced consonant enhancement processing on the acoustic feature to obtain the optimized acoustic feature; If the dialect type is the second regional dialect type, performing vowel lengthening processing on the acoustic feature to obtain the optimized acoustic feature; If the dialect type is the third regional dialect type, the acoustic feature is corrected by syllable truncation to obtain the optimized acoustic feature.
4. The method according to claim 1, characterized in that: The performing speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal includes: Based on the optimized acoustic features, determining a word segmentation result corresponding to the speech signal, the word segmentation result including each word segmentation; For each of the segmented words, when it is determined that the segmented word is a non-emerging word, based on the correspondence between the segmented words in the corpus and the transliteration results of the target language, the transliteration result of the target language corresponding to the segmented word is determined, and decoding is performed according to the transliteration result of the target language to obtain the speech recognition result; For each of the segmented words, when it is determined that the segmented word is an emerging word, the transliteration result of the target language corresponding to the segmented word is determined based on the determined semantics of the segmented word, and decoding is performed according to the transliteration result of the target language corresponding to the segmented word to obtain the speech recognition result.
5. The method according to any one of claims 1 to 4, characterized in that: The method is performed by a speech recognition model, the speech recognition result includes a confidence level of the speech recognition result, and the method further includes: When it is determined that the confidence level of the speech recognition result is less than or equal to a preset confidence level threshold, obtaining a corrected speech recognition result; Updating the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model; The speech recognition result is optimized based on the updated speech recognition model to obtain an optimized speech recognition result.
6. The method according to claim 5, characterized in that The obtaining of the corrected speech recognition result includes: Uploading the speech recognition result to a cloud server, so that the cloud server corrects the speech recognition result according to the multimodal data to obtain the corrected speech recognition result, wherein the multimodal data includes the corrected text data and the user's lip language video; Receive the corrected speech recognition result returned by the cloud server.
7. The method according to claim 5, characterized in that The updating of the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model includes: Determine, according to the corrected speech recognition result, an influencing factor whose confidence level is less than or equal to a preset confidence threshold, wherein the influencing factor includes a first influencing factor or a second influencing factor; If the influencing factors whose confidence level is less than or equal to the preset confidence threshold include the first influencing factor, generating semantics of new vocabulary in the revised speech recognition result, and updating the speech recognition model based on the semantics of the new vocabulary; If the influencing factors whose confidence level is less than or equal to the preset confidence threshold include the second influencing factor, performing speech enhancement processing on the dialect, and updating the speech recognition model based on the dialect after the speech enhancement processing; If the influencing factors whose confidence is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor, the semantics of the vocabulary are generated, and the dialect is speech enhanced, and the speech recognition model is updated based on the semantics of the vocabulary and the dialect after speech enhancement.
8. The method according to claim 7, characterized in that The determining, based on the corrected speech recognition result, the influencing factors that cause the confidence level to be less than or equal to a preset confidence threshold, includes: Determining the perplexity of the modified speech recognition result and the matching degree between the acoustic features of the speech signal and the dialect feature library; If the perplexity is greater than or equal to a preset perplexity threshold, and the matching degree is less than a preset matching degree threshold, determining that the influencing factor whose confidence degree is less than or equal to the preset confidence threshold includes the first influencing factor; If the perplexity is less than a preset perplexity threshold, and the matching degree is greater than or equal to a preset matching degree threshold, determining that the influencing factor whose confidence degree is less than or equal to the preset confidence threshold includes the second influencing factor; If the perplexity is greater than or equal to a preset perplexity threshold, and the matching degree is greater than or equal to a preset matching degree threshold, determining that the influencing factors whose confidence degree is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor.
9. The method according to any one of claims 1 to 4, characterized in that: The method is performed by a speech recognition model, the speech recognition result includes a confidence level of the speech recognition result, and the method further includes: When it is determined that the confidence level is greater than a preset confidence threshold, obtaining feedback data of the speech recognition result; Updating the speech recognition model according to the feedback data to obtain an updated speech recognition model; The updated speech recognition model is deployed on the edge for speech recognition.
10. The method according to claim 4, characterized in that The corpus is constructed by the following steps: Constructing a correspondence between speech data of the standard language of the target language and the recognized text, and a correspondence between speech data of the dialect of the target language and its corresponding recognized text, to obtain an initial corpus; If the acquired network vocabulary of the target language is not included in the initial corpus, querying the transliteration result of the target language corresponding to the network vocabulary through online networking, and adding the correspondence between the network vocabulary and the transliteration result of the target language corresponding to the network vocabulary to the initial corpus to obtain the corpus; If no correspondence between the network vocabulary and the transliteration result of the target language is found through online networking, a correspondence between the network vocabulary and its corresponding transliteration result of the target language is generated according to the determined semantics of the network vocabulary, and the correspondence is added to the initial corpus to obtain the corpus.
11. The method according to claim 1, characterized in that: The method further comprises: When it is determined that the language type corresponding to the speech signal of the target language is the standard language of the target language, a word segmentation result corresponding to the speech signal is determined based on the acoustic features of the speech signal, wherein the word segmentation result includes each word segmentation; Speech recognition is performed based on the word segmentation result corresponding to the speech signal to obtain a speech recognition result of the speech signal.
12. A speech recognition device, characterized in that: include: a dialect type identification unit, configured to identify the dialect type of the speech signal when determining that the language type corresponding to the speech signal of the target language is the dialect of the target language; an optimization processing unit, configured to perform speech optimization processing on the acoustic features of the speech signal according to a speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, wherein the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and used to improve speech recognition effect; A speech recognition unit is used to perform speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
13. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method according to any one of claims 1 to 11 by running the program in the memory.
14. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, cause the processor to implement the method as claimed in any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for extracting acoustic features in language identification system
CN103559879A
Multi-dialect identification method, apparatus and device and readable storage medium
CN110517664A
Speech recognition method and device, equipment and computer readable storage medium
CN113012683A
Voice changing method and device, and electronic equipment
CN113345451A
Voice signal processing method and device, electronic equipment and storage medium
CN115662397A