Speech recognition method, device, equipment and product
By identifying the dialect types of the Uyghur language and performing voice optimization processing, the problem of low accuracy in Uyghur speech recognition has been solved, and the speech recognition effect has been improved, especially the ability to recognize Chinese loanwords and new Internet words.
Patent Information
- Application Number
- CN202510652891.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Existing speech recognition technology has the problem of low speech recognition accuracy in the application of minority languages, especially Uyghur, mainly due to dialect differences, the influence of Chinese loanwords and new Internet words, resulting in poor recognition effect.
By identifying the dialect type of the speech signal and performing speech optimization processing based on the accent characteristics of the dialect type, including strategies such as voiced consonant enhancement, vowel extension, and syllable truncation correction, the quality of acoustic features is improved, and speech recognition is ultimately performed based on the optimized acoustic features.
The accuracy of speech recognition of minority languages, especially Uyghur, has been improved, and the recognition ability of Chinese loanwords and new Internet words has been enhanced to adapt to diverse language environments.
Smart Images

Figure CN120183383B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition, and in particular to a speech recognition method, apparatus, device and product. Background Art
[0002] With the continuous advancement and widespread application of intelligent speech recognition technology, the demand for its support for diverse language environments is also increasing. In particular, speech recognition technology for minority languages with fewer speakers but higher cultural value is gradually becoming a research hotspot.
[0003] While speech recognition technology has made significant progress in mainstream languages, it still suffers from low accuracy for languages with fewer speakers. Therefore, there is an urgent need to improve speech recognition for these less-speaking but culturally valuable languages. Summary of the Invention
[0004] Based on the above-mentioned technical status, the present application provides a speech recognition method, device, equipment and product, which can improve the speech recognition accuracy of languages with fewer users.
[0005] In order to achieve the above technical objectives, this application specifically proposes the following technical solutions:
[0006] According to a first aspect of an embodiment of the present application, a speech recognition method is provided, comprising: when determining that the language type corresponding to a speech signal of a target language is a dialect of the target language, identifying the dialect type of the speech signal; optimizing the acoustic features of the speech signal according to a speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, wherein the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and is used to improve the speech recognition effect; performing speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
[0007] According to a second aspect of an embodiment of the present application, a speech recognition device is provided, comprising: a dialect type recognition unit, for identifying the dialect type of the speech signal when determining that the language type corresponding to the speech signal of the target language is the dialect of the target language; an optimization processing unit, for performing speech optimization processing on the acoustic features of the speech signal according to a speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, wherein the speech optimization strategy is a strategy generated based on the accent characteristics of the dialect type and is used to improve the speech recognition effect; and a speech recognition unit, for performing speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
[0008] According to the third aspect of the embodiment of the present application, an electronic device is proposed, comprising a memory and a processor; the memory is connected to the processor and is used to store programs; the processor is used to implement the speech recognition method described in the first aspect and any one of the implementation methods of the first aspect by running the program in the memory.
[0009] According to a fourth aspect of an embodiment of the present application, a computer program product is proposed, comprising computer program instructions, which, when executed by a processor, enable the processor to implement the speech recognition method described in the first aspect and any one of the implementation methods of the first aspect.
[0010] The embodiments of the present application provide a speech recognition method, apparatus, device and product. The method identifies the dialect type of the speech signal by determining that the language type corresponding to the speech signal of the target language is the dialect of the target language, and then optimizes the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type to obtain the optimized acoustic features. Finally, speech recognition is performed based on the optimized acoustic features to obtain the final speech recognition result. Among them, the speech optimization strategy is a strategy generated according to the accent characteristics of the dialect type and is used to improve the speech recognition effect. Because the dialect type is identified before speech recognition is performed based on the acoustic features, and speech optimization processing is performed on the dialect to enhance the speech quality of the acoustic features, thereby improving the accuracy of subsequent speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0012] Figure 1 A flowchart of a speech recognition method provided in an embodiment of the present application.
[0013] Figure 2 A flowchart of speech recognition based on optimized acoustic features provided in an embodiment of the present application.
[0014] Figure 3 The present invention provides a flow chart for performing speech recognition when the speech signal provided in the embodiment of the present application is in the standard language of the target language.
[0015] Figure 4 A flowchart of the process of constructing a corpus provided in an embodiment of the present application.
[0016] Figure 5A flowchart of a confidence update model based on speech recognition results provided in an embodiment of the present application.
[0017] Figure 6 A flowchart of updating the model based on the corrected speech recognition results provided in an embodiment of the present application.
[0018] Figure 7 A schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application.
[0019] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solution provided by the embodiments of the present application can be applied to application scenarios in the fields of medical care, education or justice that require speech recognition of the target language, so as to improve the accuracy of speech recognition of the target language in these scenarios. For example, in a health consultation service scenario, users who use the target language can describe their health problems over the phone, and the speech recognition system automatically recognizes and gives corresponding suggestions or transfers them to appropriate medical staff. In a distance education platform, the online education platform can integrate the speech recognition function of the target language, allowing teachers to teach in the target language and generate subtitles or notes in real time to facilitate students to review the course content. In judicial proceedings, by introducing speech recognition technology in the target language, real-time recording of the trial process can be achieved to ensure that the statements of all participants are accurately recorded.
[0021] The technical solutions provided in the embodiments of the present application can be exemplarily applied to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged into software programs and run. When the hardware devices execute the processing of the technical solutions in the embodiments of the present application, or the above-mentioned software programs are run, the automatic splitting of target tasks and the automatic calling of the application program interfaces required for the tasks can be achieved, thereby completing the purpose of the target tasks. The embodiments of the present application only provide an illustrative introduction to the specific processing of the technical solutions in the present application, and do not limit the specific implementation form of the technical solutions in the present application. Any technical implementation form that can execute the processing of the technical solutions in the present application can be adopted by the embodiments of the present application.
[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0023] Before introducing this application solution, we first introduce the relevant technologies:
[0024] Speech recognition technology for minority languages is increasingly being applied in healthcare, education, and the legal field. It has already been deeply applied in many key areas, demonstrating its powerful enabling role. Currently, to improve recognition performance in minority languages, self-trained acoustic and language models are commonly used to optimize speech recognition systems.
[0025] When building a self-training acoustic model, a common approach to improving robustness and data security is to integrate general synthetic data into the customer's real-world scenario data to enhance recognition performance in general scenarios. This approach eliminates the need for customers to save historical data for iterative training, thereby reducing training time. Furthermore, in multilingual applications, acoustic models can be pre-trained in a high-resource language and then migrated to a smaller language environment with less data. This ensures recognition accuracy in the smaller language, effectively addressing the issue of poor recognition performance in smaller languages due to insufficient data.
[0026] A common approach for self-training language models is to collect relevant scene data and real-world conversation-level data and feed it into a semantic generalization system to generate corpus that captures the speaker's intent. This generated corpus is then mixed with real, general-purpose corpus to optimize the language model and, in turn, improve speech recognition accuracy.
[0027] Although the above strategies can improve speech recognition to a certain extent, the recognition results for minority languages (such as Uyghur) are still not ideal. The specific reasons are as follows:
[0028] (1) Traditional speech recognition systems are mainly used to recognize a single standard Uyghur language. However, in actual applications, users often mix Chinese words or phrases into their conversations, that is, Chinese loanwords frequently appear. Since the existing system is only optimized for standard Uyghur, it may not be able to correctly recognize Chinese content. In addition, Uyghur belongs to the Altaic language family, while Chinese is a tonal language (tone changes can change the meaning of words). There are significant differences between the two in terms of grammatical structure and pronunciation. Even if the speech recognition system attempts to add Chinese support, due to the fundamental differences between the two languages, especially the unique tonal system of Chinese, the recognition accuracy of the Chinese part is still low.
[0029] (2) New internet words in Uyghur are growing rapidly. For example, policy terms such as “infrastructure workers” or social media terms such as “live broadcast room” are emerging. These popular internet terms, constantly emerging new internet words, and new APP names are emerging words. Because of their rapid changes and diverse contexts, it is difficult for speech recognition systems with slower update speeds to capture these new words in real time, resulting in low speech recognition accuracy.
[0030] (3) Uyghur has a wide variety of dialects, each with its own unique voice, intonation, and pronunciation. This makes it difficult for speech recognition systems to accurately identify a particular language using a unified standard. For example, some dialects may contain unique phonemes or pronunciation habits that differ from standard Uyghur. Currently, most speech recognition engines only support dialects from a single region or a similar region, and cannot meet the needs of large-scale dialect scenarios.
[0031] In summary, current speech recognition technology cannot effectively meet the needs of speech recognition in some languages, thereby limiting the promotion and application of speech recognition technology to a wider user group.
[0032] In view of this, the embodiments of the present application are dedicated to providing a speech recognition method, apparatus, device, and product that, before performing speech recognition based on the acoustic features of a speech signal, identifies the dialect type of the speech signal and performs speech optimization processing for the dialect to enhance the quality of its acoustic features, thereby improving the accuracy of the subsequent speech recognition process and ensuring more accurate and reliable speech recognition results. These are described in detail in the following embodiments.
[0033] Exemplary Methods
[0034] Figure 1 This is a flow chart of a speech recognition method provided in an embodiment of the present application. Figure 1 As shown, the speech recognition method provided in this embodiment includes steps S101-S103:
[0035] S101: When it is determined that the language type corresponding to the speech signal of the target language is a dialect of the target language, identify the dialect type of the speech signal.
[0036] In this embodiment, the target language refers to the language to be speech recognized. The target language can be a minority language with a small number of speakers. In some examples, the target language can be Uyghur.
[0037] In some embodiments, step S101 includes: obtaining a speech signal of a target language; extracting acoustic features from the speech signal; determining the language type corresponding to the speech signal based on the acoustic features, the language type including a standard language or a dialect; if the language type corresponding to the speech signal is a dialect of the target language, identifying the dialect type corresponding to the speech signal.
[0038] Voice signals from users speaking the target language can be collected using a voice acquisition device. This device can be a microphone with high sensitivity and low noise, or other devices suitable for specific scenarios. For example, for voice acquisition in mobile environments, portable recording devices or smart terminals with integrated high-quality microphones (such as smartphones and tablets) can be used. In fixed environments, such as offices or conference rooms, desktop microphones or hanging microphone array systems can be used to effectively capture user voice signals.
[0039] After acquiring a speech signal, acoustic features can be extracted using acoustic feature extraction methods. These include Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), Perceptual Linear Prediction (PLP), and Zero-Crossing Rate (ZCR). Taking Mel-Frequency Cepstral Coefficients as an example, the speech signal is framed and windowed, and a Fast Fourier Transform (FFT) is calculated for each frame to obtain a spectrum. The spectrum is then mapped to the Mel scale, and the energy is calculated using a Mel filter bank. Finally, the logarithm of the filter bank output is taken and a discrete cosine transform (DCT) is performed to obtain MFCC features.
[0040] To improve the accuracy of speech recognition, after extracting the acoustic features of the speech signal, it is necessary to determine the language type of the speech signal based on the acoustic features, such as whether it belongs to a standard language or a dialect. Then, based on the specific language type, it is determined whether further preprocessing is required, such as speech signal enhancement processing, to improve the dialect accent problem in the speech signal, thereby improving the quality of the speech signal before performing speech recognition.
[0041] In some examples, a first language identification model can be used to determine the language type of a speech signal based on acoustic features. The first language identification model can be pre-trained based on speech samples (including standard languages and their corresponding dialects) and their corresponding real language types. The first language identification model can be a Spoken Language Identification (LID) model.
[0042] When the language type corresponding to the speech signal is a standard language, speech recognition can be performed directly based on the acoustic features to obtain a speech recognition result corresponding to the speech signal.
[0043] In the case where the language type corresponding to the speech signal is a dialect, dialects also have different types due to regional differences, so it is necessary to further determine the dialect type corresponding to the speech signal.
[0044] In some examples, a second language identification model can be used to determine the dialect type corresponding to a speech signal based on acoustic features. The second language identification model can be pre-trained based on speech samples (including different dialect types) and their corresponding real dialect types. The second language identification model can be a Spoken Language Identification (LID) model.
[0045] It should be noted that the first language recognition model and the second language recognition model can be used independently or in combination. When used independently, the first language recognition model first identifies the language type of the speech signal. If the language type is a dialect, the second language recognition model further identifies the specific dialect type. When used in combination, the first and second language recognition models can be combined into a comprehensive language recognition model. This comprehensive language recognition model can be pre-trained using speech samples (including standard language and different types of dialects) and their corresponding real language and dialect types.
[0046] S102: Perform speech optimization processing on the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features.
[0047] To improve speech recognition for dialects, you can set different speech optimization strategies for different dialect types to address accent issues in the speech signal, thereby improving speech signal quality and subsequent speech recognition accuracy. Speech optimization strategies are generated based on the accent characteristics of the dialect type and are used to improve speech recognition.
[0048] In some embodiments, the dialect type includes a first regional dialect type, a second regional dialect type, or a third regional dialect type. The first regional dialect type corresponds to a first voice optimization strategy, which is used to perform voiced consonant enhancement processing on acoustic features; the second regional dialect type corresponds to a second voice optimization strategy, which is used to perform vowel extension processing on acoustic features; and the third regional dialect type corresponds to a third voice optimization strategy, which is used to perform syllable truncation correction processing on acoustic features.
[0049] Taking the Uyghur language as an example, based on its geographical distribution, the Uyghur-speaking area can be divided into the first, second, and third regions, from north to south. Uyghur dialects in the first region are characterized by weak voiced consonants; those in the second region have relatively short vowels; and those in the third region often experience syllable weakening, omission, or linked speech, which can lead to inaccurate syllable truncation.
[0050] Given the differences in pronunciation characteristics of different regional dialect types, different voice optimization strategies can be set for them to perform voice optimization processing on different types of dialects, thereby improving the quality and clarity of voice signals.
[0051] Among them, according to different types of dialects, appropriate speech optimization strategies are selected to perform speech optimization processing on them, specifically including: if the dialect type is the first regional dialect type, the acoustic features are subjected to voiced consonant enhancement processing to obtain optimized acoustic features; if the dialect type is the second regional dialect type, the acoustic features are subjected to vowel extension processing to obtain optimized acoustic features; if the dialect type is the third regional dialect type, the acoustic features are subjected to syllable truncation correction processing to obtain optimized acoustic features.
[0052] Specifically, if the first regional dialect is detected, the first speech optimization strategy is activated. This strategy aims to enhance voiced consonants in the acoustic features. For example, the energy distribution of voiced consonants is adjusted to increase their spectral significance, making them more clearly distinguishable in the processed speech signal, thereby achieving optimized acoustic features.
[0053] If the second regional dialect is detected, the second speech optimization strategy is applied. This strategy aims to lengthen the vowels in the acoustic features. For example, the duration of specific vowel segments is increased and formant parameters are adjusted to ensure that the vowels are pronounced full and easy to recognize, thereby achieving an optimized acoustic feature.
[0054] If the detected dialect type is a third regional dialect, the third speech optimization strategy is adopted. This strategy aims to correct syllable truncation in acoustic features. By accurately analyzing and adjusting syllable boundaries in continuous speech streams, it addresses inaccurate syllable truncation caused by weakening, omitting, or connecting speech, ensuring that each syllable is correctly parsed and expressed, ultimately forming high-quality optimized acoustic features.
[0055] Through the above-mentioned targeted speech optimization processing method, the naturalness and clarity of the speech signal can be improved, which helps to improve the accuracy of subsequent speech recognition. Next, speech recognition can be performed based on the optimized acoustic features to obtain speech recognition results. For details, please refer to the detailed description of step S103 below:
[0056] S103: Perform speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
[0057] Since the speech signal may contain new words, it may not be correctly recognized. Therefore, this embodiment can also first decode based on the optimized acoustic features to obtain the initial recognition text, and then determine whether each word contained in the initial recognition text is a new word, and then select different processing strategies based on the judgment results to improve the accuracy of speech recognition of new words. The following is a detailed introduction to this implementation process with reference to the accompanying drawings:
[0058] Figure 2 This is a flow chart of speech recognition based on optimized acoustic features provided in the embodiment of this application. Figure 2 As shown, speech recognition based on optimized acoustic features includes the following steps S201-S203:
[0059] S201: Determine a word segmentation result corresponding to the speech signal based on the optimized acoustic features.
[0060] In some embodiments, step S201 includes: encoding by an encoder based on optimized acoustic features to obtain a hidden feature representation; decoding by a decoder based on the hidden feature representation to obtain an initial speech recognition result corresponding to the speech signal; performing word segmentation processing on the initial speech recognition result to obtain a word segmentation result corresponding to the speech signal.
[0061] The encoder may be a cross-language Transformer encoder, and the decoder may be a cross-language Transformer decoder.
[0062] The initial speech recognition result is the initial recognition text corresponding to the speech signal. By performing word segmentation processing on the initial recognition text, the initial recognition text can be divided into individual words to obtain the word segmentation result corresponding to the speech signal.
[0063] After obtaining the word segmentation results, each word in the word segmentation results needs to be determined to be an emerging word or a non-emerging word. Non-emerging words can be understood as regular words, that is, words that have been included in the pre-built corpus; correspondingly, emerging words include new words on the Internet, that is, words that are not included in the pre-built corpus.
[0064] In some implementations, the GPT-4V (GPT-4 with Vision) model can be used to determine whether each word is an emerging word. For example, the model can be fed with a command formatted as "Please classify the following words as emerging words or non-emerging words" to determine whether each word is an emerging word or not.
[0065] GPT-4V can be continuously trained on the latest online data, capturing language trends, including the emergence and popularity of new words. Because the model not only recognizes new words but also understands the cultural and semantic evolution behind them, it can parse their meaning and generate reasonable transliterations.
[0066] In some implementations, for each word segmentation result, a pre-built corpus can be used to determine whether it is an emerging word or a non-emerging word.
[0067] Afterwards, a suitable decoding strategy is selected according to different judgment results to perform decoding to obtain a speech recognition result. For details, please refer to the following description of steps S202 and S203.
[0068] S202: For each word segment, when it is determined that the word segment is a non-emerging word, based on the correspondence between the non-emerging word in the corpus and the transliteration result of the target language, determine the transliteration result in the target language corresponding to the word segment, and decode the transliteration result in the target language to obtain a speech recognition result.
[0069] Specifically, when a certain segmented word is determined to be a non-emerging word, it means that it has been included in the pre-constructed corpus. Based on the correspondence between the non-emerging word in the corpus and the transliteration result of the target language, the transliteration result of the target language corresponding to the segmented word can be determined directly, and the speech recognition result can be obtained based on the re-decoding.
[0070] S203: For each word in each segmentation, when it is determined that the word is an emerging word, a transliteration result of the target language corresponding to the word is determined based on the semantics of the determined word, and decoding is performed according to the transliteration result of the target language corresponding to the word to obtain a speech recognition result.
[0071] When a segmented word is identified as an emerging word, the transliteration result of the target language corresponding to the segmented word cannot be found in the pre-built corpus. Therefore, the GPT-4V model is needed to determine the source and semantics of the segmented word. Then, based on the source and semantics, the transliteration result of the segmented word in the target language is determined, and the transliteration result is decoded to obtain the speech recognition result.
[0072] In some examples, a language model (such as an N-Gram model) can be used to re-decode the speech recognition result based on the transliteration result of the target language.
[0073] This embodiment first determines whether the target language's speech signal corresponds to a specific dialect within that language. If so, it further identifies the specific dialect type of the speech signal. Based on the identified dialect type, it applies a speech optimization strategy pre-customized based on the dialect's accent characteristics to perform targeted optimization processing on the speech signal's acoustic features, thereby obtaining higher-quality acoustic features. Finally, the speech recognition process is performed using these optimized acoustic features to generate the final speech recognition results, thereby improving the accuracy of the speech recognition results.
[0074] See also Figure 3 In another embodiment of the present application, a process for implementing speech recognition when determining that the language type corresponding to the speech signal of the target language is the standard language of the target language is provided. Specifically, the process may include the following steps S301 and S302:
[0075] S301: When it is determined that the language type corresponding to the speech signal of the target language is the standard language of the target language, a word segmentation result corresponding to the speech signal is determined based on acoustic features of the speech signal.
[0076] The word segmentation result includes each word segmentation.
[0077] Before step S301 , it is necessary to obtain a speech signal of the target language; extract acoustic features from the speech signal; and determine the language type corresponding to the speech signal based on the acoustic features, where the language type includes a standard language or a dialect.
[0078] It should be noted that the specific implementation method of how to obtain the speech signal of the target language, extract its acoustic features, and determine the language type corresponding to the speech signal based on the acoustic features is similar to the specific implementation method in step S101. For details, please refer to the detailed introduction of the specific implementation method in step S101, which will not be repeated here.
[0079] In some embodiments, step S301 includes: encoding based on acoustic features by an encoder to obtain a hidden feature representation; decoding based on the hidden feature representation by a decoder to obtain an initial speech recognition result corresponding to the speech signal; and performing word segmentation processing on the initial speech recognition result to obtain a word segmentation result corresponding to the speech signal.
[0080] The encoder may be a cross-language Transformer encoder, and the decoder may be a cross-language Transformer decoder.
[0081] The initial speech recognition result is the initial recognition text corresponding to the speech signal. By performing word segmentation processing on the initial recognition text, the initial recognition text can be divided into individual words to obtain the word segmentation result corresponding to the speech signal.
[0082] After obtaining the word segmentation results, it is necessary to determine whether each word in the word segmentation results is an emerging word or a non-emerging word. Non-emerging words can be understood as regular words; correspondingly, emerging words can include new words on the Internet.
[0083] In some examples, the GPT-4V (GPT-4 with Vision) model can be used to determine whether each word is an emerging word based on a web corpus. For example, by inputting the instruction "Please classify the following words as emerging words or non-emerging words" into the model, the model will be able to determine whether each word is an emerging word or not.
[0084] GPT-4V can be trained on the latest online data and can capture language trends, including the emergence and popularity of new words. Because the model can not only identify new words but also understand the cultural and semantic evolution behind them, it can parse their meaning and generate reasonable transliterations.
[0085] S302: Perform speech recognition based on the word segmentation result corresponding to the speech signal to obtain a speech recognition result of the speech signal.
[0086] In some embodiments, step S302 includes: for each of the segmented words, when it is determined that the segmented word is a non-emerging word, based on the correspondence between the non-emerging word in the corpus and the transliteration result of the target language, determining the transliteration result of the target language corresponding to the segmented word, and decoding according to the transliteration result of the target language to obtain a speech recognition result; for each of the segmented words, when it is determined that the segmented word is an emerging word, determining the transliteration result of the target language corresponding to the segmented word based on the semantics of the determined segmented word, and decoding according to the transliteration result of the target language corresponding to the segmented word to obtain a speech recognition result.
[0087] Specifically, when a certain segmented word is determined to be a non-emerging word, it means that it has been included in the pre-constructed corpus. Based on the correspondence between the non-emerging word in the corpus and the transliteration result of the target language, the transliteration result of the target language corresponding to the segmented word can be determined directly, and the speech recognition result can be obtained based on the re-decoding.
[0088] When a segmented word is identified as an emerging word, the transliteration result of the target language corresponding to the segmented word cannot be found in the pre-built corpus. Therefore, the GPT-4V model is needed to determine the source and semantics of the segmented word. Then, based on the source and semantics, the transliteration result of the segmented word in the target language is determined, and the transliteration result is decoded to obtain the speech recognition result.
[0089] In some embodiments, if a certain word segmentation is determined to be a non-emerging word, and based on the correspondence between the non-emerging words in the corpus and the transliteration results of the target language, the transliteration result of the target language corresponding to the word segmentation cannot be determined, then you can choose to skip the word segmentation and output its corresponding recognition result in a blank format.
[0090] In some examples, a language model (such as an N-Gram model) can be used to re-decode the speech recognition result based on the transliteration result of the target language.
[0091] This embodiment performs word segmentation on the preliminary speech recognition results and determines whether each word is an emerging vocabulary. Different processing strategies are then adopted for different judgment results to improve the recognition accuracy of each word. Furthermore, the accuracy of the overall speech recognition results is further improved by re-decoding each processed word.
[0092] See also Figure 4 In another embodiment of the present application, a corpus construction process is also provided, which may specifically include the following steps S401 to S403:
[0093] S401: Constructing a correspondence between speech data of a standard language of a target language and recognized text, and a correspondence between speech data of a dialect of the target language and its corresponding recognized text, to obtain an initial corpus.
[0094] In some embodiments, the specific implementation process of establishing the correspondence between the speech data of the standard language and the recognized text may be as follows:
[0095] In some examples, by collecting speech data corresponding to the standard language of the target language in scenarios such as medical, judicial, educational, and daily communication, and using a double-blind labeling method, different labelers independently annotate it to obtain the correspondence between the speech data of the standard language of the standard language and the recognized text.
[0096] The sampling rate for standard language speech data is 16kHz, and the signal-to-noise ratio is ≥40dB. To increase data diversity, speech data from different fields can be collected. For example, speech data from the medical field, the judicial field, the educational field, and daily communication fields account for 30%, 25%, 20%, and 25%, respectively. The duration of each speech data sentence is controlled between 5-15 seconds, and the interval between sentences is greater than or equal to 1 second.
[0097] Furthermore, the voice data collection targets can also be users aged between 18 and 65 who are native speakers of the target language, with a gender ratio of 1:1 among these users.
[0098] In addition, the voice data may also be from different regions, such as a first region, a second region, and a third region.
[0099] In order to ensure the accuracy of the annotation results, the annotation consistency of different annotators must be ≥95%.
[0100] It should be noted that the standard language speech data can include both standard speech data and Chinese loanwords. For Chinese loanwords, they can be transliterated, and a correspondence between the Chinese loanwords and the transliterated results in the target language can be established. The standard speech data can also be directly annotated with text data in the target language, thereby establishing a correspondence between the standard language speech data and the text data.
[0101] In some embodiments, the specific implementation process of establishing the correspondence between dialect speech data and recognized text may be as follows:
[0102] To ensure the diversity of dialect voice data, we collected dialect voice data from different regions (such as the first, second, and third regions) on the network. This dialect voice data totaled 5,000 hours and included recordings in noisy environments such as hospitals, markets, and schools.
[0103] At the same time, users can mark incorrectly recognized words and phrases and upload them to the cloud server in real time. The cloud server corrects the text data corresponding to the dialect speech data based on the incorrectly recognized words and phrases marked by the user, and constructs a corresponding relationship between the dialect speech data and its corresponding corrected text data based on the corrected text data.
[0104] In addition, a dialect feature extraction model can be built based on dialect speech data. This model can determine the dialect type of a speech based on the acoustic characteristics of the speech data. This model can use XGBoost.
[0105] For example, consider a speech segment from a regional dialect. After processing, 25 MFCC features are extracted. When these features are fed into a trained XGBoost model, the model may output the following results: First regional dialect: 80%; Second regional dialect: 15%; Third regional dialect: 5%. This indicates that the model is most likely to identify the speech segment as being from the first regional dialect.
[0106] S402: If the acquired network vocabulary of the target language is not included in the initial corpus, the transliteration result of the target language corresponding to the network vocabulary is searched online, and the correspondence between the network vocabulary and the transliteration result of the target language corresponding to the network vocabulary is added to the initial corpus to obtain a corpus.
[0107] In step S402, a social media crawler system can be used to monitor content on official media websites, TikTok, Instagram, and other social platforms for content with an explosiveness index ≥ 7 (calculated based on the number of reposts, likes, or comments) and a usage rate ≥ 80%. This process aims to capture emerging buzzwords and phrases online, ensuring that highly popular and relevant content is promptly discovered and incorporated into the corpus, thereby updating and enriching the initial corpus.
[0108] For Internet vocabulary, if it is found that it is not included in the initial corpus, it is necessary to start an online query for the transliteration results of the Internet vocabulary, generate a correspondence between the Internet vocabulary and the transliteration results of the target language, and add it to the initial corpus.
[0109] S403. If no correspondence between the network vocabulary and the transliteration results of the target language is found through online networking, a correspondence between the network vocabulary and its corresponding transliteration results in the target language is generated according to the semantics of the determined network vocabulary, and the correspondence is added to the initial corpus to obtain a corpus.
[0110] If the correspondence between the Internet vocabulary and the transliteration results of the target language cannot be found through online networking, the semantics of the vocabulary can be further determined through the GPT-4V model, and the correspondence between the Internet vocabulary and its corresponding transliteration results in the target language can be generated, and added to the initial corpus to obtain the corpus.
[0111] For example, if no corresponding relationship is found for internet vocabulary, the GPT-4V model can be used to determine whether the internet vocabulary is an internet buzzword. For example, if the model finds that a new internet vocabulary is derived from the homophone of an existing vocabulary, it will be semantically labeled as follows: {"Source":"internet homophone pun", "Mapping suggestion":"new internet vocabulary → homophone of existing vocabulary"}. This mapping relationship is then written into the initial corpus.
[0112] It should be noted that for new network words, whether they appear in voice or text form, a correspondence between them and the transliteration results of the target language can be established. Therefore, the corpus in this embodiment can also be understood as a Chinese-Uyghur transliteration rule library.
[0113] This embodiment constructs a corpus that not only includes the correspondence between the speech data of the standard language and the recognized text, but also includes the correspondence between the speech data of Chinese loanwords and the corresponding transliteration results of the target language, as well as the correspondence between the speech data or text data of new Internet vocabulary and the corresponding transliteration results of the target language. Therefore, when performing speech recognition, the corpus can be used to obtain the transliteration results of the target language corresponding to Chinese loanwords, dialects and new Internet vocabulary, thereby helping the language model understand the speech data during decoding, thereby improving the accuracy of speech recognition.
[0114] See also Figure 5 In another embodiment of the present application, the speech recognition method of the embodiment of the present application can be executed by a speech recognition model, and the speech recognition model can be obtained by training with training samples and corresponding labels. The training samples include various speech data collected from the corpus construction process (such as standard Uyghur, standard Uyghur with Chinese loanwords, Uyghur dialects, standard Uyghur with new Internet vocabulary, and Uyghur dialects with new Internet vocabulary, etc.). The label of each training sample is the recognition text or transliteration result of the corresponding speech data generated according to the corpus. The output of the speech recognition model also includes the confidence of the speech recognition result. The model can be dynamically updated based on the confidence, which can specifically include the following steps S501 to S503:
[0115] S501: When it is determined that the confidence level of the speech recognition result is less than or equal to a preset confidence threshold, obtain a corrected speech recognition result.
[0116] The confidence level represents the degree of trust the speech recognition model places on its output, reflecting the model's assessment of the certainty of the recognition result. It is a probabilistic estimate used to indicate the likelihood that the result is correct. If the confidence level of a speech recognition result is less than or equal to the preset confidence threshold, it indicates that the speech recognition result is less accurate and the model lacks confidence in the result. Conversely, if the confidence level of a speech recognition result is greater than the preset confidence threshold, it indicates that the speech recognition result is more accurate and the model has greater confidence in its output.
[0117] When the confidence level of the speech recognition result is less than or equal to the preset confidence threshold, it is necessary to obtain the corrected speech recognition result and update the model based on it.
[0118] The confidence threshold can be set based on empirical values, with a value range of [0.9, 0.97]. For example, the values can be 0.94, 0.95, 0.96, etc.
[0119] In this embodiment, the speech recognition method may be executed by an edge device, such as a terminal device. Obtaining a corrected speech recognition result includes: uploading the speech recognition result to a cloud server, so that the cloud server corrects the speech recognition result based on multimodal data to obtain a corrected speech recognition result, where the multimodal data includes corrected text data and a video of the user's lip movements; and receiving the corrected speech recognition result returned by the cloud server.
[0120] The edge initiates a cloud API request to the cloud server to upload the speech recognition results with low confidence to the cloud server for processing, so that the cloud server can combine the corrected text and multiple modal data including images of user lip movements to improve the accuracy of the speech recognition results.
[0121] During the interactive process, users can manually modify the speech recognition results output by the model, thereby obtaining corrected text data. Image data includes a video of the user's lip readings as they speak, corresponding to the speech signal. Lip reading recognition technology can capture portions of the speech signal that may be obscured by noise or blur, providing supplementary information for speech recognition.
[0122] The cloud server comprehensively analyzes and corrects the speech recognition results based on this multimodal data, generating more accurate results and returning them to the edge. These corrections not only consider the speech signal itself but also incorporate text, images, and other related information, significantly improving the reliability of the recognition results.
[0123] Furthermore, the corrected speech recognition results can be used to update the speech recognition model to improve the performance of the model. For details, please refer to the detailed description of step S502 below:
[0124] S502: Update the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model.
[0125] See also Figure 6 In another embodiment of the present application, after obtaining the corrected speech recognition result, the speech recognition model is updated based on the corrected speech recognition result to obtain an updated speech recognition model, which may specifically include the following steps S601 to S604:
[0126] S601: Determine, based on the corrected speech recognition result, influencing factors whose confidence level is less than or equal to a preset confidence threshold.
[0127] The influencing factors include a first influencing factor or a second influencing factor. The first influencing factor includes the presence of words not included in the corpus. The second influencing factor includes the dialect.
[0128] In some embodiments, step S601 includes the following steps a1 to a4:
[0129] Step a1: determining the perplexity of the corrected speech recognition result through the language model, and determining the matching degree between the acoustic features of the speech signal and the dialect feature library.
[0130] Step a2: If the perplexity is greater than or equal to the preset perplexity threshold and the matching degree is less than the preset matching degree threshold, determine that the influencing factor whose confidence degree is less than or equal to the preset confidence threshold includes the first influencing factor.
[0131] If the model's perplexity for a sentence is greater than or equal to the preset perplexity threshold, it means that the sentence may contain words that are not included in the corpus, resulting in a lower confidence level in the recognition result.
[0132] If the match between the acoustic features of a sentence and the dialect feature library is less than the preset match threshold, it means that the sentence does not belong to the dialect, and the reason for the low confidence of the recognition result is not due to the dialect characteristics.
[0133] Step a3: If the perplexity is less than the preset perplexity threshold and the matching degree is greater than or equal to the preset matching degree threshold, then determining that the influencing factor with a confidence less than or equal to the preset confidence threshold includes the second influencing factor.
[0134] If the model's perplexity for a sentence is less than the preset perplexity threshold, it means that the sentence may not contain words that are not included in the corpus. Therefore, the reason for the low confidence level of the recognition result is not a new vocabulary problem.
[0135] If the match between the acoustic features of a sentence and the dialect feature library is greater than or equal to the preset match threshold, it means that the language type of the sentence is a dialect. Therefore, the reason for the low confidence level of the recognition result may be related to the characteristics of the dialect.
[0136] Step a4: If the perplexity is greater than or equal to the preset perplexity threshold, and the matching degree is greater than or equal to the preset matching degree threshold, the influencing factors whose confidence is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor.
[0137] If the model's perplexity for a sentence is greater than or equal to the preset perplexity threshold, it means that there may be words in the sentence that are not included in the corpus, which leads to a low confidence level in the recognition result.
[0138] Moreover, if the matching degree between the acoustic features of a sentence and the dialect feature library is greater than or equal to the preset matching degree threshold, it means that the language type of the sentence is a dialect. Therefore, the reason for the low confidence level of the recognition result may involve both new vocabulary problems and dialect characteristics.
[0139] S602: If the influencing factors with confidence levels less than or equal to a preset confidence threshold include a first influencing factor, generate semantics of the vocabulary, and update the speech recognition model based on the semantics of the vocabulary.
[0140] Specifically, when the influencing factors with a confidence level less than or equal to the preset confidence threshold include the first influencing factor, it indicates that there are obvious new words in the sentence, but it is unlikely to be a dialect. In this case, the GPT-4V model can be used for zero-shot classification to parse each word in the sentence to identify the new words and generate the semantics of the words, thereby updating the speech recognition model.
[0141] S603: If the influencing factors with a confidence level less than or equal to a preset confidence threshold include a second influencing factor, speech enhancement processing is performed on the dialect, and the speech recognition model is updated based on the dialect after the speech enhancement processing.
[0142] Specifically, when the influencing factors with a confidence level less than or equal to the preset confidence threshold include the second influencing factor, the sentence does not contain any obvious new vocabulary, but is likely to belong to a certain dialect. Therefore, the corresponding dialect type must first be identified, and then the acoustic features are optimized according to the speech optimization strategy corresponding to that dialect type. Subsequently, the speech recognition results are obtained based on the optimized acoustic features, and the speech recognition model is updated accordingly.
[0143] S604. If the influencing factors whose confidence level is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor, the semantics of the vocabulary are generated, and the speech signal is subjected to speech enhancement processing, and the speech recognition model is updated based on the semantics of the vocabulary and the dialect after the speech enhancement processing.
[0144] Specifically, when the influencing factors with a confidence level less than or equal to the preset confidence threshold include both the first and second factors, it indicates that the sentence may contain both new vocabulary and dialect characteristics. In this case, the two processing methods described above can be used simultaneously, and the processing results of the two methods can be combined to obtain the final speech recognition result, and the speech recognition model can be updated based on this result.
[0145] In some embodiments, there may be situations where an influencing factor with a confidence level less than or equal to a preset confidence threshold does not involve either the first influencing factor or the second influencing factor. This indicates that the sentence contains neither new vocabulary nor dialect characteristics. In this case, no special processing is required.
[0146] After obtaining the semantics of the vocabulary and / or the speech signal after speech enhancement, the revised speech recognition result can be updated according to the semantics of the vocabulary to obtain an optimized speech recognition result, and the revised speech recognition result can be updated according to the recognition result regenerated according to the speech signal after speech enhancement to obtain an optimized speech recognition result.
[0147] After that, the speech signal can be used as a training sample, and the optimized speech recognition results can be used as the corresponding labels. This data can be used to retrain the speech recognition model. Through this process, the model can learn a more accurate mapping relationship, thus obtaining an updated speech recognition model.
[0148] In some embodiments, the lexical semantics generated by the GPT-4V model can be used to help adjust the transliteration rules in the corpus. For example, for a new word derived from a homophone, its semantic label can be set as {"Source":"Internet homophone pun", "Mapping suggestion":"Internet new word → homophone of an existing word"}.
[0149] In addition, these semantic tags can also be used to optimize the word segmentation model. For example, for a new word, whose semantic category is network abbreviation, during word segmentation, the word will be treated as a whole rather than being split into individual Chinese characters, thereby improving the accuracy and rationality of word segmentation.
[0150] Continue reading Figure 5 After the speech recognition model is updated based on the corrected speech recognition result in step S502 to obtain an updated speech recognition model, step S503 may be further included.
[0151] S503: Optimize the speech recognition result based on the updated speech recognition model to obtain an optimized speech recognition result.
[0152] After the model update is complete, the optimized speech recognition model is used to process new speech signals, further improving speech recognition accuracy and generating optimized speech recognition results. Ultimately, the optimized results are returned to the edge for subsequent applications or interactive use.
[0153] This embodiment corrects the speech recognition result output by the speech recognition model when the confidence level of the speech recognition result is less than or equal to a preset confidence threshold, and based on the correction of the speech recognition result, determines whether the reason for the low confidence level is due to the large number of new words or the influence of the dialect accent, and takes targeted processing measures, which are used to update the model to continuously optimize the speech recognition performance of the model.
[0154] In some embodiments, in order to continuously optimize the performance of the speech recognition model, the speech recognition model can also be dynamically updated through a user feedback mechanism, which specifically includes the following steps: when it is determined that the confidence level is greater than a preset confidence threshold, obtaining user feedback data on the speech recognition results; updating the speech recognition model based on the feedback data to obtain an updated speech recognition model; and deploying the updated speech recognition model on the edge for speech recognition.
[0155] When the confidence level of the speech recognition result exceeds a preset threshold (e.g., 90%), the result can be displayed to the user, who can then be asked to confirm or correct the result. The user can provide feedback through simple interactions, such as clicking a "correct" or "wrong" button or directly entering the correct text.
[0156] After collecting user feedback, the data is screened, cleaned, and annotated. If the user confirms the recognition result is correct, the speech sample and its corresponding recognition result are marked as a "positive sample" to enhance the model's learning. If the user provides a revised result, the original speech and the revised content are combined into a "negative sample" to correct model errors. This newly added training data is then uploaded to the cloud server for incremental training of the speech recognition model to generate an updated speech recognition model.
[0157] After the model update is complete, the new speech recognition model is deployed to edge devices (such as smartphones, smart home devices, or in-vehicle systems) to achieve lower latency and higher operational efficiency. The performance of the deployed model can also be monitored. If the model performs abnormally (for example, a significant drop in recognition accuracy), a rollback mechanism is automatically triggered to restore to the previous version. If the model performs normally, continuous optimization and accumulation of feedback data will be carried out to further enhance the model's recognition capabilities.
[0158] This embodiment further optimizes the model's speech recognition performance by updating the model based on user feedback when the confidence level exceeds a preset confidence threshold. While leveraging the reliability of high-confidence results, continuously improving the model with user feedback effectively enhances the model's speech recognition accuracy and robustness.
[0159] Exemplary devices
[0160] Corresponding to the above-mentioned speech recognition method, an embodiment of the present application also provides a speech recognition device. Figure 7 This is a schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application. Figure 7As shown, the speech recognition device provided by the embodiment of the present application includes: a dialect type recognition unit 701, an optimization processing unit 702 and a speech recognition unit 703; wherein the dialect type recognition unit 701 is used to identify the dialect type of the speech signal when the language type corresponding to the speech signal of the target language is the dialect of the target language; the optimization processing unit 702 is used to optimize the acoustic features of the speech signal according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, and the speech optimization strategy is generated according to the accent characteristics of the dialect type and is used to improve the speech recognition effect; the speech recognition unit 703 is used to perform speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
[0161] In some embodiments, the dialect type includes a first regional dialect type, a second regional dialect type, or a third regional dialect type; the first regional dialect type corresponds to a first speech optimization strategy, and the first speech optimization strategy is used to perform voiced consonant enhancement processing on the acoustic features; the second regional dialect type corresponds to a second speech optimization strategy, and the second speech optimization strategy is used to perform vowel extension processing on the acoustic features; the third regional dialect type corresponds to a third speech optimization strategy, and the third speech optimization strategy is used to perform syllable truncation correction processing on the acoustic features.
[0162] In some embodiments, the optimization processing unit 702 performs speech optimization processing on the acoustic features according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, including: if the dialect type is the first regional dialect type, then performing voiced consonant enhancement processing on the acoustic features to obtain the optimized acoustic features; if the dialect type is the second regional dialect type, then performing vowel extension processing on the acoustic features to obtain the optimized acoustic features; if the dialect type is the third regional dialect type, then performing syllable truncation correction processing on the acoustic features to obtain the optimized acoustic features.
[0163] In some embodiments, the speech recognition unit 703 performs speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal, including: determining a word segmentation result corresponding to the speech signal based on the optimized acoustic features, the word segmentation result including each word segmentation; for each word segmentation, when it is determined that the word segmentation is a non-emerging word, based on the correspondence between the word segmentation in the corpus and the transliteration result of the target language, determining the transliteration result of the target language corresponding to the word segmentation, and decoding according to the transliteration result of the target language to obtain the speech recognition result; for each word segmentation, when it is determined that the word segmentation is an emerging word, based on the determined semantics of the word segmentation, determining the transliteration result of the target language corresponding to the word segmentation, and decoding according to the transliteration result of the target language corresponding to the word segmentation to obtain the speech recognition result.
[0164] In some embodiments, the method is executed by a speech recognition model, and the speech recognition result includes the confidence of the speech recognition result. The device also includes: a model updating unit 704, which is used to perform the following steps: when it is determined that the confidence of the speech recognition result is less than or equal to a preset confidence threshold, obtaining a corrected speech recognition result; updating the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model; optimizing the speech recognition result based on the updated speech recognition model to obtain an optimized speech recognition result.
[0165] In some embodiments, the model updating unit 704 obtains the corrected speech recognition result, including: uploading the speech recognition result to the cloud server so that the cloud server corrects the speech recognition result according to the multimodal data to obtain the corrected speech recognition result, and the multimodal data includes the corrected text data and the user's lip language video; receiving the corrected speech recognition result returned by the cloud server.
[0166] In some embodiments, the model updating unit 704 updates the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model, including: determining, according to the corrected speech recognition result, the influencing factors whose confidence is less than or equal to a preset confidence threshold, the influencing factors including a first influencing factor or a second influencing factor; if the influencing factors whose confidence is less than or equal to the preset confidence threshold include the first influencing factor, generating the semantics of the new vocabulary in the corrected speech recognition result, and updating the speech recognition model based on the semantics of the new vocabulary; if the influencing factors whose confidence is less than or equal to the preset confidence threshold include the second influencing factor, performing speech enhancement processing on the dialect, and updating the speech recognition model based on the dialect after the speech enhancement processing; if the influencing factors whose confidence is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor, generating the semantics of the vocabulary, performing speech enhancement processing on the dialect, and updating the speech recognition model based on the semantics of the vocabulary and the dialect after the speech enhancement processing.
[0167] In some embodiments, the model updating unit 704 determines the influencing factors that cause the confidence level to be less than or equal to a preset confidence threshold based on the corrected speech recognition result, including: determining the perplexity of the corrected speech recognition result and the degree of match between the acoustic features of the speech signal and the dialect feature library; if the perplexity is greater than or equal to the preset perplexity threshold, and the degree of match is less than the preset matching threshold, then determining that the influencing factors that cause the confidence level to be less than or equal to the preset confidence threshold include the first influencing factors; if the perplexity is less than the preset perplexity threshold, and the degree of match is greater than or equal to the preset matching threshold, then determining that the influencing factors that cause the confidence level to be less than or equal to the preset confidence threshold include the second influencing factors; if the perplexity is greater than or equal to the preset perplexity threshold, and the degree of match is greater than or equal to the preset matching threshold, then determining that the influencing factors that cause the confidence level to be less than or equal to the preset confidence threshold include the first influencing factors and the second influencing factors.
[0168] In some embodiments, the method is executed by a speech recognition model, and the speech recognition result includes the confidence of the speech recognition result. The model updating unit 704 is also used to perform the following steps: when it is determined that the confidence is greater than a preset confidence threshold, obtaining feedback data of the speech recognition result; updating the speech recognition model according to the feedback data to obtain an updated speech recognition model; and deploying the updated speech recognition model on the edge for speech recognition.
[0169] In some embodiments, the corpus is constructed by the following steps: constructing a correspondence between speech data of the standard language of the target language and the recognized text, and a correspondence between speech data of the dialect of the target language and its corresponding recognized text, to obtain an initial corpus; if the acquired network vocabulary of the target language is not included in the initial corpus, querying the transliteration result of the target language corresponding to the network vocabulary through online networking, and adding the correspondence between the network vocabulary and the transliteration result of the target language corresponding to the network vocabulary to the initial corpus to obtain the corpus; if the correspondence between the network vocabulary and the transliteration result of the target language is not queried through online networking, generating a correspondence between the network vocabulary and the transliteration result of the target language corresponding to the network vocabulary according to the determined semantics of the network vocabulary, and adding the correspondence to the initial corpus to obtain the corpus.
[0170] In some embodiments, the speech recognition unit 703 is further used to perform the following steps: when determining that the language type corresponding to the speech signal of the target language is the standard language of the target language, determining the word segmentation result corresponding to the speech signal based on the acoustic features of the speech signal, the word segmentation result including each word segmentation; performing speech recognition based on the word segmentation result corresponding to the speech signal to obtain a speech recognition result of the speech signal.
[0171] The speech recognition device provided in this embodiment is based on the same concept as the speech recognition method provided in the above embodiments of this application. It can execute the speech recognition method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of executing the speech recognition method. For technical details not fully described in this embodiment, please refer to the specific processing content of the speech recognition method provided in the above embodiments of this application, and will not be repeated here.
[0172] It should be noted that the functions implemented by the above dialect type determination unit 701, optimization processing unit 702, speech recognition unit 703 and model updating unit 704 can be implemented by the same or different processors respectively, and the embodiment of the present application is not limited thereto.
[0173] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, and the memory can be a memory within the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits. The functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units can be realized by designing the logical relationships between the components within the circuit. For another example, in another implementation, the hardware circuit can be implemented by a PLD. For example, an FPGA can include a large number of logic gate circuits. The connection relationships between the logic gate circuits are configured through a configuration file to realize the functions of some or all of the above units. All units of the above devices can be implemented entirely in the form of a processor calling software, or entirely in the form of hardware circuits, or partially in the form of a processor calling software, with the remaining parts implemented in the form of hardware circuits.
[0174] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP. In another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit may be fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0175] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0176] In addition, the various units in the above apparatus may be fully or partially integrated together, or may be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the various units of the apparatus. The at least one processor may be of different types, such as a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0177] An embodiment of the present application further provides a vehicle, which is equipped with the above-mentioned voice recognition system and is used to execute the steps of the above-mentioned voice recognition method.
[0178] Exemplary electronic devices
[0179] The present application embodiment provides an electronic device, see Figure 8 As shown, the electronic device includes:
[0180] Memory 200 and processor 210;
[0181] The memory 200 is connected to the processor 210 and is used to store programs;
[0182] The processor 210 is configured to implement the speech recognition method disclosed in any of the above embodiments by running the program stored in the memory 200 .
[0183] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0184] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are interconnected via a bus.
[0185] A bus may include a pathway that transfers information between components of a computer system.
[0186] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, or the like. It can also be an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware components.
[0187] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0188] Memory 200 stores programs that implement the technical solutions of the present invention and may also store an operating system and other key services. Specifically, the programs may include program code, which includes computer operating instructions. More specifically, memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, and the like.
[0189] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0190] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speakers, etc.
[0191] The communication interface 220 may include any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0192] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of any speech recognition method provided in the above embodiments of the present application.
[0193] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs the program stored in the memory through the data interface to execute the speech recognition method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be found in the above-mentioned embodiment introduction of the speech recognition method.
[0194] Exemplary computer program products and storage media
[0195] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the speech recognition method according to various embodiments of the present application described in any of the above-mentioned embodiments of this specification.
[0196] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0197] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to execute the steps of the speech recognition method according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically can implement the steps in the above speech recognition method embodiments.
[0198] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0199] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0200] The steps in the methods of each embodiment of the present application can be adjusted in sequence, merged, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0201] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be merged, divided, and deleted according to actual needs.
[0202] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or submodules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0203] The modules or submodules described as separate components may or may not be physically separate, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules may be selected to achieve the purpose of this embodiment according to actual needs.
[0204] In addition, each functional module or submodule in each embodiment of the present application may be integrated into a processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into a single module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or software functional modules or submodules.
[0205] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0206] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, software executed by a processor, or a combination of the two. The software may be stored in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0207] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0208] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. A speech recognition method, characterized in that: include: When determining that the language type corresponding to the speech signal of the target language is a dialect of the target language, identifying the dialect type of the speech signal; Optimizing the acoustic features of the speech signal according to a speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, wherein the speech optimization strategy is generated based on the accent characteristics of the dialect type and is used to improve speech recognition performance, and the optimization process includes voiced consonant enhancement, vowel extension, or syllable truncation correction. Speech recognition is performed based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
2. The method according to claim 1, characterized in that The dialect type includes a first regional dialect type, a second regional dialect type, or a third regional dialect type; The first regional dialect type corresponds to a first speech optimization strategy, and the first speech optimization strategy is used to perform voiced consonant enhancement processing on the acoustic feature; The second regional dialect type corresponds to a second voice optimization strategy, and the second voice optimization strategy is used to perform vowel extension processing on the acoustic features; The third regional dialect type corresponds to a third voice optimization strategy, and the third voice optimization strategy is used to perform syllable truncation correction processing on the acoustic features.
3. The method according to claim 2, characterized in that The performing speech optimization processing on the acoustic features according to the speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features includes: If the dialect type is the first regional dialect type, performing voiced consonant enhancement processing on the acoustic feature to obtain the optimized acoustic feature; If the dialect type is the second regional dialect type, performing vowel lengthening processing on the acoustic feature to obtain the optimized acoustic feature; If the dialect type is the third regional dialect type, a syllable truncation correction process is performed on the acoustic feature to obtain the optimized acoustic feature.
4. The method according to claim 1, wherein The performing speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal includes: Determining a word segmentation result corresponding to the speech signal based on the optimized acoustic features, the word segmentation result including each word segmentation; For each of the segmented words, when it is determined that the segmented word is a non-emerging word, based on the correspondence between the segmented words in the corpus and the transliteration results of the target language, the transliteration result in the target language corresponding to the segmented word is determined, and decoding is performed according to the transliteration result in the target language to obtain the speech recognition result; For each of the segmented words, when it is determined that the segmented word is an emerging word, the transliteration result of the target language corresponding to the segmented word is determined based on the determined semantics of the segmented word, and decoding is performed according to the transliteration result of the target language corresponding to the segmented word to obtain the speech recognition result.
5. The method according to any one of claims 1 to 4, characterized in that The method is performed by a speech recognition model, the speech recognition result includes a confidence level of the speech recognition result, and the method further includes: When it is determined that the confidence level of the speech recognition result is less than or equal to a preset confidence threshold, obtaining a corrected speech recognition result; updating the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model; The speech recognition result is optimized based on the updated speech recognition model to obtain an optimized speech recognition result.
6. The method according to claim 5, characterized in that The obtaining of the corrected speech recognition result includes: Uploading the speech recognition result to a cloud server, so that the cloud server corrects the speech recognition result according to the multimodal data to obtain the corrected speech recognition result, wherein the multimodal data includes the corrected text data and the user's lip language video; Receive the corrected speech recognition result returned by the cloud server.
7. The method according to claim 5, characterized in that The updating of the speech recognition model based on the corrected speech recognition result to obtain an updated speech recognition model includes: Determining, based on the corrected speech recognition result, an influencing factor whose confidence level is less than or equal to a preset confidence threshold, the influencing factor including the first influencing factor or the second influencing factor; If the influencing factors whose confidence level is less than or equal to the preset confidence threshold include the first influencing factor, generating semantics of new words in the revised speech recognition result, and updating the speech recognition model based on the semantics of the new words; If the influencing factors with a confidence level less than or equal to a preset confidence threshold include the second influencing factor, performing speech enhancement processing on the dialect, and updating the speech recognition model based on the dialect after the speech enhancement processing; If the influencing factors whose confidence is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor, the semantics of the vocabulary are generated, and the dialect is speech enhanced, and the speech recognition model is updated based on the semantics of the vocabulary and the dialect after speech enhancement processing.
8. The method according to claim 7, characterized in that The determining, based on the corrected speech recognition result, the factors affecting the confidence level being less than or equal to a preset confidence threshold, includes: Determining the perplexity of the corrected speech recognition result and the degree of match between the acoustic features of the speech signal and the dialect feature library; If the perplexity is greater than or equal to a preset perplexity threshold, and the matching degree is less than a preset matching degree threshold, determining that the influencing factors whose confidence degree is less than or equal to the preset confidence threshold include the first influencing factor; If the perplexity is less than a preset perplexity threshold, and the matching degree is greater than or equal to a preset matching degree threshold, determining that the influencing factor whose confidence degree is less than or equal to the preset confidence threshold includes the second influencing factor; If the perplexity is greater than or equal to a preset perplexity threshold, and the matching degree is greater than or equal to a preset matching degree threshold, determining that the influencing factors whose confidence degree is less than or equal to the preset confidence threshold include the first influencing factor and the second influencing factor.
9. The method according to any one of claims 1 to 4, characterized in that The method is performed by a speech recognition model, the speech recognition result includes a confidence level of the speech recognition result, and the method further includes: When it is determined that the confidence level is greater than a preset confidence threshold, obtaining feedback data of the speech recognition result; updating the speech recognition model according to the feedback data to obtain an updated speech recognition model; The updated speech recognition model is deployed on the edge for speech recognition.
10. The method according to claim 4, characterized in that The corpus is constructed using the following steps: Establishing a correspondence between speech data of the standard language of the target language and the recognized text, and a correspondence between speech data of the dialect of the target language and its corresponding recognized text, to obtain an initial corpus; If the acquired network vocabulary of the target language is not included in the initial corpus, querying the transliteration result of the target language corresponding to the network vocabulary through online networking, and adding the correspondence between the network vocabulary and the transliteration result of the target language corresponding to the network vocabulary to the initial corpus to obtain the corpus; If no correspondence between the network vocabulary and the transliteration result of the target language is found through online networking, a correspondence between the network vocabulary and its corresponding transliteration result of the target language is generated according to the determined semantics of the network vocabulary, and the correspondence is added to the initial corpus to obtain the corpus.
11. The method according to claim 1, wherein The method further comprises: When it is determined that the language type corresponding to the speech signal of the target language is the standard language of the target language, determining a word segmentation result corresponding to the speech signal based on the acoustic features of the speech signal, the word segmentation result including each word segmentation; Speech recognition is performed based on the word segmentation result corresponding to the speech signal to obtain a speech recognition result of the speech signal.
12. A speech recognition device, characterized in that: include: a dialect type identification unit, configured to identify the dialect type of the speech signal when determining that the language type corresponding to the speech signal of the target language is the dialect of the target language; an optimization processing unit, configured to perform speech optimization processing on the acoustic features of the speech signal according to a speech optimization strategy corresponding to the dialect type to obtain optimized acoustic features, wherein the speech optimization strategy is generated based on the accent characteristics of the dialect type and is used to improve speech recognition performance, and the optimization processing includes voiced consonant enhancement processing, vowel extension processing, or syllable truncation correction processing; A speech recognition unit is used to perform speech recognition based on the optimized acoustic features to obtain a speech recognition result of the speech signal.
13. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method according to any one of claims 1 to 11 by running the program in the memory.
14. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Speech recognition method and device, equipment and computer readable storage medium
CN113012683A
Semantic recognition method and device, equipment and storage medium
CN116913265A