Speech recognition method, system and device based on artificial intelligence and medium

By extracting the pitch period and speech spectrogram of the speech signal in the speech recognition system, homophone expansion and recognition ambiguity calculation are performed, the problem of degradation of speech recognition accuracy caused by homophone confusion is solved, and higher recognition accuracy and robustness are achieved.

CN119943032AActive Publication Date: 2025-05-06IFLYTEK LINGZHI (JIANGSU) TECH CO LTD

Patent Information

Application Number
CN202510422269.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing speech recognition systems are difficult to accurately recognize the correctness of pronunciation vocabulary when confusing homophones, resulting in a decrease in recognition accuracy.

Method used

By obtaining the speech signal to be identified, its pitch period is determined, and the speech keywords are extracted based on the speech spectrogram. Then, homophone expansion is performed on each pronunciation keyword, context semantic information is obtained, recognition ambiguity is calculated, and finally the text probability distribution is analyzed through the pre-trained speech recognition model to obtain recognition results.

Benefits of technology

In the case of homophone confusion, the accuracy of speech recognition is improved, and by efficiently filtering redundant information and accurately matching user intentions, the recognition ambiguity is reduced and the robustness of the speech recognition system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943032A_ABST
    Figure CN119943032A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method, system and device based on artificial intelligence and a medium. The method comprises the steps of obtaining a to-be-recognized voice signal; performing keyword extraction on an initial voice text corresponding to the voice signal based on a pitch period determined by an amplitude deviation between adjacent voice amplitudes in the voice signal and a voice spectrogram of the voice signal to obtain a plurality of voice keywords; determining a homonym group of each voice keyword, and performing recognition verification on each voice keyword based on context semantic information of each voice keyword in the initial voice text and semantic features of homonyms in each homonym group to obtain recognition ambiguity of each voice keyword; and recognizing text probability distribution of the voice signal according to the recognition ambiguity and the homonym group corresponding to each voice keyword, and analyzing the text probability distribution to obtain a text recognition result of the voice signal. According to the scheme, based on the text probability distribution, the speech recognition accuracy can be improved under homonym confusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and more specifically, to a speech recognition method, system, device and medium based on artificial intelligence. Background Art

[0002] Speech recognition is an important branch of the field of artificial intelligence today. The development of speech recognition is inseparable from the key link of feature extraction. The purpose of feature extraction is to extract feature information related to the speech content from the original speech signal, while eliminating noise and interference irrelevant to the speech content. With the development of deep learning technology, feature extraction methods based on neural networks have been proposed, which can more effectively extract high-level abstract features from speech signals to increase the focus on important speech information. These new feature extraction methods not only improve the accuracy of speech recognition, but also enhance its robustness in complex environments.

[0003] In existing speech recognition, feature extraction is the process of converting speech signals into feature vectors that can be processed by the recognition model. First, the speech signal is segmented into short-time frames, and then the short-time frames are converted into frequency domain signals to analyze their frequency components in order to extract feature parameters that reflect the short-time energy and spectral shape of the speech. The extracted feature parameters capture the phoneme information of the speech and provide high-quality input for the subsequent speech recognition model. However, in the existing speech recognition system, due to the sameness and similarity of the pronunciation of a large number of words, the speech recognition system may confuse homophones when converting speech signals into speech text, which in turn leads to the inability to accurately identify the correctness of the speech vocabulary during the speech recognition process. Therefore, how to improve the recognition accuracy of speech under homophone confusion has become a difficult problem faced by the industry. Summary of the invention

[0004] The present application provides a speech recognition method, system, device and medium based on artificial intelligence, which can improve the recognition accuracy of speech under homophone confusion.

[0005] In a first aspect, the present application provides a speech recognition method based on artificial intelligence, comprising the following steps: Acquire a speech signal to be recognized; Determine the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal, extract keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum of the speech signal, and obtain a plurality of speech keywords corresponding to the speech signal; Performing homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; Acquire contextual semantic information of each voice keyword in the initial voice text, perform recognition verification on each voice keyword based on all the contextual semantic information and semantic features of homophones in each homophone group, and obtain recognition ambiguity of each voice keyword in the voice recognition process; The text probability distribution of the speech signal is identified according to the recognition ambiguity and homophones corresponding to each speech keyword, and then the text probability distribution is analyzed by a pre-trained speech recognition model to obtain a text recognition result of the speech signal.

[0006] In some embodiments, determining the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal specifically includes: Performing frame processing on the speech signal to obtain a plurality of short-time signal frames corresponding to the speech signal; Calculating the amplitude of each short-time signal frame to obtain an amplitude sequence of each short-time signal frame; Calculate the amplitude deviation between adjacent speech amplitudes in each amplitude sequence; The pitch period of the speech signal is determined according to the deviation variation characteristics of all amplitude deviations.

[0007] In some embodiments, extracting keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrogram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal specifically includes: Dividing the speech signal into a plurality of time periods based on the pitch period, wherein each time period includes a pitch period; Performing Fourier transform on the speech signal in each time period to obtain a speech spectrogram of each time period; Determine the frequency band area and peak information in each speech spectrogram; A plurality of speech keywords corresponding to the speech signal are extracted from the initial speech text corresponding to the speech signal based on the frequency band area and peak information in each speech spectrogram.

[0008] In some embodiments, performing homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords specifically includes: Performing phonetic analysis on each speech keyword to identify multiple candidate homophones corresponding to each speech keyword; Select a voice keyword as the selected voice keyword; Screening multiple candidate homophones corresponding to the selected speech keyword according to the phonological features, and combining the screened candidate homophones to obtain a homophone group of the selected speech keyword; Continue to determine homophone phrases for the remaining phonetic keywords.

[0009] In some embodiments, each speech keyword is recognized and verified based on all contextual semantic information and semantic features of homophones in each homophone group, and obtaining the recognition ambiguity of each speech keyword in the speech recognition process specifically includes: Performing semantic analysis on each voice keyword based on contextual semantic information of each voice keyword to obtain key semantic features of each voice keyword; Select a voice keyword as the selected voice keyword; Performing semantic association analysis on each homophone in the homophone group corresponding to the selected voice keyword according to the key semantic feature of the selected voice keyword, and obtaining the semantic association degree between the semantic feature of each homophone in the homophone group and the key semantic feature; Determining the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic associations; Continue to determine the recognition ambiguity of the remaining speech keywords in the speech recognition process.

[0010] In some embodiments, parsing the text probability distribution through a pre-trained speech recognition model to obtain a text recognition result of the speech signal specifically includes: Get a pre-trained speech recognition model; The text probability distribution is input into a pre-trained speech recognition model for parsing, and the speech recognition model outputs a text recognition result of the speech signal.

[0011] In some embodiments, the speech signal to be recognized is obtained through an audio file.

[0012] In a second aspect, the present application provides a speech recognition system based on artificial intelligence, comprising: An acquisition module, used for acquiring a speech signal to be recognized; A processing module, configured to determine a pitch period of the speech signal based on an amplitude deviation between adjacent speech amplitudes in the speech signal, and extract keywords from an initial speech text corresponding to the speech signal through the pitch period and a speech spectrum diagram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal; The processing module is further used to perform homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; The processing module is further used to obtain contextual semantic information of each voice keyword in the initial voice text, and to perform recognition verification on each voice keyword based on all the contextual semantic information and the semantic features of homophones in each homophone group, so as to obtain the recognition ambiguity of each voice keyword in the voice recognition process; The execution module is used to identify the text probability distribution of the voice signal according to the recognition ambiguity and homophone phrases corresponding to each voice keyword, and then parse the text probability distribution through a pre-trained voice recognition model to obtain the text recognition result of the voice signal.

[0013] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores codes, and the processor is configured to obtain the codes and execute the above-mentioned artificial intelligence-based speech recognition method.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned artificial intelligence-based speech recognition method is implemented.

[0015] The technical solution provided by the embodiments disclosed in this application has the following beneficial effects: In the artificial intelligence-based speech recognition method, system, device and medium provided in the present application, first, a speech signal to be recognized is obtained; secondly, the pitch period of the speech signal is determined based on the amplitude deviation between adjacent speech amplitudes in the speech signal, and keywords are extracted from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum diagram of the speech signal to obtain multiple speech keywords corresponding to the speech signal; further, homophone expansion is performed on each of the speech keywords to obtain homophone phrases of each speech keyword; then, contextual semantic information of each speech keyword in the initial speech text is obtained, and recognition verification is performed on each speech keyword based on all contextual semantic information and semantic features of homophones in each homophone phrase to obtain recognition ambiguity of each speech keyword in the speech recognition process; finally, the text probability distribution of the speech signal is identified according to the recognition ambiguity corresponding to each speech keyword and the homophone phrase, and then the text probability distribution is parsed through a pre-trained speech recognition model to obtain a text recognition result of the speech signal.

[0016] It can be seen that the present application can improve the recognition accuracy of speech under homophone confusion; first, obtain the speech signal to be recognized to provide a data basis for subsequent speech recognition; secondly, extract keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum diagram of the speech signal, and obtain multiple speech keywords corresponding to the speech signal, so that the speech recognition system can efficiently filter redundant information, accurately match user intentions, and improve recognition accuracy; further, perform homophone expansion on each speech keyword to obtain a homophone group for each speech keyword, which is conducive to the speech recognition system to make more reasonable judgments on uncertain pronunciations, and then identify the true intention of the target speech; further, based on the contextual semantic information of each speech keyword in the initial speech text and the semantic features of the homophones in each homophone group, each speech keyword is expanded. Keywords are recognized and verified to obtain the recognition ambiguity of each voice keyword in the voice recognition process, so as to evaluate the similarity and uncertainty of the voice recognition results, and optimize the search results based on the similarity and uncertainty of the voice recognition results, so that the voice recognition system can provide recommended content that is more in line with the user's intention, thereby avoiding the confusion of homophones caused by the sameness and similarity of the pronunciation of a large number of words; then, the text probability distribution of the voice signal is identified according to the recognition ambiguity and homophone groups corresponding to each voice keyword, so as to effectively improve the voice recognition system's ability to understand the language and enable the voice signal to generate text more accurately; finally, the text probability distribution is analyzed through the pre-trained voice recognition model to obtain the text recognition result of the voice signal; in summary, the technical solution provided by the present application can improve the recognition accuracy of voice under homophone confusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an exemplary flow chart of a speech recognition method based on artificial intelligence according to some embodiments of the present application; Figure 2 is an exemplary flow chart of determining a pitch period according to some embodiments of the present application; Figure 3 is an exemplary flow chart of determining homophone phrases according to some embodiments of the present application; Figure 4 is a schematic diagram of the structure of a speech recognition system based on artificial intelligence according to some embodiments of the present application; Figure 5 It is a structural diagram of a computer device for implementing an artificial intelligence-based speech recognition method according to some embodiments of the present application. DETAILED DESCRIPTION

[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0019] refer to Figure 1 , which is an exemplary flow chart of a speech recognition method based on artificial intelligence according to some embodiments of the present application. The speech recognition method 100 based on artificial intelligence mainly includes the following steps: In step 101, a speech signal to be recognized is obtained.

[0020] In a specific implementation, the voice signal to be recognized is obtained through an audio file, where the audio file refers to a file used to store the voice signal. In this application, the voice signal refers to audio data used for voice recognition processing.

[0021] In step 102, the pitch period of the speech signal is determined based on the amplitude deviation between adjacent speech amplitudes in the speech signal, and keywords are extracted from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum diagram of the speech signal to obtain multiple speech keywords corresponding to the speech signal.

[0022] In some embodiments, reference Figure 2 As shown in FIG. 1 , this figure is an exemplary flow chart of determining the pitch period according to some embodiments of the present application. In this embodiment, determining the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal can be implemented by the following steps: First, in step 1021, the speech signal is framed to obtain a plurality of short-time signal frames corresponding to the speech signal; Secondly, in step 1022, the amplitude of each short-time signal frame is calculated to obtain an amplitude sequence of each short-time signal frame; Then, in step 1023, the amplitude deviation between adjacent speech amplitudes in each amplitude sequence is calculated; Finally, in step 1024, the pitch period of the speech signal is determined according to the deviation variation characteristics of all amplitude deviations.

[0023] In a specific implementation, first, the voice signal is framed to obtain a plurality of short-time signal frames corresponding to the voice signal, that is, the voice signal is divided according to a preset sliding window and a sliding step to obtain a plurality of short-time signal frames corresponding to the voice signal. For example, the sliding window can be set to 20ms and the sliding step to 10ms, wherein the sliding window is a Hamming window. In addition, in other embodiments, the sliding window and the sliding step can be set according to actual needs, which are not limited here. Secondly, the short-time amplitude calculation in the prior art is used to calculate the amplitude of each short-time signal frame to obtain an amplitude sequence of each short-time signal frame. In addition, in other embodiments, other calculation methods can be used to calculate the amplitude of each short-time signal frame. Value calculation, for example, short-time energy calculation, which is not limited here; then, the absolute difference of adjacent speech amplitudes in each amplitude sequence is calculated, and the absolute difference calculation result is used as the amplitude deviation, thereby obtaining all the amplitude deviations; finally, the fundamental pitch period of the speech signal is determined according to the deviation change characteristics of all the amplitude deviations, that is: all the amplitude deviations are averaged one by one in time sequence, and the calculated amplitude deviation mean is used as the deviation change feature, thereby obtaining the deviation change features under different time sequences, and the time sequence corresponding to the first local minimum deviation change feature among all the deviation change features is extracted as the fundamental pitch period of the speech signal, and the deviation change feature representation is used to measure the changing trend of the amplitude deviation under time sequence.

[0024] It should be noted that, in this embodiment, the short-time signal frame represents a speech signal within a short period of time; the amplitude sequence in this embodiment represents a set of multiple signal amplitudes; the fundamental frequency period in this application represents the basic repetition period generated by the periodic vibration of the speech signal, and the fundamental frequency period is usually expressed as the time interval between adjacent periodic signal peaks (or signal valleys). By determining the fundamental frequency period, the robustness of speech keyword extraction can be effectively improved.

[0025] In some embodiments, extracting keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrogram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal can be achieved by the following steps, namely: Dividing the speech signal into a plurality of time periods based on the pitch period, wherein each time period includes a pitch period; Performing Fourier transform on the speech signal in each time period to obtain a speech spectrogram of each time period; Determine the frequency band area and peak information in each speech spectrogram; A plurality of speech keywords corresponding to the speech signal are extracted from the initial speech text corresponding to the speech signal based on the frequency band area and peak information in each speech spectrogram.

[0026] In a specific implementation, first, the speech signal is divided into a plurality of time periods in chronological order based on the length of the fundamental pitch period, wherein each time period includes a fundamental pitch period; secondly, the speech signal in each time period is subjected to Fourier transform to obtain a speech spectrogram of each time period; then, the frequency band area and peak information in each speech spectrogram are determined, namely: the frequency band area in each speech spectrogram is obtained by extracting the Mel-frequency cepstral coefficient, and the peak information in each speech spectrogram is obtained by identifying the first-order derivative method in the peak detection algorithm. In addition, in other embodiments, other methods may be used to determine the frequency band area and peak information in each speech spectrogram, which are not limited here; finally, the initial speech text corresponding to the speech signal is extracted based on the frequency band area and peak information in each speech spectrogram. A plurality of speech keywords corresponding to the speech signal are obtained, namely: first, the frequency band area and peak information in each speech spectrogram are matched with the existing phoneme-spectrum mapping model to obtain the phoneme sequence corresponding to the speech signal, and then the n-gram language model is used to match and screen the phoneme sequence with the word library in the initial speech text to obtain a plurality of speech keywords corresponding to the speech signal, wherein the existing phoneme-spectrum mapping model is a phoneme-spectrum mapping based on Gaussian mixture model-hidden Markov model (GMM-HMM). In addition, in other embodiments, other mapping models may also be used, which are not limited here. The initial speech text in this embodiment is obtained by roughly extracting the speech signal through the hidden Markov model (HMM) in the speech recognition model, which will not be repeated here.

[0027] It should be noted that, in this embodiment, the speech spectrum diagram represents the distribution diagram of the speech signal in the frequency domain. Specifically, the speech spectrum diagram is a time-frequency-energy distribution diagram obtained after the speech signal is converted from the time domain to the frequency domain; in this embodiment, the frequency band area represents the energy distribution area within a specific frequency range in the speech spectrum diagram; in this embodiment, the peak information represents the maximum value of the amplitude in the speech spectrum diagram at the frequency point; in this application, the speech keyword represents a speech word with important information significance. Specifically, the speech keyword in this application is a word with important information significance extracted from the speech signal. The speech keyword is the key information word in the speech signal. By determining the speech keyword, the speech recognition system can efficiently filter redundant information, accurately match user intentions, and improve recognition accuracy and response speed.

[0028] In step 103, homophone expansion is performed on each of the voice keywords to obtain a homophone group of each of the voice keywords.

[0029] In some embodiments, reference Figure 3As shown, this figure is an exemplary flow chart of determining homophone groups according to some embodiments of the present application. In this embodiment, homophone expansion is performed on each of the voice keywords to obtain the homophone groups of each of the voice keywords, which can be achieved by using the following steps: First, in step 1031, a phonetic analysis is performed on each voice keyword to identify a plurality of candidate homophones corresponding to each voice keyword; Next, in step 1032, a voice keyword is selected as a selected voice keyword; Then, in step 1033, multiple candidate homophones corresponding to the selected speech keyword are screened according to the phonological features, and the screened candidate homophones are combined to obtain a homophone group of the selected speech keyword; Finally, in step 1034, the homophone phrases of the remaining speech keywords continue to be determined.

[0030] In the specific implementation, first, a phonetic analysis is performed on each voice keyword to identify multiple candidate homophones corresponding to each voice keyword, that is, a phoneme segmentation tool is used to convert each voice keyword into a standard phoneme sequence, wherein the standard phoneme sequence represents a sequence of a series of phonemes converted from the voice keyword, and for each voice keyword, the standard phoneme sequence of the voice keyword is mapped to a phoneme mapping library of a speech synthesis system to match words with similar pronunciations, thereby obtaining multiple candidate homophones corresponding to each voice keyword, wherein the phoneme segmentation tool can use the Pypinyin tool, and in other embodiments , other phoneme segmentation tools can also be used, which are not limited here; secondly, a speech keyword is selected as the selected speech keyword; then, multiple candidate homophones corresponding to the selected speech keyword are screened according to the phonological features, and the screened candidate homophones are combined to obtain the homophone group of the selected speech keyword, that is: the dynamic time warping method in the phonological feature analysis is used to calculate the pronunciation similarity of the selected speech keyword and each candidate homophone, and the candidate homophones with pronunciation similarity greater than a preset threshold are screened out, and the screened candidate homophones are combined to obtain the homophone group of the selected speech keyword.

[0031] It should be noted that in this embodiment, the candidate homophones represent a group of words that are similar in pronunciation to the target voice keywords but have different meanings; in this application, a homophone group represents a combination of multiple homophones, that is, the homophone group is composed of multiple homophones. In speech recognition, speech recognition will be affected by noise, accent, and changes in speaking speed, resulting in pronunciation deviations. Therefore, by determining the homophone group, the system can make more reasonable judgments on uncertain pronunciations, thereby identifying the true intention of the target speech.

[0032] In step 104, contextual semantic information of each voice keyword in the initial voice text is obtained, and recognition verification is performed on each voice keyword based on all the contextual semantic information and the semantic features of the homophones in each homophone group to obtain the recognition ambiguity of each voice keyword in the voice recognition process.

[0033] In a specific implementation, the contextual semantic information of each speech keyword in the initial speech text is obtained through semantic analysis based on a text corpus. For example, the topic to which the speech keyword belongs is analyzed from the corpus through a latent Dirichlet allocation topic model (LDA), and then the contextual semantic information of the speech keyword in the initial speech text is determined. This will not be repeated here. In addition, in other embodiments, other acquisition methods can also be used to obtain the contextual semantic information of each speech keyword in the initial speech text, such as language model-based and context window-based methods, which are not limited here.

[0034] It should be noted that the contextual semantic information in the present application represents the association information between the current sentence and the previous and next sentences. By determining the contextual semantic information, the speech recognition system can understand the speech input more accurately, thereby reducing ambiguity and improving the accuracy of speech recognition.

[0035] In some embodiments, each speech keyword is recognized and verified based on all contextual semantic information and semantic features of homophones in each homophone group, and the recognition ambiguity of each speech keyword in the speech recognition process can be obtained by the following steps, namely: Performing semantic analysis on each voice keyword based on contextual semantic information of each voice keyword to obtain key semantic features of each voice keyword; Select a voice keyword as the selected voice keyword; Performing semantic association analysis on each homophone in the homophone group corresponding to the selected voice keyword according to the key semantic feature of the selected voice keyword, and obtaining the semantic association degree between the semantic feature of each homophone in the homophone group and the key semantic feature; Determining the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic associations; Continue to determine the recognition ambiguity of the remaining speech keywords in the speech recognition process.

[0036] In the specific implementation, first, semantic analysis is performed on each voice keyword based on the contextual semantic information of each voice keyword to obtain the key semantic features of each voice keyword, that is, a pre-trained large language model (LLM) is obtained, and for each voice keyword, the contextual semantic information of the voice keyword is input into the large language model as an input parameter, and the large language model outputs the key semantic features of the voice keyword, and then the key semantic features of each voice keyword are obtained. In this embodiment, the large language model adopts a speech recognition model based on audio conditions (Seed-ASR), and the large language model can generate semantic features of voice keywords according to the input context information. In addition, in other embodiments, other methods can be used to perform semantic analysis on each voice keyword, which is not limited here; secondly, semantic association analysis is performed on each homophone in the homophone group corresponding to the selected voice keyword according to the key semantic features of the selected voice keyword, and the semantic association between the semantic features of each homophone in the homophone group and the key semantic features is obtained, that is, the semantic features of each homophone in the homophone group are obtained. The semantic association between the key semantic features of the selected speech keyword and the semantic features of each homonym is calculated by cosine similarity, thereby obtaining the semantic association between the semantic features of each homonym in the homonym group and the key semantic features. In addition, in other embodiments, other similarity measurement methods can also be used to calculate the semantic association between the semantic features of each homonym in the homonym group and the key semantic features, which is not limited here. Further, the recognition ambiguity of the selected speech keyword in the speech recognition process is determined according to all the semantic associations, that is, all the semantic associations are weighted and summed, and the weighted sum result is used as the recognition ambiguity of the selected speech keyword in the speech recognition process, wherein the weight value of each semantic association can be set between 0 and 1 according to the occurrence frequency of the synonyms corresponding to the speech association in the corpus, the higher the occurrence frequency, the larger the weight value is set, and vice versa, the smaller the weight value is set; finally, the recognition ambiguity of the remaining speech keywords in the speech recognition process is continued to be determined by the implementation method of "determining the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic associations".

[0037] It should be noted that the key semantic features in this embodiment are information that characterizes the content of speech keywords, such as word vector features. By determining the semantic features, the speech recognition system can better handle homophones and make correct judgments in combination with the context; in this embodiment, the semantic association degree represents the degree of semantic association between speech words, and the semantic association degree measures the degree of proximity or connection strength between speech words in meaning; in this application, the recognition ambiguity degree represents the degree of recognition confusion of speech keywords in the speech recognition process, that is, the greater the recognition ambiguity degree, the greater the recognition confusion degree of speech keywords in the speech recognition process, and the smaller the recognition ambiguity degree, the smaller the recognition confusion degree of speech keywords in the speech recognition process. The recognition ambiguity degree measures the severity of speech ambiguity faced by the system when parsing speech signals, and reflects the difficulty of the system in correctly identifying the target content. By determining the recognition ambiguity degree, the speech recognition system can evaluate the similarity and uncertainty of speech recognition results, and optimize the search results based on the similarity and uncertainty of speech recognition results, so that the speech recognition system can provide recommended content that is more in line with user intentions.

[0038] It should also be noted that the recognition verification in the present application refers to the process of verifying the recognized keywords, wherein the recognition verification is performed on each voice keyword based on all contextual semantic information and the semantic features of the homophones in each homophone group, and the recognition ambiguity of each voice keyword in the voice recognition process can be obtained by the following steps, namely: semantic analysis is performed on each voice keyword based on the contextual semantic information of each voice keyword to obtain the key semantic features of each voice keyword; a voice keyword is selected as the selected voice keyword; semantic association analysis is performed on each homophone in the homophone group corresponding to the selected voice keyword according to the key semantic features of the selected voice keyword to obtain the semantic association between the semantic features of each homophone in the homophone group and the key semantic features; the recognition ambiguity of the selected voice keyword in the voice recognition process is determined according to all the semantic associations; the recognition ambiguity of the remaining voice keywords in the voice recognition process is continued to be determined, that is, the recognition verification of each voice keyword is completed.

[0039] In step 105, the text probability distribution of the speech signal is identified according to the recognition ambiguity and homophones corresponding to each speech keyword, and then the text probability distribution is analyzed by a pre-trained speech recognition model to obtain a text recognition result of the speech signal.

[0040] In some embodiments, identifying the text probability distribution of the speech signal according to the recognition ambiguity corresponding to each speech keyword and the homophone phrase can be implemented by the following steps, namely: Obtain recognition ambiguity and homophones corresponding to each voice keyword; Extract the statistical probability of each phonetic keyword and each homonym in its homonym group in a given corpus; Weighting the statistical probability of each voice keyword based on all recognition ambiguities to obtain the weighted statistical probability of each voice keyword; A speech keyword is selected as the selected speech keyword, and the weighted statistical probability of the selected speech keyword is compared with the statistical probability of each homophone in the corresponding homophone group. When the weighted statistical probability is greater than the statistical probability of each homophone in the corresponding homophone group, the weighted statistical probability is used as the selected statistical probability corresponding to the selected speech keyword. When the weighted statistical probability is less than the statistical probability of any homophone in the corresponding homophone group, the maximum statistical probability is extracted from the corresponding homophone group as the selected statistical probability corresponding to the selected speech keyword. Continue to determine the selected statistical probabilities corresponding to the remaining voice keywords; A text probability distribution of the speech signal is obtained based on all selected statistical probabilities.

[0041] In the specific implementation, first, the recognition ambiguity and homophone groups corresponding to each voice keyword are obtained; secondly, the statistical probability of each voice keyword and each homophone in its homophone group in a given corpus is determined, that is, the statistical probability of each voice keyword and each homophone in its homophone group in a given corpus can be extracted through the data processing tool Python, and the given corpus is a corpus with the greatest correlation with the voice text; further, the statistical probability of each voice keyword is weighted based on all the recognition ambiguities to obtain the weighted statistical probability of each voice keyword, that is, for each voice keyword, the recognition ambiguity and the statistical probability corresponding to the voice keyword are multiplied, and the product calculation result is used as the weighted statistical probability of the voice keyword, thereby obtaining the weighted statistical probability of each voice keyword; further, a voice keyword is selected as the selected voice keyword, and the weighted statistical probability of the selected voice keyword is compared with the statistical probability of each homophone in the corresponding homophone group. When the weighted statistical When the probability is greater than the statistical probability of each homophone in the corresponding homophone group, the weighted statistical probability is used as the selected statistical probability corresponding to the selected voice keyword; when the weighted statistical probability is less than the statistical probability of any homophone in the corresponding homophone group, the maximum statistical probability is extracted from the corresponding homophone group as the selected statistical probability corresponding to the selected voice keyword; then, the selected statistical probability corresponding to the remaining voice keywords is continued to be determined by the determination method of "comparing the weighted statistical probability of the selected voice keyword with the statistical probability of each homophone in the corresponding homophone group; when the weighted statistical probability is greater than the statistical probability of each homophone in the corresponding homophone group, the weighted statistical probability is used as the selected statistical probability corresponding to the selected voice keyword; when the weighted statistical probability is less than the statistical probability of any homophone in the corresponding homophone group, the maximum statistical probability is extracted from the corresponding homophone group as the selected statistical probability corresponding to the selected voice keyword"; finally, all the selected statistical probabilities are combined to obtain the text probability distribution of the voice signal.

[0042] It should be noted that the statistical probability in this embodiment represents the probability of occurrence of a speech keyword in a given corpus; the weighted statistical probability in this embodiment represents the statistical probability after weighted adjustment; the selected statistical probability in this embodiment represents the selected statistical probability; the text probability distribution in this application represents a collection of multiple statistical probabilities, and the text probability distribution measures the probability distribution of the speech keyword in the speech signal as the final speech text. By determining the text probability distribution, the speech recognition system's ability to understand the language can be effectively improved, so that the speech signal can generate text more accurately.

[0043] In some embodiments, the text probability distribution is parsed by a pre-trained speech recognition model to obtain a text recognition result of the speech signal, which can be achieved by the following steps, namely: Get a pre-trained speech recognition model; The text probability distribution is input into a pre-trained speech recognition model for analysis, and a text recognition result of the speech signal is obtained based on the output of the speech recognition model.

[0044] In the specific implementation, first, a pre-trained speech recognition model is obtained, and the pre-trained speech recognition model is a speech recognition model trained based on a convolutional neural network, for example, Kaldi, DeepSpeech or Google's Speech-to-Text API; then, the text probability distribution is input into the pre-trained speech recognition model for analysis, and the text recognition result of the speech signal is obtained based on the output of the speech recognition model, wherein the text probability distribution is usually expressed as a set of probability estimates for different possible text results. When the pre-trained speech recognition model receives the input text probability distribution, the pre-trained speech recognition model compares the text probability distribution with the language and speech patterns learned during the training process to generate an output. The output of the model is a text recognition result, which represents the text content that the input speech signal is most likely to correspond to. The input text probability distribution provides the speech recognition model with different possibilities about the text that the speech signal may correspond to, and the model parses these possibilities into the most likely text recognition result through the learned speech and language patterns.

[0045] It should be noted that the text recognition result in this application refers to the recognition result of the speech signal.

[0046] In addition, in another aspect of the present application, in some embodiments, the present application provides a speech recognition system based on artificial intelligence, referring to Figure 4 , which is a schematic diagram of the structure of a speech recognition system based on artificial intelligence according to some embodiments of the present application. The speech recognition system based on artificial intelligence 200 includes: an acquisition module 201, a processing module 202 and an execution module 203, which are described as follows: Acquisition module 201, in this application, acquisition module 201 is mainly used to acquire a speech signal to be recognized; The processing module 202 in the present application is mainly used to determine the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal, and extract keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum of the speech signal to obtain multiple speech keywords corresponding to the speech signal; The processing module 202 is further used to perform homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; In addition, the processing module 202 is further used to obtain contextual semantic information of each voice keyword in the initial voice text, and to perform recognition verification on each voice keyword based on all the contextual semantic information and the semantic features of the homophones in each homophone group, so as to obtain the recognition ambiguity of each voice keyword in the voice recognition process; Execution module 203. In the present application, execution module 203 is mainly used to identify the text probability distribution of the voice signal according to the recognition ambiguity and homophones corresponding to each voice keyword, and then parse the text probability distribution through a pre-trained voice recognition model to obtain the text recognition result of the voice signal.

[0047] In addition, the present application also provides a computer device, which includes a memory and a processor, the memory stores a code, and the processor is configured to obtain the code and execute the above-mentioned artificial intelligence-based speech recognition method.

[0048] In some embodiments, reference Figure 5 , which is a schematic diagram of the structure of a computer device for implementing an artificial intelligence-based speech recognition method according to some embodiments of the present application. The artificial intelligence-based speech recognition method in the above embodiment can be Figure 5 The computer device 300 shown in the figure is implemented, and the computer device 300 includes at least one processor 301, a communication bus 302, a memory 303 and at least one communication interface 304.

[0049] The processor 301 may be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC) or one or more for controlling the execution of the artificial intelligence-based speech recognition method in the present application.

[0050] The communication bus 302 may be used to transmit information between the above-mentioned components.

[0051] The memory 303 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compressed optical disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 303 may exist independently and be connected to the processor 301 via the communication bus 302. The memory 303 may also be integrated with the processor 301.

[0052] The memory 303 is used to store the program code for executing the solution of the present application, and the execution is controlled by the processor 301. The processor 301 is used to execute the program code stored in the memory 303. The program code may include one or more software modules. The determination of the speech recognition method based on artificial intelligence in the above embodiment can be implemented by the processor 301 and one or more software modules in the program code in the memory 303.

[0053] The communication interface 304 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0054] In a specific implementation, as an embodiment, a computer device may include multiple processors, each of which may be a single-CPU processor or a multi-CPU processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0055] The above-mentioned computer device may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device or an embedded device. The embodiment of the present application does not limit the type of computer device.

[0056] In addition, the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned artificial intelligence-based speech recognition method is implemented.

[0057] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0058] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A speech recognition method based on artificial intelligence, characterized in that: The steps include: Acquire a speech signal to be recognized; Determine the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal, extract keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum of the speech signal, and obtain a plurality of speech keywords corresponding to the speech signal; Performing homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; Acquire contextual semantic information of each voice keyword in the initial voice text, perform recognition verification on each voice keyword based on all the contextual semantic information and semantic features of homophones in each homophone group, and obtain recognition ambiguity of each voice keyword in the voice recognition process; The text probability distribution of the speech signal is identified according to the recognition ambiguity and homophones corresponding to each speech keyword, and then the text probability distribution is analyzed by a pre-trained speech recognition model to obtain a text recognition result of the speech signal.

2. The method according to claim 1, characterized in that Determining the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal specifically includes: Performing frame processing on the speech signal to obtain a plurality of short-time signal frames corresponding to the speech signal; Calculating the amplitude of each short-time signal frame to obtain an amplitude sequence of each short-time signal frame; Calculate the amplitude deviation between adjacent speech amplitudes in each amplitude sequence; The pitch period of the speech signal is determined according to the deviation variation characteristics of all amplitude deviations.

3. The method according to claim 1, characterized in that Keyword extraction is performed on the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum diagram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal, specifically including: Dividing the speech signal into a plurality of time periods based on the pitch period, wherein each time period includes a pitch period; Performing Fourier transform on the speech signal in each time period to obtain a speech spectrogram of each time period; Determine the frequency band area and peak information in each speech spectrogram; A plurality of speech keywords corresponding to the speech signal are extracted from the initial speech text corresponding to the speech signal based on the frequency band area and peak information in each speech spectrogram.

4. The method according to claim 1, characterized in that Performing homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords specifically includes: Performing phonetic analysis on each speech keyword to identify multiple candidate homophones corresponding to each speech keyword; Select a voice keyword as the selected voice keyword; Screening multiple candidate homophones corresponding to the selected speech keyword according to the phonological features, and combining the screened candidate homophones to obtain a homophone group of the selected speech keyword; Continue to determine homophone phrases for the remaining phonetic keywords.

5. The method according to claim 1, characterized in that Based on all contextual semantic information and the semantic features of homophones in each homophone group, each voice keyword is recognized and verified, and the recognition ambiguity of each voice keyword in the voice recognition process is obtained, which specifically includes: Performing semantic analysis on each voice keyword based on contextual semantic information of each voice keyword to obtain key semantic features of each voice keyword; Select a voice keyword as the selected voice keyword; Performing semantic association analysis on each homophone in the homophone group corresponding to the selected voice keyword according to the key semantic feature of the selected voice keyword, and obtaining the semantic association degree between the semantic feature of each homophone in the homophone group and the key semantic feature; Determining the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic associations; Continue to determine the recognition ambiguity of the remaining speech keywords in the speech recognition process.

6. The method according to claim 1, characterized in that The text probability distribution is parsed by a pre-trained speech recognition model to obtain a text recognition result of the speech signal, which specifically includes: Get a pre-trained speech recognition model; The text probability distribution is input into a pre-trained speech recognition model for parsing, and the speech recognition model outputs a text recognition result of the speech signal.

7. The method according to claim 1, characterized in that The speech signal to be recognized is obtained through the audio file.

8. A speech recognition system based on artificial intelligence, characterized in that: include: An acquisition module, used for acquiring a speech signal to be recognized; A processing module, configured to determine a pitch period of the speech signal based on an amplitude deviation between adjacent speech amplitudes in the speech signal, and extract keywords from an initial speech text corresponding to the speech signal through the pitch period and a speech spectrum diagram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal; The processing module is further used to perform homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; The processing module is further used to obtain contextual semantic information of each voice keyword in the initial voice text, and to perform recognition verification on each voice keyword based on all the contextual semantic information and the semantic features of homophones in each homophone group, so as to obtain the recognition ambiguity of each voice keyword in the voice recognition process; The execution module is used to identify the text probability distribution of the voice signal according to the recognition ambiguity and homophone phrases corresponding to each voice keyword, and then parse the text probability distribution through a pre-trained voice recognition model to obtain the text recognition result of the voice signal.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores codes, and the processor is configured to obtain the codes and execute the artificial intelligence-based speech recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the artificial intelligence-based speech recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice search processing method and device of homonym

    CN105279227A

  • Keyword expansion method and keyword expansion system

    CN106294396A

  • Information recognition method and equipment, storage medium and terminal

    CN108304375A

  • Human-computer interaction method and device

    CN108920497A

  • Voice recognition method and device, storage medium and terminal

    CN108962232A

Cited By

  • Intelligent customer service voice interaction method and system based on voice recognition

    CN120199247A