Artificial Intelligence-Based Speech Recognition Method, System, Device, and Medium

By extracting the pitch period and speech spectrogram of the speech signal in the speech recognition system, homophone expansion and recognition verification are carried out, the inaccurate speech recognition problem caused by homophone confusion is solved, and higher recognition accuracy and language comprehension ability are achieved.

CN119943032BActive Publication Date: 2025-06-17IFLYTEK LINGZHI (JIANGSU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422269.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-17
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

It is difficult for existing speech recognition systems to accurately recognize the correctness of pronunciation vocabulary when homophones are confused.

Method used

By obtaining the speech signal to be identified, its pitch period is determined, and the speech keywords are extracted using the pitch period and the speech spectrogram. Then, each pronunciation keyword is expanded to obtain context semantic information, and the recognition ambiguity of each pronunciation keyword is calculated. Finally, the text probability distribution is analyzed through the pre-trained pronunciation recognition model to obtain the recognition results.

Benefits of technology

In the case of homophone confusion, the accuracy of speech recognition is improved, text can be generated more accurately, and the speech recognition system's ability to understand language is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943032B_ABST
    Figure CN119943032B_ABST
Patent Text Reader

Abstract

The present application provides a speech recognition method, system, device and medium based on artificial intelligence, which obtains a speech signal to be recognized; extracts keywords from an initial speech text corresponding to the speech signal based on a fundamental pitch period determined by an amplitude deviation between adjacent speech amplitudes in the speech signal and a speech spectrogram of the speech signal, to obtain a plurality of speech keywords; determines a homophone group of each speech keyword, and performs recognition verification on each speech keyword based on the context semantic information of each speech keyword in the initial speech text and the semantic features of homophones in each homophone group, to obtain the recognition ambiguity degree of each speech keyword; recognizes the text probability distribution of the speech signal according to the recognition ambiguity degree corresponding to each speech keyword and the homophone group, and analyzes the text probability distribution to obtain the text recognition result of the speech signal. The above solution can improve the recognition accuracy of speech under the confusion of homophones based on the text probability distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology. More specifically, this application relates to a speech recognition method, system, device, and medium based on artificial intelligence. Background Art

[0002] Speech recognition is an important branch in the field of artificial intelligence today. Among them, the development of speech recognition is inseparable from the key link of feature extraction. The purpose of feature extraction is to extract feature information related to speech content from the original speech signal, and at the same time eliminate noise and interference unrelated to speech content. With the development of deep learning technology, feature extraction methods based on neural networks have been proposed, which can more effectively extract high-level abstract features in speech signals to improve the attention to important speech information. These new feature extraction methods not only improve the accuracy of speech recognition, but also enhance its robustness in complex environments.

[0003] In existing speech recognition, feature extraction is the process of converting a speech signal into a feature vector that can be processed by a recognition model. First, the speech signal is segmented into short-time frames, and then the short-time frames are converted into frequency-domain signals to analyze their frequency components, so as to extract feature parameters that reflect the short-time energy and spectral shape of the speech. The feature parameters obtained by extraction capture the phoneme information of the speech and provide high-quality input for the subsequent speech recognition model. However, in existing speech recognition systems, due to the similarity and approximation of the pronunciations of a large number of words, when the speech recognition system converts a speech signal into speech text, there will be a situation of homophone confusion, which in turn leads to the inability to accurately identify the correctness of speech words during the speech recognition process. Therefore, how to improve the recognition accuracy of speech under homophone confusion has become a difficult problem faced by the industry. Summary of the Invention

[0004] This application provides a speech recognition method, system, device, and medium based on artificial intelligence, which can improve the recognition accuracy of speech under homophone confusion.

[0005] In a first aspect, this application provides a speech recognition method based on artificial intelligence, including the following steps:

[0006] Obtain a speech signal to be recognized;

[0007] Determine the fundamental period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal, and extract keywords from the initial speech text corresponding to the speech signal through the fundamental period and the speech spectrogram of the speech signal to obtain multiple speech keywords corresponding to the speech signal;

[0008] Perform homophone expansion on each speech keyword to obtain a homophone group of each speech keyword;

[0009] Obtain the context semantic information of each speech keyword in the initial speech text, and perform recognition verification on each speech keyword based on all the context semantic information and the semantic features of the homophonic words in each homophonic word group, so as to obtain the recognition ambiguity of each speech keyword in the speech recognition process;

[0010] According to the recognition ambiguity corresponding to each speech keyword and the homophonic word group, identify the text probability distribution of the speech signal, and then parse the text probability distribution through a pre-trained speech recognition model to obtain the text recognition result of the speech signal.

[0011] In some embodiments, determining the fundamental period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal specifically includes:

[0012] Perform frame splitting on the speech signal to obtain a plurality of short-time signal frames corresponding to the speech signal;

[0013] Perform amplitude calculation on each short-time signal frame to obtain the amplitude sequence of each short-time signal frame;

[0014] Calculate the amplitude deviation between adjacent speech amplitudes in each amplitude sequence;

[0015] Determine the fundamental period of the speech signal according to the deviation change characteristics of all amplitude deviations.

[0016] In some embodiments, extracting keywords from the initial speech text corresponding to the speech signal through the fundamental period and the speech spectrogram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal specifically includes:

[0017] Divide the speech signal into multiple time periods based on the fundamental period, where each time period contains one fundamental period;

[0018] Perform Fourier transform on the speech signal within each time period to obtain the speech spectrogram of each time period;

[0019] Determine the frequency band region and peak information in each speech spectrogram;

[0020] Extract a plurality of speech keywords corresponding to the speech signal from the initial speech text corresponding to the speech signal based on the frequency band region and peak information in each speech spectrogram.

[0021] In some embodiments, performing homophone expansion on each speech keyword to obtain the homophonic word group of each speech keyword specifically includes:

[0022] Perform phonetic analysis on each speech keyword to identify multiple candidate homophones corresponding to each speech keyword;

[0023] Select one speech keyword as the selected speech keyword;

[0024] Screen the multiple candidate homophones corresponding to the selected speech keyword according to phonological features, and combine the screened candidate homophones to obtain a homophone group of the selected speech keyword;

[0025] Continue to determine the homophone groups of the remaining speech keywords.

[0026] In some embodiments, identifying and verifying each speech keyword based on all context semantic information and the semantic features of the homophones in each homophone group, and obtaining the recognition ambiguity degree of each speech keyword in the speech recognition process specifically includes:

[0027] Perform semantic analysis on each speech keyword based on the context semantic information of each speech keyword to obtain the key semantic features of each speech keyword;

[0028] Select one speech keyword as the selected speech keyword;

[0029] Perform semantic association analysis on each homophone in the homophone group corresponding to the selected speech keyword according to the key semantic features of the selected speech keyword, and obtain the semantic association degree between the semantic features of each homophone in the homophone group and the key semantic features;

[0030] Determine the recognition ambiguity degree of the selected speech keyword in the speech recognition process according to all the semantic association degrees;

[0031] Continue to determine the recognition ambiguity degree of the remaining speech keywords in the speech recognition process.

[0032] In some embodiments, parsing the text probability distribution through a pre-trained speech recognition model to obtain the text recognition result of the speech signal specifically includes:

[0033] Obtain a pre-trained speech recognition model;

[0034] Input the text probability distribution into the pre-trained speech recognition model for parsing, and the speech recognition model outputs the text recognition result of the speech signal.

[0035] In some embodiments, obtain the speech signal to be recognized through an audio file.

[0036] In a second aspect, the present application provides a speech recognition system based on artificial intelligence, including:

[0037] An acquisition module, configured to acquire a voice signal to be recognized;

[0038] A processing module, configured to determine a fundamental period of the voice signal based on an amplitude deviation between adjacent voice amplitudes in the voice signal, and perform keyword extraction on an initial voice text corresponding to the voice signal through the fundamental period and a voice spectrogram of the voice signal to obtain a plurality of voice keywords corresponding to the voice signal;

[0039] The processing module is further configured to perform homophone expansion on each voice keyword to obtain a homophone group of each voice keyword;

[0040] The processing module is further configured to obtain context semantic information of each voice keyword in the initial voice text, and perform recognition verification on each voice keyword based on all the context semantic information and semantic features of homophones in each homophone group to obtain an identification ambiguity degree of each voice keyword in the voice recognition process;

[0041] An execution module, configured to identify a text probability distribution of the voice signal according to an identification ambiguity degree corresponding to each voice keyword and a homophone group, and then parse the text probability distribution through a pre-trained voice recognition model to obtain a text recognition result of the voice signal.

[0042] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to acquire the code and execute the above-mentioned artificial intelligence-based voice recognition method.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned artificial intelligence-based voice recognition method is implemented.

[0044] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects:

[0045] In the speech recognition method, system, device and medium based on artificial intelligence provided by this application, first, a speech signal to be recognized is obtained; second, a fundamental period of the speech signal is determined based on an amplitude deviation between adjacent speech amplitudes in the speech signal, and keyword extraction is performed on an initial speech text corresponding to the speech signal through the fundamental period and a speech spectrogram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal; further, homophone expansion is performed on each speech keyword to obtain a homophone group of each speech keyword; then, context semantic information of each speech keyword in the initial speech text is obtained, and recognition verification is performed on each speech keyword based on all the context semantic information and semantic features of homophones in each homophone group to obtain a recognition ambiguity degree of each speech keyword in the speech recognition process; finally, a text probability distribution of the speech signal is recognized according to the recognition ambiguity degree corresponding to each speech keyword and the homophone group, and then the text probability distribution is parsed through a pre-trained speech recognition model to obtain a text recognition result of the speech signal.

[0046] Thus, this application can improve the recognition accuracy of speech under homophone confusion; first, a speech signal to be recognized is obtained, providing a data basis for subsequent speech recognition; second, keyword extraction is performed on an initial speech text corresponding to the speech signal through the fundamental period and a speech spectrogram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal, enabling the speech recognition system to efficiently filter redundant information, accurately match the user's intention, and improve the recognition accuracy; further, homophone expansion is performed on each speech keyword to obtain a homophone group of each speech keyword, which is conducive to the speech recognition system making a more reasonable judgment on uncertain pronunciations and then recognizing the true intention of the target speech; furthermore, recognition verification is performed on each speech keyword based on the context semantic information of each speech keyword in the initial speech text and the semantic features of homophones in each homophone group to obtain a recognition ambiguity degree of each speech keyword in the speech recognition process, so as to evaluate the similarity and uncertainty of the speech recognition result, and optimize the search result based on the similarity and uncertainty of the speech recognition result, enabling the speech recognition system to provide recommended content more in line with the user's intention, thereby avoiding homophone confusion caused by the sameness and approximation of a large number of vocabulary pronunciations; then, a text probability distribution of the speech signal is recognized according to the recognition ambiguity degree corresponding to each speech keyword and the homophone group to effectively improve the language understanding ability of the speech recognition system and enable the speech signal to generate text more accurately; finally, the text probability distribution is parsed through a pre-trained speech recognition model to obtain a text recognition result of the speech signal; in summary, the technical solution provided by this application can improve the recognition accuracy of speech under homophone confusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is an exemplary flowchart of an artificial intelligence-based speech recognition method shown in some embodiments of the present application;

[0048] Figure 2 is an exemplary flowchart of determining a fundamental period shown in some embodiments of the present application;

[0049] Figure 3 is an exemplary flowchart of determining homophone phrases shown in some embodiments of the present application;

[0050] Figure 4 is a schematic structural diagram of an artificial intelligence-based speech recognition system shown in some embodiments of the present application;

[0051] Figure 5 is a schematic structural diagram of a computer device for implementing an artificial intelligence-based speech recognition method shown in some embodiments of the present application. Detailed Embodiments

[0052] To better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0053] Reference Figure 1 , this figure is an exemplary flowchart of an artificial intelligence-based speech recognition method shown in some embodiments of the present application. The artificial intelligence-based speech recognition method 100 mainly includes the following steps:

[0054] In step 101, a speech signal to be recognized is obtained.

[0055] In specific implementation, the speech signal to be recognized is obtained through an audio file. The audio file refers to a file for storing a speech signal. In the present application, the speech signal represents audio data for speech recognition processing.

[0056] In step 102, based on the amplitude deviation between adjacent speech amplitudes in the speech signal, the fundamental period of the speech signal is determined. Through the fundamental period and the speech spectrogram of the speech signal, keyword extraction is performed on the initial speech text corresponding to the speech signal to obtain multiple speech keywords corresponding to the speech signal.

[0057] In some embodiments, referring to Figure 2 shown, this figure is an exemplary flowchart of determining a fundamental period shown in some embodiments of the present application. In this embodiment, the determination of the fundamental period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal can be implemented by the following steps:

[0058] First, in step 1021, frame the speech signal to obtain a plurality of short-time signal frames corresponding to the speech signal;

[0059] Secondly, in step 1022, calculate the amplitude of each short-time signal frame to obtain the amplitude sequence of each short-time signal frame;

[0060] Then, in step 1023, calculate the amplitude deviation between adjacent speech amplitudes in each amplitude sequence;

[0061] Finally, in step 1024, determine the fundamental period of the speech signal according to the deviation change characteristics of all amplitude deviations.

[0062] When specifically implemented, first, frame the speech signal to obtain a plurality of short-time signal frames corresponding to the speech signal, that is, divide the speech signal according to a preset sliding window and sliding step to obtain a plurality of short-time signal frames corresponding to the speech signal. For example, the sliding window can be set to 20 ms and the sliding step to 10 ms. The sliding window is a Hamming window. In addition, in other embodiments, the sliding window and sliding step can also be set according to actual needs, which is not limited here; secondly, use the short-time amplitude calculation in the prior art to calculate the amplitude of each short-time signal frame to obtain the amplitude sequence of each short-time signal frame. In addition, in other embodiments, other calculation methods can also be used to calculate the amplitude of each short-time signal frame, such as short-time energy calculation, which is not limited here; then, calculate the absolute difference between adjacent speech amplitudes in each amplitude sequence, and use the absolute difference calculation result as the amplitude deviation, so as to obtain all amplitude deviations; finally, determine the fundamental period of the speech signal according to the deviation change characteristics of all amplitude deviations, that is, calculate the mean value of all amplitude deviations one by one in chronological order, and use the calculated mean value of amplitude deviations as the deviation change characteristics, so as to obtain the deviation change characteristics in different chronological orders, and extract the chronological order corresponding to the first local minimum deviation change characteristic among all deviation change characteristics as the fundamental period of the speech signal. The deviation change characteristic represents a trend for measuring the change of amplitude deviation in chronological order.

[0063] It should be noted that in this embodiment, the short-time signal frame represents the speech signal within a short time period; in this embodiment, the amplitude sequence represents a set containing multiple signal amplitudes; in this application, the fundamental period represents the basic repetition period generated by the periodic vibration of the speech signal. The fundamental period is usually expressed as the time interval between adjacent periodic signal peaks (or signal valleys). By determining the fundamental period, the robustness of speech keyword extraction can be effectively improved.

[0064] In some embodiments, the extraction of multiple speech keywords corresponding to the speech signal from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrogram of the speech signal can be implemented by the following steps, that is:

[0065] Divide the speech signal into multiple time segments based on the pitch period, where each time segment contains one pitch period;

[0066] Perform Fourier transform on the speech signal in each time segment to obtain the speech spectrogram of each time segment;

[0067] Determine the frequency band region and peak information in each speech spectrogram;

[0068] Extract multiple speech keywords corresponding to the speech signal from the initial speech text corresponding to the speech signal based on the frequency band region and peak information in each speech spectrogram.

[0069] Specifically, first, divide the speech signal into multiple time segments in chronological order based on the length of the pitch period, where each time segment contains one pitch period; second, perform Fourier transform on the speech signal in each time segment to obtain the speech spectrogram of each time segment; then, determine the frequency band region and peak information in each speech spectrogram, that is: use Mel Frequency Cepstral Coefficients to extract the frequency band region in each speech spectrogram, and use the first derivative method in the peak detection algorithm to identify the peak information in each speech spectrogram. In addition, in other embodiments, other methods can also be used to determine the frequency band region and peak information in each speech spectrogram, which are not limited here; finally, extract multiple speech keywords corresponding to the speech signal from the initial speech text corresponding to the speech signal based on the frequency band region and peak information in each speech spectrogram, that is: first, match the frequency band region and peak information in each speech spectrogram with the existing phoneme-spectrum mapping model to obtain the phoneme sequence corresponding to the speech signal, and then use the n-gram language model to match and screen the phoneme sequence with the word library in the initial speech text to obtain multiple speech keywords corresponding to the speech signal. The existing phoneme-spectrum mapping model is a phoneme-spectrum mapping based on Gaussian Mixture Model - Hidden Markov Model (GMM-HMM). In addition, in other embodiments, it can also be other mapping models, which are not limited here. In this embodiment, the initial speech text is roughly extracted from the speech signal through the Hidden Markov Model (HMM) in the speech recognition model, which will not be elaborated here.

[0070] It should be noted that in this embodiment, the voice spectrogram represents the distribution map of the voice signal in the frequency domain. Specifically, the voice spectrogram is a time-frequency-energy distribution map obtained after the voice signal is converted from the time domain to the frequency domain; in this embodiment, the frequency band region represents the energy distribution region within a specific frequency range in the voice spectrogram; in this embodiment, the peak information represents the maximum value of the amplitude at the frequency point in the voice spectrogram; in this application, the voice keyword represents a voice word with important information significance. Specifically, the voice keyword in this application is a word with important information significance extracted from the voice signal. The voice keyword is the key information word in the voice signal. By determining the voice keyword, the voice recognition system can efficiently filter redundant information, accurately match the user's intention, and improve the recognition accuracy and response speed.

[0071] In step 103, perform homophone expansion on each of the voice keywords to obtain homophone groups for each of the voice keywords.

[0072] In some embodiments, as shown in Figure 3 the figure is an exemplary flowchart for determining homophone groups according to some embodiments of the present application. In this embodiment, performing homophone expansion on each of the voice keywords to obtain homophone groups for each of the voice keywords can be implemented by the following steps:

[0073] First, in step 1031, perform phonetic analysis on each voice keyword to identify multiple candidate homophones corresponding to each voice keyword;

[0074] Secondly, in step 1032, select one voice keyword as the selected voice keyword;

[0075] Then, in step 1033, screen the multiple candidate homophones corresponding to the selected voice keyword according to phonological features, and combine the screened candidate homophones to obtain the homophone group of the selected voice keyword;

[0076] Finally, in step 1034, continue to determine the homophone groups of the remaining voice keywords.

[0077] In specific implementation, first, phonetic analysis is performed on each voice keyword to identify multiple candidate homophones corresponding to each voice keyword, that is: a phoneme segmentation tool is used to convert each voice keyword into a standard phoneme sequence, and the standard phoneme sequence represents a sequence in which the voice keyword is converted into a series of phonemes. For each voice keyword, the standard phoneme sequence of the voice keyword is mapped to a phoneme mapping library of a speech synthesis system to match words with similar pronunciations, so as to obtain multiple candidate homophones corresponding to each voice keyword. Among them, the phoneme segmentation tool can use the Pypinyin tool. In other embodiments, other phoneme segmentation tools can also be used, which are not limited here; second, a voice keyword is selected as the selected voice keyword; then, the multiple candidate homophones corresponding to the selected voice keyword are screened according to phonological features, and the screened candidate homophones are combined to obtain a homophone group of the selected voice keyword, that is: the dynamic time warping method in phonological feature analysis is used to calculate the pronunciation similarity between the selected voice keyword and each candidate homophone, and the candidate homophones with pronunciation similarity greater than a preset threshold are screened out, and the screened candidate homophones are combined to obtain a homophone group of the selected voice keyword.

[0078] It should be noted that in this embodiment, the candidate homophones refer to a group of words with similar pronunciations but different meanings from the target voice keyword; in this application, the homophone group refers to a combination of multiple homophones, that is, the homophone group is composed of multiple homophones. In speech recognition, speech recognition will be affected by noise, accents, and speech rate changes, resulting in pronunciation deviations. Therefore, by determining the homophone group, the system can make a more reasonable judgment on uncertain pronunciations, so as to recognize the true intention of the target voice.

[0079] In step 104, the context semantic information of each voice keyword in the initial voice text is obtained, and each voice keyword is identified and verified based on all the context semantic information and the semantic features of the homophones in each homophone group, so as to obtain the recognition ambiguity of each voice keyword in the speech recognition process.

[0080] In specific implementation, the context semantic information of each voice keyword in the initial voice text is obtained through semantic analysis based on a text corpus. For example, the topic to which the voice keyword belongs is analyzed from the corpus through the Latent Dirichlet Allocation (LDA) topic model, and then the context semantic information of the voice keyword in the initial voice text is judged. Details are not described here. In addition, in other embodiments, other acquisition methods can also be used to obtain the context semantic information of each voice keyword in the initial voice text. For example, methods based on language models and methods based on context windows are not limited here.

[0081] It should be noted that in this application, the context semantic information represents the association information between the current sentence and the sentences before and after. By determining the context semantic information, it is beneficial for the speech recognition system to more accurately understand the speech input, thereby reducing ambiguity and improving the accuracy of speech recognition.

[0082] In some embodiments, based on all the context semantic information and the semantic features of the homophonic words in each homophonic word group, the recognition and verification of each speech keyword are performed, and the recognition ambiguity degree of each speech keyword in the speech recognition process can be achieved by the following steps, that is:

[0083] Perform semantic analysis on each speech keyword based on the context semantic information of each speech keyword to obtain the key semantic features of each speech keyword;

[0084] Select a speech keyword as the selected speech keyword;

[0085] Perform semantic association analysis on each homophonic word in the homophonic word group corresponding to the selected speech keyword according to the key semantic features of the selected speech keyword to obtain the semantic association degree between the semantic features of each homophonic word in the homophonic word group and the key semantic features;

[0086] Determine the recognition ambiguity degree of the selected speech keyword in the speech recognition process according to all the semantic association degrees;

[0087] Continue to determine the recognition ambiguity degree of the remaining speech keywords in the speech recognition process.

[0088] In specific implementation, first, semantic analysis is performed on each speech keyword based on the context semantic information of each speech keyword to obtain the key semantic features of each speech keyword, that is: obtain a pre-trained large language model (LLM). For each speech keyword, use the context semantic information of the speech keyword as an input parameter to input into the large language model, and the large language model outputs the key semantic features of the speech keyword, thereby obtaining the key semantic features of each speech keyword. In this embodiment, the large language model uses a speech recognition model based on audio conditions (Seed-ASR). The large language model can generate the semantic features of speech keywords according to the input context information. In addition, in other embodiments, other methods can also be used to perform semantic analysis on each speech keyword, which is not limited here; secondly, semantic association analysis is performed on each homophone in the homophone group corresponding to the selected speech keyword according to the key semantic features of the selected speech keyword to obtain the semantic association degree between the semantic features of each homophone in the homophone group and the key semantic features, that is: obtain the semantic features of each homophone in the homophone group, and calculate the semantic association degree between the key semantic features of the selected speech keyword and the semantic features of each homophone through cosine similarity, thereby obtaining the semantic association degree between the semantic features of each homophone in the homophone group and the key semantic features. In addition, in other embodiments, other similarity measurement methods can also be used to calculate the semantic association degree between the semantic features of each homophone in the homophone group and the key semantic features, which is not limited here; further, determine the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic association degrees, that is: perform weighted summation on all the semantic association degrees, and use the weighted summation result as the recognition ambiguity of the selected speech keyword in the speech recognition process. Among them, the weight value of each semantic association degree can be set between 0 and 1 according to the occurrence frequency of the synonym corresponding to the speech association degree in the corpus. The higher the occurrence frequency, the larger the weight value is set, and vice versa, the smaller the weight value is set; finally, continue to determine the recognition ambiguity of the remaining speech keywords in the speech recognition process through the implementation method of "determine the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic association degrees".

[0089] It should be noted that in this embodiment, the key semantic feature is the information representing the content of the speech keyword, such as the word vector feature. By determining the semantic feature, the speech recognition system can better handle homophones and make correct judgments in combination with the context. In this embodiment, the semantic association degree represents the degree of semantic association between speech words, and it measures the degree of proximity or the strength of the connection in meaning between speech words. In this application, the recognition ambiguity degree represents the degree of recognition confusion of the speech keyword in the speech recognition process, that is, the greater the recognition ambiguity degree, the greater the recognition confusion degree of the speech keyword in the speech recognition process, and the smaller the recognition ambiguity degree, the smaller the recognition confusion degree of the speech keyword in the speech recognition process. The recognition ambiguity degree measures the severity of the speech ambiguity faced by the system when parsing the speech signal and reflects the difficulty of the system in correctly recognizing the target content. By determining the recognition ambiguity degree, the speech recognition system can evaluate the similarity and uncertainty of the speech recognition result and optimize the search result based on the similarity and uncertainty of the speech recognition result, so that the speech recognition system can provide recommended content more in line with the user's intention.

[0090] It should also be noted that in this application, the recognition verification represents the process of verifying the recognized keyword. Among them, based on all the context semantic information and the semantic features of the homophones in each homophone group, the recognition verification of each speech keyword is performed, and the recognition ambiguity degree of each speech keyword in the speech recognition process can be realized by the following steps, that is: performing semantic analysis on each speech keyword based on the context semantic information of each speech keyword to obtain the key semantic feature of each speech keyword; selecting a speech keyword as the selected speech keyword; performing semantic association analysis on each homophone in the homophone group corresponding to the selected speech keyword according to the key semantic feature of the selected speech keyword to obtain the semantic association degree between the semantic features of each homophone in the homophone group and the key semantic feature; determining the recognition ambiguity degree of the selected speech keyword in the speech recognition process according to all the semantic association degrees; continuing to determine the recognition ambiguity degree of the remaining speech keywords in the speech recognition process, that is, completing the recognition verification of each speech keyword.

[0091] In step 105, according to the recognition ambiguity degree corresponding to each speech keyword and the homophone group, the text probability distribution of the speech signal is recognized, and then the text probability distribution is parsed by a pre-trained speech recognition model to obtain the text recognition result of the speech signal.

[0092] In some embodiments, the recognition of the text probability distribution of the speech signal according to the recognition ambiguity degree corresponding to each speech keyword and the homophone group can be realized by the following steps, that is:

[0093] Obtain the recognition ambiguity degree corresponding to each speech keyword and the homophone group;

[0094] Extract the statistical probabilities of each homophone in each homophone group of each speech keyword in the given corpus;

[0095] Weight the statistical probabilities of each speech keyword based on all the recognition ambiguities to obtain the weighted statistical probabilities of each speech keyword;

[0096] Select a speech keyword as the selected speech keyword, compare the weighted statistical probability of the selected speech keyword with the statistical probabilities of each homophone in the corresponding homophone group. When the weighted statistical probability is greater than the statistical probability of each homophone in the corresponding homophone group, use the weighted statistical probability as the selected statistical probability corresponding to the selected speech keyword. When the weighted statistical probability is less than the statistical probability of any homophone in the corresponding homophone group, extract the maximum statistical probability from the corresponding homophone group as the selected statistical probability corresponding to the selected speech keyword;

[0097] Continue to determine the selected statistical probabilities corresponding to the remaining speech keywords;

[0098] Obtain the text probability distribution of the speech signal based on all the selected statistical probabilities.

[0099] In specific implementation, first, obtain the recognition ambiguity and homophone phrases corresponding to each speech keyword; second, determine the statistical probability of each homophone in each speech keyword and its homophone phrases in a given corpus, that is, the statistical probability of each homophone in each speech keyword and its homophone phrases in the given corpus can be extracted by the data processing tool Python, and the given corpus is a corpus with the highest relevance to the speech text; further, based on all the recognition ambiguities, weight the statistical probability of each speech keyword to obtain the weighted statistical probability of each speech keyword, that is: for each speech keyword, calculate the product of the recognition ambiguity corresponding to the speech keyword and the statistical probability, and use the product calculation result as the weighted statistical probability of the speech keyword, thereby obtaining the weighted statistical probability of each speech keyword; still further, select a speech keyword as the selected speech keyword, compare the weighted statistical probability of the selected speech keyword with the statistical probabilities of each homophone in the corresponding homophone phrase, when the weighted statistical probability is greater than the statistical probability of each homophone in the corresponding homophone phrase, use the weighted statistical probability as the selected statistical probability corresponding to the selected speech keyword, when the weighted statistical probability is less than the statistical probability of any homophone in the corresponding homophone phrase, then extract the maximum statistical probability from the corresponding homophone phrase as the selected statistical probability corresponding to the selected speech keyword; then, continue to determine the selected statistical probabilities corresponding to the remaining speech keywords through the determination method of "comparing the weighted statistical probability of the selected speech keyword with the statistical probabilities of each homophone in the corresponding homophone phrase, when the weighted statistical probability is greater than the statistical probability of each homophone in the corresponding homophone phrase, use the weighted statistical probability as the selected statistical probability corresponding to the selected speech keyword, when the weighted statistical probability is less than the statistical probability of any homophone in the corresponding homophone phrase, then extract the maximum statistical probability from the corresponding homophone phrase as the selected statistical probability corresponding to the selected speech keyword"; finally, combine all the selected statistical probabilities to obtain the text probability distribution of the speech signal.

[0100] It should be noted that in this embodiment, the statistical probability represents the occurrence probability of the speech keyword in the given corpus; the weighted statistical probability in this embodiment represents the statistical probability after weighted adjustment; the selected statistical probability in this embodiment represents the selected statistical probability; the text probability distribution in this application represents a set of multiple statistical probabilities, and the text probability distribution measures the distribution of the possibility that the speech keyword in the speech signal is the final speech text. By determining the text probability distribution, the language understanding ability of the speech recognition system can be effectively improved, so that the speech signal can generate text more accurately.

[0101] In some embodiments, parsing the text probability distribution through a pre-trained speech recognition model to obtain the text recognition result of the speech signal can be achieved by the following steps, that is:

[0102] Obtain a pre-trained speech recognition model;

[0103] Input the text probability distribution into the pre-trained speech recognition model for parsing, and obtain the text recognition result of the speech signal based on the output of the speech recognition model.

[0104] In specific implementation, first, obtain a pre-trained speech recognition model, where the pre-trained speech recognition model is a speech recognition model trained based on a convolutional neural network. For example, Kaldi, DeepSpeech, or Google's Speech-to-Text API; then, input the text probability distribution into the pre-trained speech recognition model for parsing, and obtain the text recognition result of the speech signal based on the output of the speech recognition model. Among them, the text probability distribution usually manifests as a set of probability estimates for different possible text results. When the pre-trained speech recognition model receives the input text probability distribution, the pre-trained speech recognition model will compare the text probability distribution with the language and speech patterns it learned during the training process, thereby generating an output. The output of the model is the text recognition result, and this recognition result represents the text content most likely corresponding to the input speech signal. The input text probability distribution provides the speech recognition model with different possibilities regarding the text that the speech signal may correspond to, and the model resolves these possibilities into the most likely text recognition result through the learned speech and language patterns.

[0105] It should be noted that in this application, the text recognition result represents the recognition result of the speech signal.

[0106] In addition, on the other hand of this application, in some embodiments, this application provides an artificial intelligence-based speech recognition system. Refer to Figure 4 , this figure is a schematic structural diagram of an artificial intelligence-based speech recognition system shown according to some embodiments of this application. The artificial intelligence-based speech recognition system 200 includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described as follows:

[0107] The acquisition module 201 is mainly used in this application to acquire the speech signal to be recognized;

[0108] The processing module 202 is mainly used in this application to determine the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal, and extract keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrogram of the speech signal to obtain multiple speech keywords corresponding to the speech signal;

[0109] The processing module 202 is further configured to perform homophone expansion on each of the voice keywords to obtain a homophone phrase group for each of the voice keywords;

[0110] In addition, the processing module 202 is further configured to obtain the context semantic information of each voice keyword in the initial voice text, and perform recognition verification on each voice keyword based on all the context semantic information and the semantic features of the homophones in each homophone phrase group, so as to obtain the recognition ambiguity degree of each voice keyword in the speech recognition process;

[0111] The execution module 203. In this application, the execution module 203 is mainly configured to identify the text probability distribution of the voice signal according to the recognition ambiguity degree and the homophone phrase group corresponding to each voice keyword, and then parse the text probability distribution through a pre-trained speech recognition model to obtain the text recognition result of the voice signal.

[0112] In addition, this application also provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned speech recognition method based on artificial intelligence.

[0113] In some embodiments, refer to Figure 5 This figure is a schematic structural diagram of a computer device for implementing the speech recognition method based on artificial intelligence according to some embodiments of this application. The speech recognition method based on artificial intelligence in the above embodiments can be implemented by Figure 5 The computer device shown. The computer device 300 includes at least one processor 301, a communication bus 302, a memory 303, and at least one communication interface 304.

[0114] The processor 301 may be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the speech recognition method based on artificial intelligence in this application.

[0115] The communication bus 302 can be used to transmit information between the above components.

[0116] The memory 303 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 303 can exist independently and be connected to the processor 301 through the communication bus 302. The memory 303 can also be integrated with the processor 301.

[0117] Among them, the memory 303 is used to store the program code for executing the solution of this application, and is controlled by the processor 301 for execution. The processor 301 is used to execute the program code stored in the memory 303. The program code can include one or more software modules. The determination of the speech recognition method based on artificial intelligence in the above embodiments can be implemented by one or more software modules in the processor 301 and the program code in the memory 303.

[0118] The communication interface 304 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0119] In a specific implementation, as an embodiment, the computer device can include multiple processors, and each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0120] The computer device described above can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.

[0121] In addition, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned speech recognition method based on artificial intelligence is implemented.

[0122] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0123] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A speech recognition method based on artificial intelligence, characterized in that: The steps include: Acquire a speech signal to be recognized; Determine the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal, extract keywords from the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum of the speech signal, and obtain a plurality of speech keywords corresponding to the speech signal; Performing homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; Acquire contextual semantic information of each voice keyword in the initial voice text, perform recognition verification on each voice keyword based on all the contextual semantic information and semantic features of homophones in each homophone group, and obtain recognition ambiguity of each voice keyword in the voice recognition process, wherein the recognition ambiguity indicates the degree of recognition confusion of the voice keyword in the voice recognition process; Identify the text probability distribution of the voice signal according to the recognition ambiguity and homophones corresponding to each voice keyword, and then parse the text probability distribution through a pre-trained voice recognition model to obtain a text recognition result of the voice signal; Among them, each speech keyword is recognized and verified based on all contextual semantic information and semantic features of homophones in each homophone group, and the recognition ambiguity of each speech keyword in the speech recognition process is obtained, which specifically includes: Performing semantic analysis on each voice keyword based on contextual semantic information of each voice keyword to obtain key semantic features of each voice keyword; Select a voice keyword as the selected voice keyword; Performing semantic association analysis on each homophone in the homophone group corresponding to the selected voice keyword according to the key semantic feature of the selected voice keyword, and obtaining the semantic association degree between the semantic feature of each homophone in the homophone group and the key semantic feature; Determining the recognition ambiguity of the selected speech keyword in the speech recognition process according to all the semantic associations; Continue to determine the recognition ambiguity of the remaining speech keywords in the speech recognition process.

2. The method according to claim 1, characterized in that Determining the pitch period of the speech signal based on the amplitude deviation between adjacent speech amplitudes in the speech signal specifically includes: Performing frame processing on the speech signal to obtain a plurality of short-time signal frames corresponding to the speech signal; Calculating the amplitude of each short-time signal frame to obtain an amplitude sequence of each short-time signal frame; Calculate the amplitude deviation between adjacent speech amplitudes in each amplitude sequence; The pitch period of the speech signal is determined according to the deviation variation characteristics of all amplitude deviations.

3. The method according to claim 1, characterized in that Keyword extraction is performed on the initial speech text corresponding to the speech signal through the pitch period and the speech spectrum diagram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal, specifically including: Dividing the speech signal into a plurality of time periods based on the pitch period, wherein each time period includes a pitch period; Performing Fourier transform on the speech signal in each time period to obtain a speech spectrogram of each time period; Determine the frequency band area and peak information in each speech spectrogram; A plurality of speech keywords corresponding to the speech signal are extracted from the initial speech text corresponding to the speech signal based on the frequency band area and peak information in each speech spectrogram.

4. The method according to claim 1, characterized in that Performing homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords specifically includes: Performing phonetic analysis on each speech keyword to identify multiple candidate homophones corresponding to each speech keyword; Select a voice keyword as the selected voice keyword; Screening multiple candidate homophones corresponding to the selected speech keyword according to the phonological features, and combining the screened candidate homophones to obtain a homophone group of the selected speech keyword; Continue to determine homophone phrases for the remaining phonetic keywords.

5. The method according to claim 1, characterized in that The text probability distribution is parsed by a pre-trained speech recognition model to obtain a text recognition result of the speech signal, which specifically includes: Get a pre-trained speech recognition model; The text probability distribution is input into a pre-trained speech recognition model for parsing, and the speech recognition model outputs a text recognition result of the speech signal.

6. The method according to claim 1, characterized in that The speech signal to be recognized is obtained through the audio file.

7. A speech recognition system based on artificial intelligence, which adopts the method described in any one of claims 1 to 6 for speech recognition, characterized in that: The system includes: An acquisition module, used for acquiring a speech signal to be recognized; A processing module, configured to determine a pitch period of the speech signal based on an amplitude deviation between adjacent speech amplitudes in the speech signal, and extract keywords from an initial speech text corresponding to the speech signal through the pitch period and a speech spectrum diagram of the speech signal to obtain a plurality of speech keywords corresponding to the speech signal; The processing module is further used to perform homophone expansion on each of the voice keywords to obtain a homophone group of each of the voice keywords; The processing module is further used to obtain contextual semantic information of each voice keyword in the initial voice text, and to perform recognition verification on each voice keyword based on all the contextual semantic information and the semantic features of homophones in each homophone group, so as to obtain the recognition ambiguity of each voice keyword in the voice recognition process; The execution module is used to identify the text probability distribution of the voice signal according to the recognition ambiguity and homophone phrases corresponding to each voice keyword, and then parse the text probability distribution through a pre-trained voice recognition model to obtain the text recognition result of the voice signal.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores codes, and the processor is configured to obtain the codes and execute the artificial intelligence-based speech recognition method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the artificial intelligence-based speech recognition method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Voice recognition method and device, storage medium and terminal

    CN108962232A

  • Homonym instruction recognition processing method, equipment and device

    CN117113981A