A speech recognition method and device based on a microcomputer
By extracting hierarchical time domain and frequency domain features on microcomputers and combining multi-class cascaded filter processing, the problem of low accuracy of traditional speech recognition in complex environments is solved, and efficient speech recognition in noisy environments is achieved.
Patent Information
- Application Number
- CN202510635399.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Traditional speech recognition methods have low recognition accuracy and significantly reduced performance when background noise is high or signal quality is not high.
A microcomputer-based speech recognition method is adopted to extract hierarchical time domain and frequency domain features, combine the time domain acoustic model and the frequency domain acoustic model to generate phoneme sequences, and use a multi-class cascade filter to process the speech signal when the phoneme sequence is inconsistent, improving the spectrum characteristics of the signal.
The accuracy and reliability of speech recognition are significantly improved in complex environments, and the anti-interference ability of the system is enhanced, especially in the context of low signal-to-noise ratio and high noise.
Smart Images

Figure CN120148484B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech processing, and particularly relates to a speech recognition method and device based on a microcomputer. Background Art
[0002] As a low-power, compact and high-computing-capability computing platform, a microcomputer is widely used in embedded systems. Compared with traditional computers, a microcomputer is small in size and low in cost, and is suitable for scenarios that require convenient and real-time processing, especially in portable devices. Therefore, how to utilize the advantages of a microcomputer to improve the processing ability and accuracy of speech recognition technology has become an important research direction.
[0003] Most traditional speech recognition methods are based on fixed acoustic models and language models. These methods to some extent ignore the hierarchical characteristics of signals, resulting in low recognition accuracy in complex environments. Especially in the face of relatively high background noise or poor signal quality, the performance of speech recognition systems often drops significantly. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a speech recognition method and device based on a microcomputer to solve the technical problem that the performance of a speech recognition system often drops significantly in the face of relatively high background noise or poor signal quality.
[0005] The first aspect of the embodiments of the present invention provides a speech recognition method based on a microcomputer. The speech recognition method based on a microcomputer includes:
[0006] Obtain a speech signal to be processed, and extract hierarchical time-domain features from the speech signal to be processed;
[0007] Extract hierarchical frequency-domain features from the speech signal to be processed;
[0008] Input the hierarchical time-domain features into a time-domain acoustic model to obtain a first phoneme sequence, and input the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a second phoneme sequence;
[0009] When the first phoneme sequence and the second phoneme sequence are the same, input the first phoneme sequence or the second phoneme sequence into a language model to obtain a speech recognition result output by the language model; the speech recognition result is a word sequence;
[0010] When the first phoneme sequence and the second phoneme sequence are different, process the speech signal to be processed with a multi-stage cascade filter to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of the signal; wherein, the response of each stage cascade filter is the product of time-domain impulse responses;
[0011] Identify the speech recognition result according to the spectral coefficients.
[0012] Further, the step of obtaining the speech signal to be processed and extracting the hierarchical time-domain features from the speech signal to be processed includes:
[0013] Collect the original speech signal;
[0014] Perform denoising processing, pre-emphasis, frame segmentation processing, and windowing processing on the original speech signal to obtain a plurality of the speech signals to be processed;
[0015] Calculate the short-time energy, short-time zero-crossing rate, and fundamental period of the speech signal to be processed;
[0016] Perform multi-resolution time-domain analysis on the speech signal to be processed through wavelet transform, and extract the time-domain features at different time scales;
[0017] Use the short-time energy, the short-time zero-crossing rate, the fundamental period, and the time-domain features at different time scales as the hierarchical time-domain features.
[0018] Further, the step of extracting the hierarchical frequency-domain features from the speech signal to be processed includes:
[0019] Perform short-time Fourier transform on the speech signal to be processed to obtain spectral information;
[0020] Calculate the Mel frequency cepstral coefficients of the speech signal to be processed;
[0021] Perform multi-resolution frequency-domain analysis on the spectral information through wavelet transform, and extract the frequency-domain features at different frequencies;
[0022] Use the Mel frequency cepstral coefficients and the frequency-domain features at different frequencies as the hierarchical frequency-domain features.
[0023] Further, the step of, when the first phoneme sequence and the second phoneme sequence are different, processing the speech signal to be processed by using a multi-stage cascade filter to obtain spectral coefficients includes:
[0024] Extract the different phonemes in the first phoneme sequence and the second phoneme sequence, and match the speech signals to be processed corresponding to the different phonemes; the different phonemes refer to the different phonemes existing at the same sequential position in the first phoneme sequence and the second phoneme sequence;
[0025] Input the speech signal to be processed corresponding to the different phonemes into the multi-stage cascade filter to obtain the spectral coefficients.
[0026] Further, the step of inputting the speech signal to be processed corresponding to the different phoneme into a multi-stage cascade filter to obtain the spectral coefficients includes:
[0027] Input the speech signal to be processed corresponding to the different phoneme into a multi-stage cascade filter to obtain composite responses corresponding to multiple frequency bands;
[0028] Calculate the energy values of the composite responses corresponding to multiple frequency bands respectively;
[0029] Perform logarithmic compression processing on the energy values corresponding to multiple frequency bands to obtain logarithmic energies corresponding to multiple frequency bands;
[0030] Perform discrete cosine transform on the logarithmic energies to obtain spectral coefficients.
[0031] Further, the step of identifying the speech recognition result according to the spectral coefficients includes:
[0032] Combine the spectral coefficients with the hierarchical frequency domain features and input them into a frequency domain acoustic model to obtain a third phoneme sequence output by the frequency domain acoustic model;
[0033] If the third phoneme sequence is the same as the first phoneme sequence, input the first phoneme sequence or the third phoneme sequence into a language model to obtain the speech recognition result output by the language model;
[0034] If the third phoneme sequence is the same as the second phoneme sequence, input the second phoneme sequence or the third phoneme sequence into a language model to obtain the speech recognition result output by the language model.
[0035] Further, the step of, if the third phoneme sequence is the same as the second phoneme sequence, inputting the second phoneme sequence or the third phoneme sequence into a language model to obtain the speech recognition result output by the language model includes:
[0036] If the third phoneme sequence is the same as the second phoneme sequence, input the second phoneme sequence or the third phoneme sequence into a language model to obtain an initial recognition result output by the language model;
[0037] Perform spelling correction and grammar check on the initial recognition result to obtain the speech recognition result.
[0038] The second aspect of the embodiments of the present invention provides a speech recognition device based on a microcomputer, including:
[0039] An acquisition unit, configured to acquire a speech signal to be processed and extract hierarchical time domain features in the speech signal to be processed;
[0040] An extraction unit for extracting hierarchical frequency-domain features from the speech signal to be processed;
[0041] A calculation unit for inputting the hierarchical time-domain features into a time-domain acoustic model to obtain a first phoneme sequence, and inputting the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a second phoneme sequence;
[0042] A first determination unit for, when the first phoneme sequence and the second phoneme sequence are the same, inputting the first phoneme sequence or the second phoneme sequence into a language model to obtain a speech recognition result output by the language model; the speech recognition result is a word sequence;
[0043] A second determination unit for, when the first phoneme sequence and the second phoneme sequence are different, processing the speech signal to be processed using a multi-stage cascade filter to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of the signal; wherein, the response of each stage cascade filter is the product of time-domain impulse responses;
[0044] An identification unit for identifying a speech recognition result according to the spectral coefficients.
[0045] A third aspect of an embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps in the speech recognition method based on a microcomputer described in the first aspect above when executing the computer program.
[0046] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium storing a computer program, where the computer program implements the steps in the speech recognition method based on a microcomputer described in the first aspect above when executed by a processor.
[0047] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: By extracting the hierarchical time-domain features and frequency-domain features of the speech signal to be processed, the key information of the speech can be better captured in various noise environments. The time-domain features can retain the temporal characteristics in the speech signal, while the frequency-domain features help to reveal the spectral characteristics of the speech. Through the dual extraction of these two features, this method can effectively reduce the interference of background noise on the speech signal and improve the speech recognition performance in complex environments. The time-domain acoustic model and the frequency-domain acoustic model are respectively used to process the hierarchical time-domain features and frequency-domain features, and two phoneme sequences are generated. By comparing the similarity of these two sequences, the recognition deviation caused by low signal quality can be effectively reduced. When the phoneme sequences are consistent, the input language model further optimizes the recognition result, thereby improving the accuracy and reliability of speech recognition. When the first phoneme sequence and the second phoneme sequence are inconsistent, the present invention uses a multi-stage cascade filter to further process the speech signal to be processed. This multi-stage cascade filter can process signals in different frequency ranges respectively, thereby effectively improving the spectral characteristics of the signal and enhancing the accuracy of speech recognition. The response of each stage cascade filter is optimized by the product of the time-domain impulse responses, so that the system can suppress noise in different frequency ranges in a targeted manner, and thus achieve better recognition results in complex noise environments. Through the extraction of hierarchical time-domain and frequency-domain features and the processing of the multi-stage cascade filter, this method performs particularly well in low signal-to-noise ratio environments. Traditional speech recognition systems are prone to recognition errors or performance degradation when faced with noise and signal distortion, while the present invention significantly enhances the anti-interference ability of the system and improves the speech recognition accuracy in harsh environments through a multi-level and all-round feature extraction and processing strategy. In summary, the present invention significantly improves the performance of the speech recognition system in low signal-to-noise ratio and high-noise backgrounds through the combination of hierarchical feature extraction and multi-stage cascade filter, and effectively solves the limitations of traditional speech recognition technology in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the related art descriptions will be briefly introduced below. Obviously, the drawings in the following descriptions are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 FIG. shows a schematic flowchart of a speech recognition method based on a microcomputer provided by the present invention;
[0050] Figure 2 FIG. shows a schematic diagram of a speech recognition device based on a microcomputer provided by an embodiment of the present invention;
[0051] Figure 3 FIG. 2 shows a schematic diagram of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0052] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0053] Embodiments of the present invention provide a speech recognition method and device based on a microcomputer to solve the technical problem that in the case of relatively high background noise or low signal quality, the performance of the speech recognition system often significantly degrades.
[0054] First, the present invention provides a speech recognition method based on a microcomputer. Please refer to Figure 1 , Figure 1 FIG. 3 shows a schematic flowchart of a speech recognition method based on a microcomputer provided by the present invention. As Figure 1 shown, the speech recognition method based on a microcomputer may include the following steps:
[0055] Step 101: Obtain a speech signal to be processed, and extract hierarchical time-domain features from the speech signal to be processed;
[0056] Time-domain features are directly extracted from the time-domain waveform of the audio signal. These features can reflect the time-varying process of the speech signal and are particularly important for short-time speech recognition. Especially when the signal noise is high, retaining the time-domain information helps reduce noise interference. Among them, the specific extraction logic of the hierarchical time-domain features is as follows:
[0057] Specifically, step 101 specifically includes steps 1011 to 1015:
[0058] Step 1011: Collect the original speech signal;
[0059] Through a microphone or other audio acquisition device, obtain the original speech signal to be processed. The original speech signal is usually a continuous analog audio signal, which contains the speech content of the speaker and may also include interference components such as environmental noise.
[0060] Step 1012: Perform denoising processing, pre-emphasis, framing processing, and windowing processing on the original speech signal to obtain a plurality of the speech signals to be processed;
[0061] Noise reduction processing: To remove background noise in the speech signal, various noise reduction techniques (such as spectral subtraction, Wiener filtering, etc.) can be adopted. The purpose of noise reduction processing is to improve the signal quality and make subsequent feature extraction more accurate.
[0062] Pre-emphasis: Pre-emphasis is to compensate for the spectral imbalance of the speech signal. Usually, the high-frequency part of the speech signal is relatively weak. Therefore, by performing high-pass filtering on the signal (i.e., pre-emphasis), the high-frequency components can be enhanced, making the high-frequency part of the signal smoother and enhancing the clarity of the signal.
[0063] Framing processing: Since the speech signal is a non-stationary signal (i.e., the signal changes over time), the signal needs to be divided into short-time windows (usually time frames of 20 ms to 40 ms). The signal within each frame can be regarded as a stationary signal, which is convenient for further feature analysis.
[0064] Windowing processing: Windowing processing is to reduce the "boundary effect" of the signal during the framing process (i.e., the start and end parts of the frame may not be smooth, resulting in spectral leakage). Usually, a window function (such as Hamming window, Hanning window, etc.) is applied to each frame of the signal to smooth the boundary and improve the accuracy of spectral analysis.
[0065] Step 1013: Calculate the short-time energy, short-time zero-crossing rate, and fundamental period of the speech signal to be processed;
[0066] Short-time energy is an index describing the intensity of the signal within a certain frame, and the calculation method is the sum of the squares of each sampling point of the frame signal. Short-time energy can reflect the intensity change of the signal and help distinguish speech signals from silence / noise.
[0067] Short-time zero-crossing rate refers to the number of times the signal waveform crosses zero within a certain frame. It can reflect the periodicity and frequency characteristics of the signal. Especially when processing noisy or reverberant speech signals, the short-time zero-crossing rate is a distinguishing feature between speech and noise. Signals with a higher zero-crossing rate are usually related to noise, while the zero-crossing rate of speech signals is lower and regular.
[0068] Fundamental period refers to the duration of the periodic fundamental part in the speech signal, which is usually used to identify the vocal cord vibration period of the speech. In the speech signal, the fundamental period reflects the fundamental frequency (F0) of the speech and is very important for features such as pitch and intonation.
[0069] Step 1014: Perform multi-resolution time-domain analysis on the speech signal to be processed through wavelet transform, and extract time-domain features at different time scales;
[0070] Wavelet transform is a time-frequency analysis method that can simultaneously analyze the high-frequency and low-frequency components of a signal at different time scales. Different from the traditional Fourier transform, wavelet transform has good time-frequency localization ability and can finely capture the changes of the signal in different frequency bands and time scales.
[0071] By performing multi-resolution analysis on the signal, the characteristics of the signal can be extracted from different time scales. For example, wavelet transform can provide a higher time resolution in the low-frequency part and a higher frequency resolution in the high-frequency part. In this way, different levels of speech information can be captured, which is particularly important for speech recognition in complex backgrounds.
[0072] The time-domain features extracted by wavelet transform include signal energy, instantaneous frequency, etc. at different frequency bands. These features are very useful in capturing the time changes of the signal, the changes of speech phonemes, etc. For speech signals with strong noise or poor signal quality, wavelet transform can help the system extract more detailed and multi-level features, thereby improving the recognition accuracy.
[0073] Wavelet transform provides signal information at different time scales by decomposing the signal into sub-signals with multiple frequency bandwidths. Different from the global analysis of Fourier transform, wavelet transform can analyze the changes of the signal locally and is particularly suitable for processing non-stationary signals (such as speech signals). The calculation formula of wavelet transform is as follows:
[0074]
[0075] where represents the wavelet function, represents the scale factor, represents the translation parameter, represents the speech signal to be processed at the t-th moment, represents the wavelet transform result.
[0076] In time-domain feature extraction, using the multi-resolution analysis of wavelet transform can view the changes of the signal at different time scales. For example, on a longer time scale, we can capture the long-term changes of syllables, and on a shorter time scale, we can identify the rapid changes in speech, such as instantaneous phoneme changes.
[0077] Step 1015: Use the short-time energy, the short-time zero-crossing rate, the pitch period, and the time-domain features at different time scales as the hierarchical time-domain features.
[0078] The above-extracted features (short-time energy, short-time zero-crossing rate, pitch period, time-domain features extracted by wavelet transform) together constitute the hierarchical time-domain features. These features capture the information of the speech signal from aspects such as the energy change, frequency characteristics, periodicity, and time scale of the signal. By combining these features from different sources, the system can describe the speech signal in multiple dimensions, enhancing the robustness against noise interference and providing richer and more diverse information for subsequent speech recognition. This hierarchical feature extraction method helps to improve the system performance in different speech recognition tasks.
[0079] In the embodiments of steps 1011 to 1015, the purpose is to construct a multi-level and multi-dimensional information representation through a series of preprocessing and feature extraction methods, so as to effectively process speech signals in complex environments. Through operations such as denoising, pre-emphasis, framing, and windowing, the system can improve the signal quality; through the extraction of features such as short-time energy, zero-crossing rate, and pitch period, as well as the multi-resolution analysis of wavelet transform, the system can extract the time-domain features of speech from multiple perspectives, thus providing a solid foundation for subsequent speech recognition. Hierarchical time-domain feature extraction analyzes the speech signal from multiple levels and different time scales by combining techniques such as short-time energy, zero-crossing rate, autocorrelation analysis, and wavelet transform. This multi-level and all-round feature extraction strategy can fully capture the diverse information in the speech signal and has strong robustness, especially in speech signals with complex background noise or different speakers.
[0080] Step 102: Extract the hierarchical frequency-domain features from the speech signal to be processed;
[0081] Different from time-domain features, frequency-domain features are obtained by performing a Fourier transform (such as short-time Fourier transform) on the audio signal. Frequency-domain features can reflect the energy distribution of different frequency components and are usually particularly effective for identifying features such as pitch and timbre of sound. Hierarchical frequency-domain features mean extracting features in different frequency ranges, which can capture the changes in the speech signal more precisely. The specific extraction logic of the hierarchical frequency-domain features is as follows:
[0082] Specifically, step 102 specifically includes steps 1021 to 1024:
[0083] Step 1021: Perform a short-time Fourier transform on the speech signal to be processed to obtain spectral information;
[0084] The short-time Fourier transform is a commonly used method for converting a signal from the time domain to the frequency domain. The short-time Fourier transform processes the signal by framing, and then performs a Fourier transform on each frame to obtain the representation of each frame in the frequency domain. The short-time Fourier transform can show the changes of the signal in both the time and frequency dimensions. Through the short-time Fourier transform, the obtained spectral information describes the frequency distribution of the speech signal in each frame. The spectral information of each frame contains the various frequency components and their corresponding energy distributions within that time period, and can effectively reflect the frequency characteristics of the speech signal. The spectral information is the basis for subsequent feature extraction (such as Mel-frequency cepstral coefficients) and speech recognition. By analyzing the spectrum, the frequency changes of the speech signal can be deeply understood, which helps to distinguish different speech phonemes or background noises.
[0085] Step 1022: Calculate the Mel-frequency cepstral coefficients of the speech signal to be processed;
[0086] The Mel-frequency cepstral coefficients are generated through the following steps:
[0087] First, use the STFT to calculate the spectrum of the speech signal.
[0088] Use a set of Mel filters to filter the spectrum. The frequency response of the Mel filters simulates the perceptual differences of the human ear to different frequencies. The Mel filter bank divides the spectrum into multiple sub-bands, with the bandwidth of each sub-band being smaller in the high-frequency part and larger in the low-frequency part.
[0089] Take the logarithm of the energy value output by each filter to simulate the non-linear response of the human ear to the sound intensity.
[0090] Finally, compress the log energy spectrum through the discrete cosine transform to obtain the MFCC coefficients. Usually, the first 12 coefficients are used to represent the features of the speech signal.
[0091] The Mel-frequency cepstral coefficients effectively compress the frequency information of the speech signal by simulating the auditory mechanism of the human ear, while retaining the features useful for recognition in the speech.
[0092] Step 1023: Perform multi-resolution frequency domain analysis on the spectral information through wavelet transform to extract the frequency domain features at different frequencies;
[0093] Wavelet transform is a frequency domain analysis tool capable of multi - resolution analysis. Different from the Fourier transform, wavelet transform has the ability of localization in both time and frequency, and can capture the high - frequency and low - frequency information of a signal simultaneously at different frequency scales. Through wavelet transform, a signal can be analyzed in different frequency ranges. In the low - frequency part, wavelet transform provides a higher time resolution, while in the high - frequency part, it provides a higher frequency resolution. This enables wavelet transform to capture both the stable characteristics on a long - time scale and the rapid change characteristics on a short - time scale in a signal. By performing wavelet transform on the spectral information, the system can extract the frequency domain characteristics in different frequency ranges. These characteristics can help the system better identify the detailed changes in the speech signal. Especially in a noisy environment, wavelet transform can provide more refined frequency domain characteristics.
[0094] Wavelet transform is a traditional technology and can be implemented with reference to step 1014, which will not be elaborated here.
[0095] Step 1024: Take the Mel - Frequency Cepstral Coefficients and the frequency domain characteristics at different frequencies as the hierarchical frequency domain characteristics.
[0096] Combine the Mel - Frequency Cepstral Coefficients (MFCC) and the frequency domain characteristics in different frequency ranges extracted by wavelet transform to form hierarchical frequency domain characteristics. This kind of hierarchical frequency domain characteristics combines the global frequency characteristics of speech by MFCC and the analysis results of local frequency details by wavelet transform.
[0097] By combining these two kinds of characteristics, the system can simultaneously process and understand the details of the speech signal at multiple frequency levels. MFCC provides an overall description of the speech signal spectrum, while wavelet transform can capture the more refined frequency changes in the signal. Especially in a complex environment, this kind of hierarchical frequency domain characteristics can improve the robustness and accuracy of the speech recognition system.
[0098] In the embodiments of steps 1021 to 1024, the hierarchical frequency domain characteristics of the speech signal are extracted with the aim of obtaining the frequency information of speech at multiple frequency scales. The short - time Fourier transform provides the basic spectral information. The Mel - Frequency Cepstral Coefficients (MFCC) simulate the auditory characteristics of the human ear and can effectively extract the frequency domain characteristics of speech. Wavelet transform provides a more refined frequency domain analysis and can capture the local and global frequency characteristics in the signal based on multi - resolution. By combining these frequency domain characteristics from different sources, the system can analyze and understand the speech signal more comprehensively and accurately, especially improving the performance of speech recognition in a noisy environment.
[0099] Step 103: Input the hierarchical time - domain characteristics into the time - domain acoustic model to obtain the first phoneme sequence, and input the hierarchical frequency domain characteristics into the frequency - domain acoustic model to obtain the second phoneme sequence;
[0100] This step is to input the features extracted from the time domain into an acoustic model to generate a phoneme sequence. The phoneme sequence represents the order of the basic speech units (phonemes) in speech and is an intermediate output in speech recognition.
[0101] Input the features extracted from the frequency domain into another acoustic model (which may also be a deep neural network) to generate a phoneme sequence. Since the frequency domain features are sensitive to the frequency components of speech, this model can better process the speech information related to frequency.
[0102] Among them, the time domain acoustic model and the frequency domain acoustic model are traditional models, and the specific processing logics are not elaborated here.
[0103] Step 104: When the first phoneme sequence is the same as the second phoneme sequence, input the first phoneme sequence or the second phoneme sequence into the language model to obtain the speech recognition result output by the language model; the speech recognition result is a word sequence;
[0104] The system will compare the phoneme sequences generated in the time domain and the frequency domain. If the two sequences are the same, it indicates that the features extracted from these two different perspectives have obtained the same result in speech recognition, further proving the reliability of the recognition process. If the first and second phoneme sequences are consistent, then input any one of the phoneme sequences into the language model for subsequent language understanding and reasoning. The language model can output the final speech recognition result (word sequence) through context information. The language model adopts a traditional model, and its processing process is not elaborated here.
[0105] Step 105: When the first phoneme sequence is different from the second phoneme sequence, process the speech signal to be processed using a multi-stage cascade filter to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of the signal; where the response of each stage cascade filter is the product of the time domain impulse responses;
[0106] If the phoneme sequences extracted from the time domain and the frequency domain are inconsistent, it may mean that there is strong noise in the signal or some features that are difficult to identify by traditional methods. Therefore, the system chooses to introduce a multi-stage cascade filter to further process the speech signal. This filter is usually used to separate signals in different frequency ranges so as to perform more refined processing on signals in different frequency bands. The response of each stage filter is the product of the time domain impulse responses, which helps to filter specific frequency bands in the signal, reduce noise and enhance effective speech information. Each stage cascade filter focuses on processing different frequency segments in the speech signal. By cascading multiple filters, high-frequency noise or low-frequency interference can be gradually filtered out, and the frequency components crucial for speech recognition can be retained.
[0107] The product of the time-domain impulse responses means that the impulse response function of the filter combines the responses of different frequency bands through cascading operations to form the final frequency response. Such a design can effectively process complex speech signals, especially in environments with high noise levels.
[0108] Specifically, step 105 specifically includes steps 1051 to 1052:
[0109] Step 1051: Extract the different phonemes in the first phoneme sequence and the second phoneme sequence, and match the speech signal to be processed corresponding to the different phonemes; the different phonemes refer to the different phonemes existing at the same sequential position in the first phoneme sequence and the second phoneme sequence;
[0110] The phoneme sequence is the basic unit of speech. By comparing the phonemes in the "first phoneme sequence" and the "second phoneme sequence", it can be found that there are differences between the two at certain positions.
[0111] Step 1052: Input the speech signal to be processed corresponding to the different phonemes into a multi-stage cascaded filter to obtain the spectral coefficients.
[0112] By cascading multiple filters, the signal can be processed step by step, thereby extracting different frequency components of the signal. The cascaded filter can perform fine analysis on the signal in different frequency bands, enabling the features in different frequency ranges to be enhanced or suppressed more effectively.
[0113] Through the processing of the multi-stage cascaded filter, the system will obtain the spectral coefficients corresponding to these different phonemes. The spectral coefficients reflect the energy distribution of the signal in each frequency band. The spectral coefficients play a core role in the processing and feature extraction of speech signals. Especially when comparing different phonemes or identifying subtle differences, the spectral coefficients are very important basic features.
[0114] In the embodiments of steps 1051 to 1052, for different phonemes (different phonemes), the corresponding speech signal segments are extracted, and the frequency characteristics of these signal segments are further analyzed through a multi-stage cascaded filter, thereby obtaining the spectral coefficients. This process helps the system focus on processing the parts that are different at the phoneme level, which is helpful to improve the recognition accuracy of speech signals. Especially when processing phonemes with subtle differences or speech in a noisy environment, through the efficient frequency-domain analysis of the multi-stage cascaded filter, the important features in the speech can be finely captured, enhancing the robustness of the system.
[0115] Specifically, step 1052 specifically includes steps A1 to A4:
[0116] Step A1: Input the speech signal to be processed corresponding to the different phonemes into a multi-stage cascaded filter to obtain the composite responses corresponding to multiple frequency bands;
[0117] As mentioned previously, the differential phonemes are obtained by comparing different phonemes in the phoneme sequence. The speech signal segments corresponding to these differential phonemes will then be input into a multi-stage cascade filter.
[0118] A cascade filter is usually composed of multiple filters connected in cascade. Each filter processes signals in different frequency bands, can decompose the speech signal, and extract details in different frequency ranges. Through the processing of the multi-stage cascade filter, "composite responses" corresponding to multiple frequency bands can be obtained, that is, the signals after being processed by the filter in each frequency band.
[0119] The composite response usually refers to the output results of the signal in each frequency band after passing through the filter, reflecting the frequency components of the signal in each frequency band. The composite responses of multiple frequency bands together constitute a complete description of the frequency-domain characteristics of the speech signal.
[0120] The calculation process of the multi-stage cascade filter is as follows:
[0121]
[0122] where is a function of time and the center frequency , represents the polynomial factor in the time domain part, represents the i-th bandwidth.
[0123] represents the exponential decay term of the filter, which controls the bandwidth of the filter. This term determines the attenuation rate of the filter for the signal, and thus affects the bandwidth (frequency selectivity) of the filter. A larger corresponds to a wider bandwidth, and a smaller corresponds to a narrower bandwidth. Here, is the bandwidth of the -th filter, which will be adjusted according to different frequency ranges. For example, the filters in the low-frequency region may have a smaller bandwidth, while the filters in the high-frequency region may have a larger bandwidth.
[0124] is the cosine modulation term of the filter, which controls the center frequency of the filter. This term represents the selectivity of the filter for a specific frequency, is the center frequency of the -th filter, which determines the "position" of the frequency response of the filter. Due to the periodicity of the cosine function, the filter response shows a center frequency The oscillatory behavior is such that the response of the filter is strongest near the center frequency and gradually weakens as the frequency moves away.
[0125] Denotes the bandwidth parameter of the
[0126] th filter, which determines the frequency selectivity of the filter. The larger the bandwidth, the wider the response region of the filter, and vice versa. In practical applications, the appropriate bandwidth can be selected according to the frequency range processed by the filter. Usually, the bandwidth in the low-frequency region may be smaller, while the bandwidth in the high-frequency region may be larger to adapt to the frequency selectivity of the cochlea. Denotes the center frequency of the th filter. Each filter has a center frequency, which determines the frequency at which the filter has the strongest response. The selection of the center frequency is usually adjusted according to the spectral characteristics of the signal. In this cascaded filter structure, the center frequencies of each filter
[0127] Denotes the index of the filter, representing the position of the current filter in the cascading process. For example, is the first filter, is the second filter, and so on. The change of
[0128] determines the different parameters of the filter (such as order, bandwidth, center frequency) and controls the range and manner in which each filter acts on the signal.
[0129] This formula describes the cascaded structure of multiple filters. The time-domain response of each filter consists of a polynomial factor, an exponential decay factor, and a cosine modulation factor. By adjusting the bandwidth and the center frequency of each filter, detailed signal processing can be carried out for different frequency ranges to simulate the frequency response characteristics of the human ear in different frequency bands. By cascading multiple filters, we can provide more complex frequency selectivity, thereby better analyzing and processing signals.
[0130] A cascaded filter means that the outputs of multiple filters are passed step by step to the next filter. Each filter processes different parts of the signal (different frequency ranges), thus generating a composite response in the final output. Here, each cascaded filter analyzes a specific frequency band or characteristic, and the final output is the combined response of all these frequency bands. By cascading multiple filters, the processing ability of the human ear for complex frequency distributions can be simulated, especially when the filtering responses in different frequency bands have different characteristics.
[0131] The response of each individual filter is the product of the time-domain impulse responses and usually behaves as a band-pass filter, whose bandwidth and center frequency determine its selectivity for different frequency bands. By cascading multiple filters, the response of the filters becomes more refined and can better simulate the complex auditory process of the human ear.
[0132] The gain function of each filter (determined by an exponential decay term and a cosine term) controls the frequency response of the signal in the frequency band of that filter. Since the responses of multiple filters are different (e.g., different bandwidths and center frequencies ), they can perform a detailed analysis of the signal in different frequency ranges, similar to the responses of different parts of the cochlea (different basilar membrane regions) to different frequencies.
[0133] Cascaded filters can provide more complex frequency selectivity: filters in the low-frequency region may have a higher frequency resolution, while the responses of filters in the high-frequency region may be more relaxed. By applying different bandwidths and center frequencies in different frequency bands, the different perceptual characteristics of the human ear for different frequency intervals can be simulated.
[0134] This method can enhance the frequency resolution of the filter in different frequency segments, perform a more detailed analysis of the signal in the low-frequency region, and may process the smooth response of the signal through a wide-bandwidth filter in the high-frequency region.
[0135] The cochlea does not simply perform a single-level frequency analysis on the input signal. In fact, different positions on the basilar membrane of the cochlea have different induction characteristics for different frequencies, and each position acts like a frequency-selective filter. By cascading multiple filters, the processing of signals in different frequency bands at different positions of the cochlea can be simulated.
[0136] Cascaded filters simulate this multi-level filtering process, where each filter acts on a different frequency range and forms a complex frequency response in the final output.
[0137] By using cascaded filters, we can better simulate the frequency selectivity and multi-stage processing mechanism of the cochlea. Multiple filters provide more delicate frequency analysis and enable higher-precision feature extraction within a wider frequency band. Additionally, when processing signals in different frequency ranges, cascaded filters can utilize different bandwidth and center frequency parameters, enhancing the robustness of the entire system in a noisy environment.
[0138] Cascaded filters can not only refine the frequency selectivity of signals but also simulate the cochlea's ability to perform multi-stage frequency processing in different frequency bands. The advantage of this method lies in the collaborative work of multiple filters, which can provide a more refined and flexible frequency response, thereby improving performance and robustness in areas such as speech recognition and audio processing.
[0139] Step A2: Calculate the energy values of the composite responses corresponding to multiple frequency bands respectively;
[0140] For the composite response of each frequency band, its energy value needs to be calculated. Energy is usually represented by the sum of the squares of the signal or calculated by the square of its spectral amplitude. The energy value can reflect the signal strength or importance in that frequency band.
[0141] The purpose of calculating the energy value is to capture the energy distribution of the signal in each frequency band. Frequency bands with high energy usually contain more useful information, while those with low energy may contain noise or unimportant components.
[0142] Step A3: Perform logarithmic compression on the energy values corresponding to multiple frequency bands to obtain the logarithmic energies corresponding to multiple frequency bands;
[0143] Logarithmic compression is mainly used to simulate the way the human ear perceives sound. The human ear's perception of sound has a logarithmic characteristic, that is, it is more sensitive to large-amplitude changes and less responsive to small-amplitude changes. By performing logarithmic compression on the energy values, the energy differences between different frequency bands can be made more in line with the human ear's perception law.
[0144] By performing logarithmic compression on the energy, it is possible to better balance the energy differences between frequency bands and transform the energy values into a form more suitable for subsequent analysis. This helps to compress information and reduce unnecessary fluctuations, making the features more stable, especially in the low-energy part, and avoiding excessive amplification of weak signals.
[0145] Step A4: Perform discrete cosine transform on the logarithmic energy to obtain spectral coefficients.
[0146] The discrete cosine transform can effectively extract the principal components of a signal and remove redundant information. After being processed by the discrete cosine transform, the obtained spectral coefficients are the low-dimensional representation of the speech signal, usually including the principal components of the logarithmic energy in each frequency band. The spectral coefficients are important features in the speech signal and can effectively describe the spectral structure of the speech.
[0147] The discrete cosine transform processing can extract the most representative part of the signal, remove most of the redundant information, and at the same time retain the features related to speech recognition. The obtained spectral coefficients not only reduce the amount of data, but also retain the key frequency-domain features in the speech signal.
[0148] In the embodiments of steps A1 to A4, the input speech signal is subjected to multi-band filtering processing to obtain the "composite response" in multiple frequency bands. Calculate the energy value of the composite response in each frequency band to capture the intensity distribution of the signal in different frequency bands. Perform logarithmic compression on the energy value to simulate the perception characteristics of the human ear and make the distribution of the energy more in line with the auditory perception law. Convert the logarithmic energy into spectral coefficients through DCT, and finally obtain a concise and effective feature representation. These steps work together to extract the spectral coefficients, which can be used as the input for subsequent speech recognition or analysis tasks and can effectively describe the frequency-domain characteristics of the speech signal. Through these operations, the key information of the speech signal is retained, and the redundant part is compressed, thereby improving the efficiency and accuracy of system processing and recognition.
[0149] Step 106: Identify the speech recognition result according to the spectral coefficients.
[0150] The spectral coefficients after being processed by the multi-stage cascade filter reflect the intensity and distribution of the signal at different frequencies. These spectral coefficients will be further used as the input of the speech recognition model for final phoneme recognition and output the corresponding speech recognition result.
[0151] Specifically, step 106 specifically includes steps 1061 to 1063:
[0152] Step 1061: Combine the spectral coefficients with the hierarchical frequency-domain features and input them into the frequency-domain acoustic model to obtain the third phoneme sequence output by the frequency-domain acoustic model;
[0153] The spectral coefficients are obtained through multi-stage cascaded filters, energy calculation, logarithmic compression, and DCT transformation in the previous steps. The hierarchical frequency-domain features refer to the feature information in multiple frequency bands and different frequency ranges, which involves different expressions of high-frequency, low-frequency, or mid-frequency components. Combining the spectral coefficients with the hierarchical frequency-domain features is actually a more detailed characterization of the frequency-domain information of the speech signal. The combined frequency-domain features are input into the frequency-domain acoustic model, and the model analyzes the signal and outputs a phoneme sequence, which is called the third phoneme sequence.
[0154] Step 1062: If the third phoneme sequence is the same as the first phoneme sequence, input the first phoneme sequence or the third phoneme sequence into the language model to obtain the speech recognition result output by the language model.
[0155] According to the goal of speech recognition, it is necessary to first determine whether the third phoneme sequence output by the frequency-domain acoustic model matches the first phoneme sequence. If the third phoneme sequence is exactly the same as the first phoneme sequence, it means that the phoneme sequence extracted from the speech signal is consistent with the original phoneme sequence, indicating that the model processes the signal more accurately. After obtaining the phoneme sequence, the system needs to generate corresponding words or sentences based on these phoneme sequences. At this time, the system inputs the first phoneme sequence or the third phoneme sequence into the language model. By considering factors such as grammar and context, the language model can further optimize the speech recognition result and produce a reasonable speech recognition output. The output of the language model is the final speech recognition result, which not only depends on the phoneme sequence recognized by the acoustic model but also needs to consider the grammatical structure and context meaning of the sentence, so that the final recognition result is more natural and accurate.
[0156] Step 1063: If the third phoneme sequence is the same as the second phoneme sequence, input the second phoneme sequence or the third phoneme sequence into the language model to obtain the speech recognition result output by the language model.
[0157] If the third phoneme sequence matches the second phoneme sequence, it means that they are relatively consistent in speech features, which may reflect different speech signals or language variations. At this time, the system needs to perform recognition based on this new phoneme sequence. Similar to the previous step, input the second phoneme sequence or the third phoneme sequence into the language model for processing. The language model judges the most likely word or sentence based on the input phoneme sequence and outputs the final speech recognition result. The recognition result obtained at this time may be different from the previous result because the input phoneme sequences are different. The language model will generate the final recognition result by combining the second phoneme sequence or the third phoneme sequence according to factors such as context and grammar rules.
[0158] In the embodiments of steps 1061 to 1063, first, a phoneme recognition is performed on the speech signal through a frequency-domain acoustic model to obtain a third phoneme sequence. If the third phoneme sequence is the same as the first phoneme sequence, then this sequence or the original phoneme sequence is input into the language model to obtain the final speech recognition result. If the third phoneme sequence is the same as the second phoneme sequence, then the second phoneme sequence or the third phoneme sequence is input into the language model to generate the final speech recognition result. The design of these steps enables the system to flexibly adjust according to different situations during the phoneme recognition process, combines the advantages of the acoustic model and the language model, and improves the accuracy and robustness of speech recognition.
[0159] Specifically, step 1063 specifically includes steps B1 to B2:
[0160] Step B1: If the third phoneme sequence is the same as the second phoneme sequence, then the second phoneme sequence or the third phoneme sequence is input into the language model to obtain the initial recognition result output by the language model;
[0161] After being processed by the language model, the output obtained by the system is an initial recognition result. This result represents the most likely word or sentence predicted based on the current phoneme sequence and the language model.
[0162] Step B2: Perform spelling correction and grammar checking on the initial recognition result to obtain the speech recognition result.
[0163] Spelling correction: The initial recognition result is the recognition output obtained based on the language model. However, during the speech recognition process, due to reasons such as noise, accent, unclear pronunciation, etc., the recognition result may contain spelling mistakes. The spelling correction stage will check and correct the words in the initial recognition result.
[0164] Grammar checking: In addition to spelling mistakes, the initial result of speech recognition may also have grammar errors. Although the language model will consider some basic context information, in more complex sentence structures, inappropriate grammar or expressions may still occur. Grammar checking will review the syntactic structure of the recognition result to ensure that the sentence conforms to grammar rules. For example, it may correct problems such as subject-verb disagreement, misused tenses, or word order. Grammar checking generally uses rule-based grammar analysis or statistical models to perform.
[0165] After spelling correction and grammar checking, the recognition result will be more accurate and natural. The final output is the speech recognition result obtained after two-step correction, and at this time, the accuracy and fluency of speech recognition are significantly improved.
[0166] In the embodiments from step B1 to step B2, through spelling correction and grammar checking, the system can better handle common errors in speech recognition (such as misspellings, inappropriate word choices, unsmooth grammar, etc.), thereby improving the naturalness, accuracy, and readability of the recognition results. This multi-level processing method not only depends on the phoneme sequence output of the acoustic model but also combines the semantic inference of the language model and the correction in the post-processing stage, making the final result closer to human language understanding.
[0167] In the embodiments corresponding to steps 101 to 106, by extracting the hierarchical time-domain features and frequency-domain features of the speech signal to be processed, the key information of the speech can be better captured in various noise environments. The time-domain features can retain the temporal characteristics in the speech signal, while the frequency-domain features help to reveal the spectral characteristics of the speech. Through the dual extraction of these two types of features, this method can effectively reduce the interference of background noise on the speech signal and improve the speech recognition performance in complex environments. The time-domain acoustic model and the frequency-domain acoustic model are respectively used to process the hierarchical time-domain features and frequency-domain features, and two phoneme sequences are generated. By comparing the similarity of these two sequences, the recognition deviation caused by low signal quality can be effectively reduced. When the phoneme sequences are consistent, the input language model further optimizes the recognition result, thereby improving the accuracy and reliability of speech recognition. When the first phoneme sequence and the second phoneme sequence are inconsistent, the present invention uses a multi-stage cascade filter to further process the speech signal to be processed. This multi-stage cascade filter can process signals in different frequency ranges respectively, thereby effectively improving the spectral characteristics of the signal and enhancing the accuracy of speech recognition. The response of each stage cascade filter is optimized through the product of the time-domain impulse responses, enabling the system to specifically suppress noise in different frequency ranges, thus achieving better recognition results in complex noise environments. Through the extraction of hierarchical time-domain and frequency-domain features and the processing of the multi-stage cascade filter, this method performs particularly well in low signal-to-noise ratio environments. Traditional speech recognition systems are prone to recognition errors or performance degradation when facing noise and signal distortion, while the present invention significantly enhances the anti-interference ability of the system and improves the speech recognition accuracy in harsh environments through a multi-level and all-round feature extraction and processing strategy. In summary, the present invention significantly improves the performance of the speech recognition system in low signal-to-noise ratio and high-noise backgrounds by combining hierarchical feature extraction and a multi-stage cascade filter, effectively solving the limitations of traditional speech recognition technologies in complex environments.
[0168] Such as Figure 2 The present invention provides a speech recognition device based on a microcomputer. Please refer to Figure 2 , Figure 2 shows a schematic diagram of a speech recognition device based on a microcomputer provided by the present invention. Such as Figure 2A voice recognition device based on a microcomputer includes:
[0169] An acquisition unit 21, configured to acquire a voice signal to be processed and extract hierarchical time-domain features from the voice signal to be processed;
[0170] An extraction unit 22, configured to extract hierarchical frequency-domain features from the voice signal to be processed;
[0171] A calculation unit 23, configured to input the hierarchical time-domain features into a time-domain acoustic model to obtain a first phoneme sequence, and input the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a second phoneme sequence;
[0172] A first judgment unit 24, configured to, when the first phoneme sequence and the second phoneme sequence are the same, input the first phoneme sequence or the second phoneme sequence into a language model to obtain a voice recognition result output by the language model; the voice recognition result is a word sequence;
[0173] A second judgment unit 25, configured to, when the first phoneme sequence and the second phoneme sequence are different, process the voice signal to be processed using a multi-stage cascade filter to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of the signal; wherein, the response of each stage cascade filter is the product of time-domain impulse responses;
[0174] An identification unit 26, configured to identify a voice recognition result according to the spectral coefficients.
[0175] A voice recognition device based on a microcomputer provided by the present invention can better capture the key information of speech in various noise environments by extracting the hierarchical time-domain features and frequency-domain features of the speech signal to be processed. The time-domain features can retain the timing characteristics in the speech signal, while the frequency-domain features help to reveal the spectral characteristics of the speech. Through the dual extraction of these two features, this method can effectively reduce the interference of background noise on the speech signal and improve the speech recognition performance in complex environments. The time-domain acoustic model and the frequency-domain acoustic model are respectively used to process the hierarchical time-domain features and frequency-domain features, and two phoneme sequences are generated. By comparing the similarity of these two sequences, the recognition deviation caused by low signal quality can be effectively reduced. When the phoneme sequences are consistent, the input language model further optimizes the recognition result, thereby improving the accuracy and reliability of speech recognition. When the first phoneme sequence and the second phoneme sequence are inconsistent, the present invention uses a multi-stage cascade filter to further process the speech signal to be processed. The multi-stage cascade filter can process signals in different frequency ranges respectively, thereby effectively improving the spectral characteristics of the signal and enhancing the accuracy of speech recognition. The response of each stage cascade filter is optimized by the product of the time-domain impulse responses, so that the system can specifically suppress noise in different frequency ranges, and thus achieve better recognition results in complex noise environments. Through the extraction of hierarchical time-domain and frequency-domain features and the processing of the multi-stage cascade filter, this method performs particularly well in low signal-to-noise ratio environments. Traditional speech recognition systems are prone to recognition errors or performance degradation when facing noise and signal distortion, while the present invention significantly enhances the anti-interference ability of the system and improves the speech recognition accuracy in harsh environments through a multi-level and all-round feature extraction and processing strategy. In summary, the present invention significantly improves the performance of the speech recognition system in low signal-to-noise ratio and high-noise backgrounds through the combination of hierarchical feature extraction and multi-stage cascade filter, and effectively solves the limitations of traditional speech recognition technology in complex environments.
[0176] Figure 3 It is a schematic diagram of a terminal device provided by an embodiment of the present invention. As Figure 3 shown, a terminal device 3 in this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a voice recognition program based on a microcomputer. When the processor 30 executes the computer program 32, the steps in the above-mentioned embodiments of various voice recognition methods based on a microcomputer are implemented, such as Figure 1 the steps 101 to 106 shown. Alternatively, when the processor 30 executes the computer program 32, the functions of each unit in the above-mentioned device embodiments are implemented, such as Figure 2 the functions of the units shown.
[0177] Exemplarily, the computer program 32 may be divided into one or more units, which are stored in the memory 31 and executed by the processor 30 to implement the present invention. The one or more units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the terminal device 3. For example, the specific functions of the computer program 32 that may be divided into each unit are as follows:
[0178] An acquisition unit, configured to acquire a voice signal to be processed and extract hierarchical time-domain features in the voice signal to be processed;
[0179] An extraction unit, configured to extract hierarchical frequency-domain features in the voice signal to be processed;
[0180] A calculation unit, configured to input the hierarchical time-domain features into a time-domain acoustic model to obtain a first phoneme sequence, and input the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a second phoneme sequence;
[0181] A first determination unit, configured to, when the first phoneme sequence is the same as the second phoneme sequence, input the first phoneme sequence or the second phoneme sequence into a language model to obtain a speech recognition result output by the language model; the speech recognition result is a word sequence;
[0182] A second determination unit, configured to, when the first phoneme sequence is different from the second phoneme sequence, process the voice signal to be processed by using a multi-stage cascade filter to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of the signal; wherein, the response of each stage cascade filter is the product of time-domain impulse responses;
[0183] An identification unit, configured to identify a speech recognition result according to the spectral coefficients.
[0184] The terminal device includes, but is not limited to, the processor 30 and the memory 31. Those skilled in the art can understand that Figure 3 This is only an example of the terminal device 3, and does not constitute a limitation on the terminal device 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal device may further include input / output devices, network access devices, buses, etc.
[0185] The processor 30 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0186] The memory 31 may be an internal storage unit of the terminal device 3, such as a hard disk or memory of the terminal device 3. The memory 31 may also be an external storage device of the terminal device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the terminal device 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the terminal device 3. The memory 31 is used to store the computer program and other programs and data required by the roaming control device. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0187] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0188] It should be noted that the content such as information interaction and execution process between the above devices / units, due to being based on the same concept as the method embodiments of the present invention, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.
[0189] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0190] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor can implement the steps in the above method embodiments.
[0191] An embodiment of the present invention provides a computer program product, which when running on a mobile terminal enables the mobile terminal to execute the steps in the above method embodiments.
[0192] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.
[0193] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0194] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0195] In the embodiments provided by the present invention, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0196] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units.
[0197] It should be understood that when used in the specification and appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0198] It should also be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.
[0199] As used in the specification and appended claims of the present invention, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0200] In addition, in the description of the specification and the appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0201] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present invention means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of the present invention. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other some embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0202] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A voice recognition method based on a microcomputer, characterized in that, The microcomputer-based speech recognition method includes: Obtain a speech signal to be processed, and extract hierarchical time-domain features in the speech signal to be processed; Extract hierarchical frequency-domain features in the speech signal to be processed; Input the hierarchical time-domain features into a time-domain acoustic model to obtain a first phoneme sequence, and input the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a second phoneme sequence; When the first phoneme sequence and the second phoneme sequence are the same, input the first phoneme sequence or the second phoneme sequence into a language model to obtain a speech recognition result output by the language model; the speech recognition result is a word sequence; When the first phoneme sequence and the second phoneme sequence are different, extract the different phonemes in the first phoneme sequence and the second phoneme sequence, and match the speech signal to be processed corresponding to the different phonemes; the different phonemes refer to the different phonemes at the same sequential position in the first phoneme sequence and the second phoneme sequence; Input the speech signal to be processed corresponding to the different phonemes into a multi-stage cascade filter to obtain composite responses corresponding to multiple frequency bands; Calculate the energy values of the composite responses corresponding to multiple frequency bands respectively; Perform logarithmic compression processing on the energy values corresponding to multiple frequency bands to obtain logarithmic energies corresponding to multiple frequency bands; Perform discrete cosine transform on the logarithmic energies to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of signals; wherein, the response of each stage cascade filter is the product of time-domain impulse responses; Identify the speech recognition result according to the spectral coefficients.
2. The speech recognition method based on a microcomputer according to claim 1, characterized in that The steps of obtaining the speech signal to be processed and extracting hierarchical time-domain features in the speech signal to be processed include: Collect an original speech signal; Perform denoising processing, pre-emphasis, frame segmentation processing, and windowing processing on the original speech signal to obtain multiple speech signals to be processed; Calculate the short-time energy, short-time zero-crossing rate, and fundamental pitch period of the speech signal to be processed; Perform multi-resolution time-domain analysis on the speech signal to be processed through wavelet transform to extract time-domain features at different time scales; Use the short-time energy, the short-time zero-crossing rate, the fundamental pitch period, and the time-domain features at different time scales as the hierarchical time-domain features.
3. The speech recognition method based on a microcomputer according to claim 1, wherein The steps of extracting hierarchical frequency-domain features in the speech signal to be processed include: Perform short-time Fourier transform on the speech signal to be processed to obtain spectral information; Calculate the Mel frequency cepstral coefficients of the speech signal to be processed; Perform multi-resolution frequency-domain analysis on the spectral information through wavelet transform to extract frequency-domain features at different frequencies; Use the Mel frequency cepstral coefficients and the frequency-domain features at different frequencies as the hierarchical frequency-domain features.
4. The speech recognition method based on a microcomputer according to claim 1, characterized in that The steps of identifying the speech recognition result according to the spectral coefficients include: Input the spectral coefficients combined with the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a third phoneme sequence output by the frequency-domain acoustic model; If the third phoneme sequence is the same as the first phoneme sequence, input the first phoneme sequence or the third phoneme sequence into a language model to obtain a speech recognition result output by the language model; If the third phoneme sequence is the same as the second phoneme sequence, input the second phoneme sequence or the third phoneme sequence into the language model to obtain the speech recognition result output by the language model.
5. The speech recognition method based on a microcomputer according to claim 4, characterized in that, The step of, if the third phoneme sequence is the same as the second phoneme sequence, inputting the second phoneme sequence or the third phoneme sequence into the language model to obtain the speech recognition result output by the language model includes: If the third phoneme sequence is the same as the second phoneme sequence, input the second phoneme sequence or the third phoneme sequence into the language model to obtain the initial recognition result output by the language model; Perform spelling correction and grammar checking on the initial recognition result to obtain the speech recognition result.
6. A voice recognition device based on a microcomputer, characterized in that, The speech recognition device based on a microcomputer includes: An acquisition unit configured to acquire a speech signal to be processed and extract hierarchical time-domain features from the speech signal to be processed; An extraction unit configured to extract hierarchical frequency-domain features from the speech signal to be processed; A calculation unit configured to input the hierarchical time-domain features into a time-domain acoustic model to obtain a first phoneme sequence, and input the hierarchical frequency-domain features into a frequency-domain acoustic model to obtain a second phoneme sequence; A first determination unit configured to, when the first phoneme sequence is the same as the second phoneme sequence, input the first phoneme sequence or the second phoneme sequence into the language model to obtain the speech recognition result output by the language model; the speech recognition result is a word sequence; A second determination unit configured to, when the first phoneme sequence is different from the second phoneme sequence, extract the different phonemes in the first phoneme sequence and the second phoneme sequence, and match the speech signal to be processed corresponding to the different phonemes; the different phonemes refer to the different phonemes at the same sequential position in the first phoneme sequence and the second phoneme sequence; input the speech signal to be processed corresponding to the different phonemes into a multi-stage cascade filter to obtain composite responses corresponding to multiple frequency bands; respectively calculate the energy values of the composite responses corresponding to the multiple frequency bands; perform logarithmic compression processing on the energy values corresponding to the multiple frequency bands to obtain the logarithmic energies corresponding to the multiple frequency bands; perform discrete cosine transform on the logarithmic energies to obtain spectral coefficients; the multi-stage cascade filter is used to process different frequency ranges of the signal; wherein, the response of each stage cascade filter is the product of time-domain impulse responses; An identification unit configured to identify the speech recognition result according to the spectral coefficients.
7. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and a speech recognition program based on a microcomputer stored on the memory and executable on the processor, the speech recognition program based on a microcomputer being configured to implement the steps in the speech recognition method based on a microcomputer according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps in the speech recognition method based on a microcomputer according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speech recognition model construction method and device, speech recognition method and device and storage medium
CN116013256A
Audio processing method and system for speech spectrum reconstruction
CN119541475A