A speech recognition method

By using a lightweight GMM acoustic model and an efficient matching algorithm in low-cost embedded devices, the problems of insufficient computing power and excessive power consumption are solved, efficient and accurate speech recognition is achieved, the device operating time is extended, and hardware costs are reduced.

CN120071911BActive Publication Date: 2025-10-21QINGDAO LINGLANG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510213818.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-10-21
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Low-cost embedded devices face the problems of insufficient computing power and excessive power consumption when performing speech recognition tasks. Traditional speech recognition methods are difficult to meet the needs of actual applications.

Method used

A lightweight GMM acoustic model is used for speech recognition. By obtaining speech frames and a candidate word library, the trained GMM model is used to determine speech features and generation probabilities. Combined with lightweight feature extraction and efficient matching algorithms, the computational complexity and energy consumption are reduced.

Benefits of technology

It achieves efficient and accurate speech recognition on low-power embedded devices, extending the device's operating time, reducing hardware costs, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071911B_ABST
    Figure CN120071911B_ABST
Patent Text Reader

Abstract

The application discloses a speech recognition method, and particularly relates to the technical field of speech recognition. The method comprises the following steps: obtaining at least one speech frame corresponding to a to-be-recognized speech signal and a candidate word library; determining speech features of each speech frame and inputting the speech features into a trained GMM acoustic model to obtain component probabilities of each speech frame belonging to each phoneme GMM; determining at least one target candidate word and corresponding HMM probabilities from the candidate word library according to the component probabilities of each speech frame belonging to each phoneme GMM; and determining generation probabilities of each speech frame corresponding to each target candidate word according to the HMM probabilities corresponding to each target candidate word and the component probabilities of each speech frame belonging to each phoneme GMM, so as to obtain a speech recognition result corresponding to the to-be-recognized speech signal. In the application, the speech recognition model is a simplified GMM model, the complexity of recognition operation is reduced, therefore, the energy consumption of a device can be reduced in an embedded system, and the demand of a low-power embedded device is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of speech recognition technology, and in particular to a speech recognition method. Background Art

[0002] With the rapid development of the Internet of Things and smart home technologies, embedded devices are being used more and more widely in daily life. As one of the important ways of human-computer interaction, the demand for speech recognition in embedded devices is also growing.

[0003] However, the chips used in many low-cost embedded devices face challenges with insufficient computing power and excessive power consumption when performing speech recognition tasks. Traditional speech recognition algorithms are inefficient on these devices, making them difficult to meet practical application requirements. Furthermore, traditional speech recognition methods often consume high power and are unsuitable for battery-powered embedded devices. Therefore, there is an urgent need for a speech recognition method specifically for low-power embedded devices. Summary of the Invention

[0004] The purpose of this application is to solve at least one of the above technical deficiencies.

[0005] In one aspect, an embodiment of the present application provides a speech recognition method, which is applied to a recognition chip in a low-power embedded device. The method includes:

[0006] Obtaining at least one speech frame corresponding to the speech signal to be recognized and a candidate word library, wherein the candidate word library includes each candidate word and an HMM probability corresponding to each candidate word;

[0007] Determine the speech features of each speech frame and input the speech features of each speech frame into the trained GMM acoustic model to obtain the component probability of each speech frame belonging to each phoneme GMM;

[0008] Determine at least one target candidate word from the candidate word library and an HMM probability corresponding to the at least one target candidate word based on the component probability of each speech frame belonging to each phoneme GMM;

[0009] Determine the generation probability of each speech frame corresponding to each target candidate word based on the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM;

[0010] According to the generation probability of each speech frame corresponding to each target candidate word, the speech recognition result corresponding to the speech signal to be recognized is obtained.

[0011] Optionally, obtaining at least one speech frame corresponding to the speech signal to be recognized includes:

[0012] Acquire an initial speech signal and extract the short-time energy features of the initial speech signal;

[0013] If the short-time energy feature is greater than a preset energy threshold, the initial speech signal is used as the speech signal to be recognized, and signal segmentation processing is performed on the speech signal to be recognized to obtain at least one speech frame corresponding to the speech signal to be recognized;

[0014] If the short-term energy characteristic is not greater than the preset energy threshold, the processing frequency and voltage of the low-power embedded device are adjusted according to a preset dynamic adjustment rule.

[0015] Optionally, obtaining a speech recognition result corresponding to the speech signal to be recognized based on the generation probability of each target candidate word corresponding to each speech frame includes:

[0016] For each speech frame, the target candidate word with the highest generation probability among the generation probabilities of each candidate word corresponding to the speech frame is taken as the speech recognition result corresponding to the speech frame;

[0017] The speech recognition results corresponding to each speech frame are combined according to the signal segmentation order to obtain the speech recognition results corresponding to the speech signal to be recognized.

[0018] Optionally, the HMM probability includes a forward probability and a backward probability. According to the HMM probability corresponding to each candidate word and the component probability of the speech frame belonging to each phoneme GMM, the generation probability of each speech frame corresponding to each target candidate word is determined, including:

[0019] For each target candidate word, multiply the forward probability and backward probability corresponding to the target candidate word to obtain the product probability corresponding to the candidate word;

[0020] For each speech frame, the component probability of the speech frame belonging to each phoneme GMM is added to the product probability corresponding to each candidate word to obtain the generation probability of the speech frame corresponding to each target candidate word.

[0021] Optionally, the GMM acoustic model is trained in the following way:

[0022] Obtain at least one sample data point and an initial GMM acoustic model, where the initial GMM acoustic model is obtained based on at least one Gaussian distribution;

[0023] Clustering at least one sample data point based on a preset clustering algorithm to obtain a cluster center, and determining the initial values ​​of each parameter in the initial GMM acoustic model according to the cluster center;

[0024] The initial values ​​of the parameters in the initial GMM acoustic model are iteratively trained according to at least one sample data point until the log-likelihood value determined based on the at least one sample data point meets the preset requirement, and the trained initial GMM acoustic model is used as the GMM acoustic model.

[0025] Optionally, the parameters include the number, mean, covariance, and weight of each Gaussian distribution included; iteratively training the initial values ​​of the parameters in the initial GMM acoustic model based on at least one sample data point until a log-likelihood value determined based on the at least one sample data point meets a preset requirement, including:

[0026] According to the initial values ​​of the parameters in the initial GMM acoustic model, the posterior probability of each sample data point belonging to each Gaussian distribution is determined;

[0027] According to the posterior probability that each sample data point belongs to each Gaussian distribution, the initial values ​​of each parameter in the initial GMM acoustic model are updated and adjusted to obtain the adjusted values ​​of each parameter;

[0028] Determine the log-likelihood value corresponding to this training based on the adjusted values ​​of each parameter and at least one sample data point;

[0029] If the log-likelihood value corresponding to this training does not meet the preset requirements, the adjusted values ​​of each parameter are used as the initial values ​​of each parameter to continue iterative training until the log-likelihood value determined based on at least one sample data point meets the preset requirements.

[0030] Optionally, at least one sample data point is determined by:

[0031] Acquire a sample speech signal and preprocess the initial speech signal to obtain a preprocessed sample speech signal;

[0032] Performing frame processing on the preprocessed sample speech signal according to a preset frame length to obtain at least one sample speech frame;

[0033] Feature extraction is performed on each sample speech frame to obtain features corresponding to each sample speech frame, and the features corresponding to each sample speech frame are used as at least one sample data point.

[0034] Optionally, the preprocessing includes at least one of signal denoising, pre-emphasis, and normalization, and performing feature extraction on each sample speech frame to obtain features corresponding to each sample speech frame includes:

[0035] For each sample speech frame, a fast Fourier transform is performed on the sample speech frame to obtain a spectrum corresponding to the sample speech frame;

[0036] Performing a Mel filter on the spectrum corresponding to the sample speech frame to obtain a Mel spectrum corresponding to the sample speech frame, and performing a logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum corresponding to the sample speech frame;

[0037] Perform discrete cosine transform on the logarithmic Mel spectrum corresponding to the sample speech frame to obtain the features corresponding to the sample speech frame.

[0038] Optionally, the GMM acoustic model is expressed by the following formula:

[0039]

[0040] Where λ represents the parameter set in the GMM acoustic model, K is the number of Gaussian distributions included in the GMM acoustic model, is the weight of the kth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the kth Gaussian distribution, is the covariance matrix of the kth Gaussian distribution.

[0041] Optionally, after obtaining the speech recognition result corresponding to the speech signal to be recognized, the method further includes:

[0042] The speech recognition results corresponding to the speech signal to be recognized are matched and converted, and the converted speech recognition results are transmitted to the universal interface of the low-power embedded device through the communication protocol.

[0043] On the other hand, an embodiment of the present application provides a speech recognition device, which is applied to a recognition chip in a low-power embedded device, and includes:

[0044] A data acquisition module is used to acquire at least one speech frame corresponding to the speech signal to be recognized and a candidate word library, wherein the candidate word library includes each candidate word and the HMM probability corresponding to each candidate word;

[0045] The component probability determination module is used to determine the speech features of each speech frame and input the speech features of each speech frame into the trained GMM acoustic model to obtain the component probability of each speech frame belonging to each phoneme GMM;

[0046] A generation probability determination module is configured to determine at least one target candidate word and an HMM probability corresponding to at least one target candidate word from a candidate word library based on the component probability of each speech frame belonging to each phoneme GMM; and determine a generation probability of each speech frame corresponding to each target candidate word based on the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM;

[0047] The recognition result determination module is used to obtain the speech recognition result corresponding to the speech signal to be recognized based on the generation probability of each speech frame corresponding to each target candidate word.

[0048] In another aspect, an embodiment of the present application provides an electronic device, including a processor and a memory:

[0049] The memory is configured to store machine-readable instructions, which, when executed by the processor, cause the processor to perform any one of the speech recognition methods.

[0050] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:

[0051] In this application, the GMM model used for speech recognition is a lightweight model with a simplified structure, which reduces the number of parameters in the model and thus reduces the complexity of the model during recognition operations. Furthermore, due to the low computational requirements of lightweight models, they can significantly reduce the energy consumption of mobile devices or embedded systems, allowing the speech recognition function to run for longer periods of time, meeting the application requirements of low-power embedded devices. Furthermore, they eliminate the need for frequent charging or battery replacement, extending battery life and enhancing the user experience.

[0052] Furthermore, due to the reduced size and computational complexity of the GMM model, operations can be completed more quickly when processing speech data, shortening the response time of the recognition process. Furthermore, by combining lightweight feature extraction algorithms with efficient matching algorithms, speech recognition tasks can maintain high efficiency and accuracy even on lower-performance chips (such as SoCs). Furthermore, the GMM model reduces the difficulty of training through algorithm optimization, reducing the amount of computation required for recognition during runtime, and thus significantly lowering the required hardware costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0054] Figure 1 A flowchart of a speech recognition method provided in an embodiment of the present application;

[0055] Figure 2 A schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application;

[0056] Figure 3A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present invention.

[0058] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0059] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0060] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0061] Specifically, such as Figure 1 As shown, the method is applied to an identification chip in a low-power embedded device, and the method may include:

[0062] Step S101 : obtaining at least one speech frame corresponding to a speech signal to be recognized and a candidate word library, wherein the candidate word library includes candidate words and HMM (Hidden Markov Model) probabilities corresponding to each candidate word.

[0063] Optionally, the speech recognition method provided in this application is applied in a recognition chip in a low-power embedded device, that is, when the low-power embedded device receives a speech signal to be recognized, the method provided in this application can be used to recognize the speech to be recognized.

[0064] In practical applications, a candidate word library can be pre-stored, which includes candidate words that can be used as speech recognition results, and the HMM probability corresponding to each candidate word. The HMM probability establishes a word or multiple factor sequence to generate a feature sequence, which specifically includes a transition probability and an observation probability, and then obtains the forward probability and backward probability based on the transition probability and the observation probability.

[0065] The transition probability is the probability of transitioning from one HMM state to another. These probabilities describe the conversion relationship between the phonemes or syllables of a word. The observation probability is the probability of generating a specific speech frame given a given HMM state. These probabilities describe how each HMM state generates the feature vector of the speech signal. Furthermore, an initial probability is assigned to the HMM state sequence of each candidate word, and a forward algorithm is used to calculate the forward probability of each HMM state under a given speech frame sequence. A backward algorithm is used to calculate the backward probability of each HMM state under a given speech frame sequence. The forward algorithm recursively calculates the cumulative probability from the initial state to each state, while taking into account both the state transition probability and the observation probability. The backward algorithm recursively calculates the cumulative probability from each state to the final state, while also taking into account both the state transition probability and the observation probability.

[0066] In an optional embodiment of the present application, obtaining at least one speech frame corresponding to the speech signal to be recognized includes:

[0067] Acquire an initial speech signal and extract the short-time energy features of the initial speech signal;

[0068] If the short-time energy feature is greater than a preset energy threshold, the initial speech signal is used as the speech signal to be recognized, and signal segmentation processing is performed on the speech signal to be recognized to obtain at least one speech frame corresponding to the speech signal to be recognized;

[0069] If the short-term energy characteristic is not greater than the preset capability threshold, the processing frequency and voltage of the low-power embedded device are adjusted according to a preset dynamic adjustment rule.

[0070] Optionally, when an initial voice information signal is received, it can be determined whether the initial voice signal needs to be processed for voice recognition, thereby avoiding processing of voice signals that do not need to be processed, thereby wasting power consumption of the device.

[0071] Specifically, the short-time energy characteristics of the received initial speech signal can be extracted and then compared with a preset energy threshold. If the short-time energy characteristics are greater than the preset energy threshold, the initial speech signal is considered to be a signal to be recognized. The recognition chip in the low-power embedded device then performs speech recognition on the speech signal to be recognized. For example, the recognition chip can first perform signal segmentation processing on the speech signal to be recognized to obtain at least one speech frame corresponding to the speech signal to be recognized. Conversely, if the short-time energy characteristics are not greater than the preset capacity threshold, the processing frequency and voltage of the low-power embedded device are dynamically adjusted according to preset dynamic adjustment rules, causing the speech processing chip to enter sleep mode, thereby reducing overall power consumption.

[0072] The short-term energy feature of the initial speech signal refers to the energy of the speech signal input within a certain time window, which can be obtained by calculating the sum of the squares of the signals:

[0073]

[0074] in, is the short-term energy at time t, is the speech signal, N is the frame length (usually 20-40 milliseconds of speech signal), and t is the current frame start time.

[0075] Step S102 : determining the speech features of each speech frame, and inputting the speech features of each speech frame into a trained GMM (Gaussian Mixture Model) acoustic model to obtain the component probability of each speech frame belonging to each phoneme GMM.

[0076] Optionally, the recognition chip in the low-power embedded device can include a trained GMM acoustic model. In this case, the speech features of each speech frame can be determined, and then the speech features of each speech frame are input into the trained GMM acoustic model to obtain the component probability of each speech frame belonging to each phoneme GMM. In the embodiment of the present application, a method such as Mel-Frequency Cepstral Coefficients (MFCC) can be used to extract features from each speech frame to obtain the speech features of each speech frame.

[0077] In an optional embodiment of the present application, the GMM acoustic model is trained in the following manner:

[0078] Obtain at least one sample data point and an initial GMM acoustic model, where the initial GMM acoustic model is obtained based on at least one Gaussian distribution;

[0079] Clustering at least one sample data point based on a preset clustering algorithm to obtain a cluster center, and determining the initial values ​​of each parameter in the initial GMM acoustic model according to the cluster center;

[0080] The initial values ​​of the parameters in the initial GMM acoustic model are iteratively trained according to at least one sample data point until the log-likelihood value determined based on the at least one sample data point meets the preset requirement, and the trained initial GMM acoustic model is used as the GMM acoustic model.

[0081] Among them, the GMM acoustic model is a Gaussian mixture model composed of several Gaussian distributions (also called normal distributions). It is a statistical model that can approximate complex data distributions by combining multiple Gaussian distributions, thereby achieving data modeling and classification.

[0082] Specifically, at least one sample data point and an initial GMM acoustic model can be obtained. The initial GMM acoustic model is derived based on at least one Gaussian distribution and has multiple parameters. In this case, the at least one sample data point can be clustered using a preset clustering algorithm (such as the K-Means clustering algorithm) to determine cluster centers. Initial values ​​for the parameters in the initial GMM acoustic model can then be set based on the determined cluster centers.

[0083] Furthermore, the initial values ​​of the parameters in the initial GMM acoustic model are iteratively trained based on at least one sample data point until the log-likelihood value determined based on at least one sample data point reaches the preset requirements, and the trained initial GMM acoustic model is used as the GMM acoustic model in this application. Among them, the preset requirements for the log-likelihood value to be reached can be set as needed, such as being set to the log-likelihood value reaching the maximum value. In actual training, in order to reduce the complexity and amount of calculation in the recognition process, the size of the final acoustic model can be reduced by reducing the Gaussian distribution clusters, increasing the number of iterations and strengthening the smoothing process during the training process, thereby reducing the complexity and amount of calculation in the recognition process.

[0084] In an optional embodiment of the present application, the parameters include the number, mean, covariance, and weights of each Gaussian distribution included; iteratively training the initial values ​​of the parameters in the initial GMM acoustic model based on at least one sample data point until the log-likelihood value determined based on the at least one sample data point meets a preset requirement, including:

[0085] According to the initial values ​​of the parameters in the initial GMM acoustic model, the posterior probability of each sample data point belonging to each Gaussian distribution is determined;

[0086] According to the posterior probability that each sample data point belongs to each Gaussian distribution, the initial values ​​of each parameter in the initial GMM acoustic model are updated and adjusted to obtain the adjusted values ​​of each parameter;

[0087] Determine the log-likelihood value corresponding to this training based on the adjusted values ​​of each parameter and at least one sample data point;

[0088] If the log-likelihood value corresponding to this training does not meet the preset requirements, the adjusted values ​​of each parameter are used as the initial values ​​of each parameter to continue iterative training until the log-likelihood value determined based on at least one sample data point meets the preset requirements.

[0089] Optionally, each parameter in the initial GMM acoustic model may include the number, mean, covariance, and weight of each Gaussian distribution included. Training the initial GMM acoustic model is to adjust the initial values ​​of the parameters in the initial GMM acoustic model based on at least one sample data point until the log-likelihood value reaches a preset requirement. Specifically, the posterior probability that the sample data point belongs to each Gaussian distribution can be first determined, and then the initial values ​​of each parameter in the initial GMM acoustic model can be updated and adjusted based on the determined posterior probability to obtain the adjusted values ​​of each parameter. Then, based on the adjusted values ​​of each parameter and the sample data points, the log-likelihood value corresponding to this training is determined, and it is judged whether the determined log-likelihood value meets the preset requirement. If the preset requirement is not met, the adjusted values ​​of each parameter are used as the initial values ​​of each parameter, and then the adjusted values ​​of each parameter are continued to be adjusted based on the sample data points until the log-likelihood value meets the preset requirement.

[0090] In order to better understand the training process of the GMM acoustic model in the embodiment of this application, the training process is described in detail below with reference to specific formulas. In this application, Gaussian distribution is the basis of the GMM acoustic model, and its probability density function (PDF) is:

[0091]

[0092] where μ is the mean, is the variance. This formula describes the distribution of data points around a certain center point. Since the GMM acoustic model is a Gaussian mixture model composed of several Gaussian distributions, the GMM acoustic model is the weighted sum of multiple Gaussian distributions. Its formula can be expressed as:

[0093]

[0094] Where λ represents the parameter set in the GMM acoustic model, K is the number of Gaussian distributions included in the GMM acoustic model, is the weight of the kth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the kth Gaussian distribution, is the covariance matrix of the kth Gaussian distribution.

[0095] In this application, the values ​​of the parameters in the GMM acoustic model can be obtained by training the expectation maximization (EM) algorithm, which can specifically include the following two steps:

[0096] E-step (expectation step): Determine the likelihood value of each sample data point belonging to each Gaussian distribution, which characterizes the possibility that each sample data point comes from each Gaussian distribution (that is, the possibility that the sample data point comes from a specific component), that is:

[0097]

[0098] in, is the weight of the kth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the kth Gaussian distribution, is the covariance matrix of the kth Gaussian distribution.

[0099] Furthermore, the likelihood value of each Gaussian distribution can be determined based on the likelihood value of each data point belonging to each Gaussian component, that is:

[0100]

[0101] in, is the weight of the jth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the jth Gaussian distribution, is the covariance matrix of the jth Gaussian distribution.

[0102] Furthermore, the posterior probability can be obtained based on the likelihood value of each sample data point belonging to each Gaussian distribution and the likelihood value of each Gaussian distribution:

[0103]

[0104] in, is the weight of the kth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the kth Gaussian distribution, is the covariance matrix of the kth Gaussian distribution, is the weight of the jth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the jth Gaussian distribution, is the covariance matrix of the jth Gaussian distribution, is the posterior probability that the nth data point belongs to the kth Gaussian distribution.

[0105] M step (maximization step): Based on the results of the E step, update the mean of the initial GMM acoustic model , covariance and the weights of the included Gaussian distributions ,in:

[0106]

[0107]

[0108]

[0109] in, is the posterior probability that the nth data point belongs to the kth Gaussian distribution, N represents the Gaussian distribution, represents the nth data point, k represents the kth Gaussian distribution, is the mean vector of the kth Gaussian distribution.

[0110] Furthermore, according to the determined mean , covariance and the weights of the included Gaussian distributions Determine the log-likelihood value of the data and determine whether the log-likelihood value is maximum. If so, end model training; otherwise, update and calculate the new mean, variance, and weight, and continue model training until the log-likelihood value of the data is maximum.

[0111] In an optional embodiment of the present application, at least one sample data point is determined by:

[0112] Acquire a sample speech signal and preprocess the initial speech signal to obtain a preprocessed sample speech signal;

[0113] Performing frame processing on the preprocessed sample speech signal according to a preset frame length to obtain at least one sample speech frame;

[0114] Feature extraction is performed on each sample speech frame to obtain features corresponding to each sample speech frame, and the features corresponding to each sample speech frame are used as at least one sample data point.

[0115] Optionally, when determining the sample data points, a sample speech signal may be obtained and then preprocessed to improve the robustness of the feature. Furthermore, the preprocessed sample speech signal may be framed according to a preset frame length to obtain at least one sample speech frame. Specifically, the sample speech signal may be divided into overlapping short frames, typically with each frame length of 25-50ms and an overlap rate of 50%.

[0116] The frame processing can be determined based on the following formula:

[0117]

[0118] in, is the mth frame signal, is the window function, N is the frame length, and n is the index of the discrete time series.

[0119] Accordingly, feature extraction can be performed on each sample speech frame, and the extracted features corresponding to each sample speech frame can be used as at least one sample data point. When extracting features from each sample speech frame, methods such as Mel-Frequency Cepstral Coefficients (MFCCs) can be used to extract features from the speech data. MFCCs simulate the human ear's perception of sound and convert the speech signal into a set of coefficients that can describe the signal's characteristics.

[0120] In an optional embodiment of the present application, the preprocessing includes at least one of signal denoising, pre-emphasis, and normalization, and feature extraction is performed on each sample speech frame to obtain features corresponding to each sample speech frame, including:

[0121] For each sample speech frame, a fast Fourier transform is performed on the sample speech frame to obtain a spectrum corresponding to the sample speech frame;

[0122] Performing a Mel filter on the spectrum corresponding to the sample speech frame to obtain a Mel spectrum corresponding to the sample speech frame, and performing a logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum corresponding to the sample speech frame;

[0123] Perform discrete cosine transform on the logarithmic Mel spectrum corresponding to the sample speech frame to obtain the features corresponding to the sample speech frame.

[0124] Optionally, the preprocessing of the sample speech signal may include at least one of signal denoising, pre-emphasis, and normalization. The pre-emphasis may be determined based on the following formula:

[0125] y(n)=x(n)−αx(n−1)

[0126] Wherein, x(n) is the sample speech signal, y(n) is the sample speech signal after pre-emphasis, and α is the pre-emphasis coefficient, which can usually be 0.95-0.97.

[0127] Furthermore, each sample speech frame may be subjected to a fast Fourier transform to obtain a spectrum corresponding to the sample speech frame. Specifically, the fast Fourier transform is performed using the following formula:

[0128]

[0129] in, is the spectrum of the mth frame, n is the index of the discrete time series, j is the imaginary unit, N is the frame length, and k is the kth Mel filter.

[0130] Accordingly, after obtaining the spectrum corresponding to the sample speech frame, the spectrum corresponding to the sample speech frame is subjected to Mel filtering to obtain the Mel spectrum corresponding to the sample speech frame, and the Mel spectrum is logarithmically transformed to obtain the logarithmic Mel spectrum corresponding to the sample speech frame. The Mel filtering can be determined based on the following formula:

[0131]

[0132] in, is the Mel spectrum of the mth frame, is the frequency response of the kth Mel filter, I represents the number of frequency components of the spectrum, is the spectrum of the mth frame.

[0133] The logarithmic Mel filter can be determined based on the following formula:

[0134]

[0135] in, is the logarithmic Mel spectrum of the mth frame, is the Mel spectrum of the mth frame.

[0136] Furthermore, a discrete cosine transform may be performed on the logarithmic Mel spectrum corresponding to the sample speech frame to obtain features corresponding to the sample speech frame, which may be determined based on the following formula:

[0137]

[0138] in, is the MFCC coefficient of the mth frame, K is the number of Mel filters, is the logarithmic Mel spectrum of the mth frame.

[0139] In this application, the first 12-13 MFCC coefficients are used because they contain most of the speech information. In actual applications, in order to improve the accuracy and robustness of recognition, the MFCC coefficients can be further processed, such as mean normalization, difference, etc.

[0140] Step S103 , determining at least one target candidate word and the HMM probability corresponding to the at least one target candidate word from the candidate word library according to the component probability of each speech frame belonging to each phoneme GMM.

[0141] Optionally, after obtaining the component probabilities of each speech frame belonging to each phoneme GMM, at least one target candidate word and the corresponding HMM probability can be determined from the candidate word library according to the component probabilities of each speech frame belonging to each phoneme GMM. For example, the input speech signal to be recognized is "Xiaojia Xiaojia", and "Xiaojia Xiaojia" will be split into multiple phonemes (i.e., speech frames), and the target candidate words "Xiao", "Jia", "Xiao", "Jia", and the HMM probabilities corresponding to "Xiao", "Jia", "Xiao", "Jia" respectively can be determined from the candidate word library according to the component probabilities of each split speech frame belonging to each phoneme GMM.

[0142] Among them, when there are many candidate words in the candidate word library, in order to make the calculation faster and improve the efficiency, the words in the candidate word library can be classified, and their respective category indexes can be established. Then, when determining the target candidate words, the category to which they belong can be determined first according to the index, and then the component probabilities of each speech frame belonging to each phoneme GMM can be compared with the HMM probabilities of the candidate words in this category to determine the target candidate words.

[0143] Step S104, determine the generation probability of each speech frame corresponding to each target candidate word according to the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM.

[0144] Optionally, for each speech frame, the generation probability of each target candidate word produced by this speech frame can be determined according to the component probability of this speech frame belonging to each phoneme GMM and each HMM probability, that is, the probability that each target candidate word is the speech recognition result of this speech frame.

[0145] In an optional embodiment of the present application, the HMM probability includes the forward probability and the backward probability. Determining the generation probability of each speech frame corresponding to each target candidate word according to the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM includes:

[0146] For each target candidate word, multiply the forward probability and the backward probability corresponding to the candidate word to obtain the product probability corresponding to the target candidate word;

[0147] For each speech frame, add the component probabilities of the speech frame belonging to each phoneme GMM to the product probabilities corresponding to each target candidate word respectively to obtain the generation probability of the speech frame corresponding to each target candidate word.

[0148] Optionally, each target candidate word includes a forward probability and a backward probability. In this case, the forward probability and the backward probability included in each target candidate word can be multiplied to obtain the product probability corresponding to each target candidate word. Accordingly, for each speech frame, the component probability of the speech frame belonging to each phoneme GMM can be added to the product probability corresponding to each target candidate word to obtain the probability of each target candidate word being the speech recognition result of the speech frame.

[0149] Step S105 , obtaining a speech recognition result corresponding to the speech signal to be recognized according to the generation probability of each speech frame corresponding to each target candidate word.

[0150] In an optional embodiment of the present application, obtaining a speech recognition result corresponding to the speech signal to be recognized according to the generation probability of each speech frame corresponding to each target candidate word includes:

[0151] For each speech frame, the target candidate word with the highest generation probability among the generation probabilities of the speech frame corresponding to each target candidate word is taken as the speech recognition result corresponding to the speech frame;

[0152] The speech recognition results corresponding to each speech frame are combined according to the signal segmentation order to obtain the speech recognition results corresponding to the speech signal to be recognized.

[0153] Optionally, for a speech frame, after obtaining the generation probability of the speech frame corresponding to each target candidate word, the generation probability of the speech frame corresponding to each target candidate word is compared, and the target candidate word corresponding to the maximum generation probability is used as the speech recognition result corresponding to the speech frame.

[0154] Furthermore, after obtaining the speech recognition result corresponding to each speech frame, the speech recognition results corresponding to the speech frames are combined according to the order in which the signal was segmented, and the combined speech recognition results are used as the speech recognition results corresponding to the speech signal to be recognized. For example, after segmenting the speech signal to be recognized, speech frame 1, speech frame 2, and speech frame 3 are obtained in sequence. At this time, the speech recognition result corresponding to speech frame 1 can be used as the first, the speech recognition result corresponding to speech frame 2 as the second, and the speech recognition result corresponding to speech frame 3 as the third, and the combined results can be used to obtain the speech recognition result corresponding to the speech signal to be recognized.

[0155] In an optional embodiment of the present application, after obtaining the speech recognition result corresponding to the speech signal to be recognized, the method further includes:

[0156] The speech recognition results corresponding to the speech signal to be recognized are matched and converted, and the converted speech recognition results are transmitted to the universal interface of the low-power embedded device through the communication protocol.

[0157] Optionally, after obtaining the speech recognition result corresponding to the speech signal to be recognized, the speech recognition result can be matched and converted, and then the converted speech recognition result is transmitted to the general interface of the low-power embedded device through the communication protocol for subsequent processing.

[0158] In this application, the GMM model used for speech recognition is a lightweight model with a simplified structure, which reduces the number of parameters in the model and thus reduces the complexity of the model during recognition operations. Furthermore, due to the low computational requirements of lightweight models, they can significantly reduce the energy consumption of mobile devices or embedded systems, allowing the speech recognition function to run for longer periods of time, meeting the application requirements of low-power embedded devices. Furthermore, they eliminate the need for frequent charging or battery replacement, extending battery life and enhancing the user experience.

[0159] Furthermore, due to the reduced size and computational complexity of the GMM model, operations can be completed more quickly when processing speech data, shortening the response time of the recognition process. Furthermore, by combining lightweight feature extraction algorithms with efficient matching algorithms, speech recognition tasks can maintain high efficiency and accuracy even on lower-performance chips (such as SoCs). Furthermore, the GMM model reduces the difficulty of training through algorithm optimization, reducing the amount of computation required for recognition during runtime, and thus significantly lowering the required hardware costs.

[0160] The embodiment of the present application provides a speech recognition device, which is applied to a recognition chip in a low-power embedded device, such as Figure 2 As shown, the apparatus may include: a data acquisition module 201, a component probability determination module 202, a generation probability determination module 203, and a recognition result determination module 204, wherein:

[0161] A data acquisition module is used to acquire at least one speech frame corresponding to the speech signal to be recognized and a candidate word library, wherein the candidate word library includes each candidate word and the HMM probability corresponding to each candidate word;

[0162] The component probability determination module is used to determine the speech features of each speech frame and input the speech features of each speech frame into the trained GMM acoustic model to obtain the component probability of each speech frame belonging to each phoneme GMM;

[0163] A generation probability determination module is configured to determine at least one target candidate word and an HMM probability corresponding to at least one target candidate word from a candidate word library based on the component probability of each speech frame belonging to each phoneme GMM; and determine a generation probability of each speech frame corresponding to each target candidate word based on the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM;

[0164] The recognition result determination module is used to obtain the speech recognition result corresponding to the speech signal to be recognized based on the generation probability of each speech frame corresponding to each target candidate word.

[0165] Optionally, when acquiring at least one speech frame corresponding to the speech signal to be recognized, the data acquisition module is specifically configured to:

[0166] Acquire an initial speech signal and extract the short-time energy features of the initial speech signal;

[0167] If the short-time energy feature is greater than a preset energy threshold, the initial speech signal is used as the speech signal to be recognized, and signal segmentation processing is performed on the speech signal to be recognized to obtain at least one speech frame corresponding to the speech signal to be recognized;

[0168] If the short-term energy characteristic is not greater than the preset energy threshold, the processing frequency and voltage of the low-power embedded device are adjusted according to a preset dynamic adjustment rule.

[0169] Optionally, when the recognition result determination module obtains the speech recognition result corresponding to the speech signal to be recognized based on the generation probability of each target candidate word corresponding to each speech frame, it is specifically used to:

[0170] For each speech frame, the target candidate word with the highest generation probability among the generation probabilities of each candidate word corresponding to the speech frame is taken as the speech recognition result corresponding to the speech frame;

[0171] The speech recognition results corresponding to each speech frame are combined according to the signal segmentation order to obtain the speech recognition results corresponding to the speech signal to be recognized.

[0172] Optionally, the HMM probability includes a forward probability and a backward probability. The generation probability determination module determines the generation probability of each speech frame corresponding to each target candidate word based on the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM, specifically for:

[0173] For each target candidate word, multiply the forward probability and backward probability corresponding to the target candidate word to obtain the product probability corresponding to the target candidate word;

[0174] For each speech frame, the component probability of the speech frame belonging to each phoneme GMM is added to the product probability corresponding to each target candidate word to obtain the generation probability of the speech frame corresponding to each target candidate word.

[0175] Optionally, the device further includes a model training module, which trains the GMM acoustic model in the following manner:

[0176] Obtain at least one sample data point and an initial GMM acoustic model, where the initial GMM acoustic model is obtained based on at least one Gaussian distribution;

[0177] Clustering at least one sample data point based on a preset clustering algorithm to obtain a cluster center, and determining the initial values ​​of each parameter in the initial GMM acoustic model according to the cluster center;

[0178] The initial values ​​of the parameters in the initial GMM acoustic model are iteratively trained according to at least one sample data point until the log-likelihood value determined based on the at least one sample data point meets the preset requirement, and the trained initial GMM acoustic model is used as the GMM acoustic model.

[0179] Optionally, the parameters include the number, mean, covariance, and weight of each Gaussian distribution included; the model training module iteratively trains the initial values ​​of the parameters in the initial GMM acoustic model based on at least one sample data point until a log-likelihood value determined based on the at least one sample data point meets a preset requirement, specifically for:

[0180] According to the initial values ​​of the parameters in the initial GMM acoustic model, the posterior probability of each sample data point belonging to each Gaussian distribution is determined;

[0181] According to the posterior probability that each sample data point belongs to each Gaussian distribution, the initial values ​​of each parameter in the initial GMM acoustic model are updated and adjusted to obtain the adjusted values ​​of each parameter;

[0182] Determine the log-likelihood value corresponding to this training based on the adjusted values ​​of each parameter and at least one sample data point;

[0183] If the log-likelihood value corresponding to this training does not meet the preset requirements, the adjusted values ​​of each parameter are used as the initial values ​​of each parameter to continue iterative training until the log-likelihood value determined based on at least one sample data point meets the preset requirements.

[0184] Optionally, at least one sample data point is determined by:

[0185] Acquire a sample speech signal and preprocess the initial speech signal to obtain a preprocessed sample speech signal;

[0186] Performing frame processing on the preprocessed sample speech signal according to a preset frame length to obtain at least one sample speech frame;

[0187] Feature extraction is performed on each sample speech frame to obtain features corresponding to each sample speech frame, and the features corresponding to each sample speech frame are used as at least one sample data point.

[0188] Optionally, the preprocessing includes at least one of signal denoising, pre-emphasis, and normalization. The model training module performs feature extraction on each sample speech frame to obtain features corresponding to each sample speech frame, specifically for:

[0189] For each sample speech frame, a fast Fourier transform is performed on the sample speech frame to obtain a spectrum corresponding to the sample speech frame;

[0190] Performing a Mel filter on the spectrum corresponding to the sample speech frame to obtain a Mel spectrum corresponding to the sample speech frame, and performing a logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum corresponding to the sample speech frame;

[0191] Perform discrete cosine transform on the logarithmic Mel spectrum corresponding to the sample speech frame to obtain the features corresponding to the sample speech frame.

[0192] Optionally, the GMM acoustic model is expressed by the following formula:

[0193]

[0194] Where λ represents the parameter set in the GMM acoustic model, K is the number of Gaussian distributions included in the GMM acoustic model, is the weight of the kth Gaussian distribution, N represents the Gaussian distribution, is the mean vector of the kth Gaussian distribution, is the covariance matrix of the kth Gaussian distribution.

[0195] Optionally, after obtaining the speech recognition result corresponding to the speech signal to be recognized, the recognition result determination module is further configured to:

[0196] The speech recognition results corresponding to the speech signal to be recognized are matched and converted, and the converted speech recognition results are transmitted to the universal interface of the low-power embedded device through the communication protocol.

[0197] The speech recognition device of this embodiment can execute the speech recognition method shown in the embodiment of this application. The implementation principle is similar and will not be repeated here.

[0198] An embodiment of the present application provides an electronic device, which includes: a processor; and a memory, wherein the memory is configured to store machine-readable instructions, which, when executed by the processor, enable the processor to perform a speech recognition method.

[0199] The present application embodiment provides an electronic device, such as Figure 3 As shown, Figure 3The electronic device shown includes a processor 2001 and a memory 2003. The processor 2001 and the memory 2003 are connected, for example, via a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the number of transceivers 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation on the embodiments of the present application.

[0200] Processor 2001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0201] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI bus or an EISA bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0202] The memory 2003 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0203] The memory 2003 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 2001. The processor 2001 is used to execute the application code stored in the memory 2003 to implement Figure 2 The illustrated embodiment provides operations of the speech recognition device.

[0204] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0205] The above description is only a partial embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A speech recognition method, characterized in that: The method is applied to an identification chip in a low-power embedded device, and the method includes: Obtaining at least one speech frame corresponding to a speech signal to be recognized and a candidate word library, wherein the candidate word library includes each candidate word and an HMM probability corresponding to each candidate word; Determine the speech features of each speech frame, and input the speech features of each speech frame into the trained GMM acoustic model to obtain the component probability of each speech frame belonging to each phoneme GMM; Determining at least one target candidate word from the candidate word library and an HMM probability corresponding to the at least one target candidate word according to the component probability of each of the speech frames belonging to each phoneme GMM; Determining a generation probability of each of the speech frames corresponding to each of the target candidate words according to the HMM probability corresponding to each of the target candidate words and the component probability of the speech frame belonging to each phoneme GMM; Obtaining a speech recognition result corresponding to the speech signal to be recognized according to a generation probability of each of the speech frames corresponding to each of the target candidate words; The HMM probability includes a forward probability and a backward probability, and determining the generation probability of each speech frame corresponding to each target candidate word according to the HMM probability corresponding to each target candidate word and the component probability of the speech frame belonging to each phoneme GMM includes: For each target candidate word, multiply the forward probability and the backward probability corresponding to the target candidate word to obtain the product probability corresponding to the candidate word; For each of the speech frames, the component probabilities of the speech frame belonging to each phoneme GMM are added to the product probabilities corresponding to each of the target candidate words to obtain the generation probability of the speech frame corresponding to each of the target candidate words.

2. The method according to claim 1, characterized in that The obtaining of at least one speech frame corresponding to the speech signal to be recognized includes: Acquiring an initial speech signal and extracting a short-time energy feature of the initial speech signal; If the short-time energy feature is greater than a preset energy threshold, the initial speech signal is used as the speech signal to be recognized, and signal segmentation processing is performed on the speech signal to be recognized to obtain at least one speech frame corresponding to the speech signal to be recognized; If the short-term energy characteristic is not greater than a preset energy threshold, the processing frequency and voltage of the low-power embedded device are adjusted according to a preset dynamic adjustment rule.

3. The method according to claim 1, characterized in that Obtaining a speech recognition result corresponding to the speech signal to be recognized according to the generation probability of each of the speech frames corresponding to each of the target candidate words includes: For each of the speech frames, taking the target candidate word with the highest generation probability among the generation probabilities of the speech frame corresponding to each of the target candidate words as the speech recognition result corresponding to the speech frame; The speech recognition results corresponding to each of the speech frames are combined according to the signal segmentation order to obtain the speech recognition results corresponding to the speech signal to be recognized.

4. The method according to claim 2, characterized in that The GMM acoustic model is trained in the following way: Obtain at least one sample data point and an initial GMM acoustic model, where the initial GMM acoustic model is obtained based on at least one Gaussian distribution; Clustering the at least one sample data point based on a preset clustering algorithm to obtain a cluster center, and determining initial values ​​of various parameters in the initial GMM acoustic model according to the cluster center; The initial values ​​of the parameters in the initial GMM acoustic model are iteratively trained according to the at least one sample data point until the log-likelihood value determined based on the at least one sample data point meets the preset requirement, and the trained initial GMM acoustic model is used as the GMM acoustic model.

5. The method according to claim 4, characterized in that The parameters include the number, mean, covariance, and weight of each Gaussian distribution included; and iteratively training the initial values ​​of the parameters in the initial GMM acoustic model based on the at least one sample data point until a log-likelihood value determined based on the at least one sample data point meets a preset requirement, including: Determining, based on initial values ​​of parameters in the initial GMM acoustic model, the posterior probability that each of the sample data points belongs to each Gaussian distribution; updating and adjusting the initial values ​​of the parameters in the initial GMM acoustic model according to the posterior probability that each of the sample data points belongs to each Gaussian distribution, to obtain adjusted values ​​of the parameters; Determining a log-likelihood value corresponding to the current training based on the adjusted values ​​of each parameter and the at least one sample data point; If the log-likelihood value corresponding to this training does not meet the preset requirements, the adjusted values ​​of each parameter are used as the initial values ​​of each parameter to continue iterative training until the log-likelihood value determined based on the at least one sample data point meets the preset requirements.

6. The method according to claim 5, characterized in that The at least one sample data point is determined by: Acquiring a sample speech signal and preprocessing the initial speech signal to obtain a preprocessed sample speech signal; Performing frame processing on the preprocessed sample speech signal according to a preset frame length to obtain at least one sample speech frame; Feature extraction is performed on each of the sample speech frames to obtain features corresponding to each of the sample speech frames, and the features corresponding to each of the sample speech frames are used as the at least one sample data point.

7. The method according to claim 6, characterized in that The preprocessing includes at least one of signal denoising, pre-emphasis, and normalization. The feature extraction of each of the sample speech frames to obtain features corresponding to each of the sample speech frames includes: For each of the sample speech frames, performing a fast Fourier transform on the sample speech frame to obtain a frequency spectrum corresponding to the sample speech frame; Performing a Mel filter on the spectrum corresponding to the sample speech frame to obtain a Mel spectrum corresponding to the sample speech frame, and performing a logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum corresponding to the sample speech frame; Performing discrete cosine transform on the logarithmic Mel spectrum corresponding to the sample speech frame to obtain features corresponding to the sample speech frame.

8. The method according to claim 1, characterized in that The GMM acoustic model is expressed by the following formula: ; Among them, λ represents the parameter set in the GMM acoustic model, K is the number of Gaussian distributions included in the GMM acoustic model, It is k The weight of a Gaussian distribution, N represents the Gaussian distribution, It is k The mean vector of a Gaussian distribution, It is k The covariance matrix of a Gaussian distribution.

9. The method according to claim 1, characterized in that After obtaining the speech recognition result corresponding to the speech signal to be recognized, the method further includes: The speech recognition result corresponding to the speech signal to be recognized is matched and converted, and the converted speech recognition result is transmitted to the universal interface of the low-power embedded device through a communication protocol.

Citation Information

Patent Citations

  • Speech recognition method, speech recognition device, computer equipment, and storage medium

    CN107331384A

  • Voice recognition method, voice recognition device and voice recognition program

    JP2012032538A