Speech recognition method and device, electronic equipment and computer readable storage medium
By building feature extraction models and speech enhancement technology on local devices, the problem of low robustness of speech recognition with limited computing resources is solved, real-time speech recognition on low-power chips is realized, multilingual support and environmental adaptability are improved, and global application needs are met.
Patent Information
- Application Number
- CN202510633894.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
AI Technical Summary
The existing offline speech recognition technology is less robust on local devices with limited computing resources, and it is difficult to achieve high-precision real-time speech recognition. It has a great impact on multilingual and multi-dialect support, environmental noise and accent differences, and the traditional acoustic-language model cascade architecture is insufficient.
The feature extraction model is used to form a convolutional layer, multiple first modules and feedforward neural networks, and control the number of modules and the number of hidden layer neurons. Through convolution processing and feature fusion, combined with acoustic decoding models and language models, locally deployed speech recognition is realized, and the generative adversarial network is used for speech enhancement and endpoint detection, and a prefix tree is built to accelerate phoneme sequence search.
Real-time voice recognition is realized on low-power chips, reducing computing complexity, improving recognition robustness and accuracy, meeting the needs of global applications, reducing dependence on the network, and protecting user privacy.
Smart Images

Figure CN120356460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet of Things technology. Specifically, this application relates to a voice recognition method, device, electronic device, and computer-readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence and voice recognition technology, offline voice recognition technology has gradually become a research hotspot. Traditional voice recognition systems usually rely on cloud processing and require real-time network connection, which has many limitations in practical applications, such as network latency, privacy leakage risk, and high bandwidth requirements. Therefore, offline voice recognition technology emerged, aiming to achieve efficient voice processing on local devices and reduce dependence on the network.
[0003] However, high-precision voice recognition has high requirements for computing resources and large parameters, making it difficult to run in real time on resource-constrained local devices. Therefore, there is a problem that the robustness of voice recognition is low when computing resources are limited. Summary of the Invention
[0004] Embodiments of this application provide a technical problem that the robustness of voice recognition is low when computing resources are limited.
[0005] According to the first aspect of the embodiments of this application, a voice recognition method is provided, which is applied to a terminal. The method includes: in response to receiving a voice collection instruction, obtaining target voice data; Inputting the target voice data into a feature extraction model deployed locally to obtain acoustic features output by the feature extraction model; the acoustic features are used to obtain the target text of the target voice data; Wherein, the feature extraction model includes a convolutional layer, a plurality of first modules, and a feedforward neural network. The number of first modules is less than a first preset threshold, and the number of neurons in the hidden layer of the feedforward neural network is less than a second preset threshold; Inputting the target voice data into a feature extraction model deployed locally to obtain acoustic features output by the feature extraction model, including: Inputting the target voice data into the convolutional layer for convolutional processing to obtain a first feature of the target voice data output by the convolutional layer; Inputting the first feature into a plurality of first modules respectively. Each first module determines the correlation between each feature value in the first feature based on its own parameters. For each feature value in the first feature, using the correlation between the feature value and each feature value, each feature value is fused with the corresponding feature value in the first feature to obtain a fused feature value corresponding to the feature value; The second feature is input into a feed-forward neural network for non-linear representation to obtain a third feature output by the feed-forward neural network, and an acoustic feature is obtained according to the third feature; wherein, each feature value in the second feature is a fusion feature value corresponding to each feature value in the first feature.
[0006] In a possible implementation, the feature extraction model is a sub-model of a locally deployed speech recognition model, and the speech recognition model further includes an acoustic decoding model and a language model as sub-models; after obtaining the acoustic feature output by the feature extraction model, the acoustic feature is input into the acoustic decoding model to obtain the phoneme probability distribution of the target speech data output by the acoustic decoding model; the phoneme probability distribution is input into the language model to obtain the target text of the target speech data output by the language model.
[0007] In another possible implementation, the speech recognition model is trained in the following manner: A teacher model and a student model are constructed; both the teacher model and the student model include a feature extraction model, an acoustic decoding model, and a language model to be trained, and the model parameters of the teacher model are more than those of the student model; Using the sample speech data as the training sample and the sample text corresponding to the sample speech data as the training label, the teacher model is trained to obtain the trained teacher model; Using the sample speech data as the training sample, the sample text corresponding to the sample speech data as the training label, and the text probability distribution output by the trained teacher model based on the training sample as the soft label, the student model is trained to obtain the trained student model; the text probability distribution is used to represent the probability distribution of each character corresponding to each moment in the training sample; The trained student model is pruned, and the pruned student model is used as the speech recognition model.
[0008] In yet another possible implementation, the importance degree of each neuron in the trained student model is determined, and the neurons with an importance degree less than a preset importance threshold are pruned; Using the sample speech data as the training sample and the sample text corresponding to the sample speech data as the training label, the pruned student model is trained to obtain a first model; The first model is quantized to obtain a second model; The calculation accuracy of the second model is evaluated. If the evaluation result meets the preset conditions, the second model is used as the speech recognition model; If the evaluation result does not meet the preset conditions, training is performed using the sample speech data as the training sample and the sample text corresponding to the sample speech data as the training label until the evaluation result meets the preset conditions.
[0009] In yet another possible implementation, inputting the phoneme probability distribution into a language model to obtain the target text of the target speech data output by the language model includes: Constructing a prefix tree based on the phoneme probability distribution; the i-th level of the prefix tree corresponds to the i-th moment in the target speech data, where i ∈ N and N is the number of moments in the target speech data. Each node in each level of the prefix tree except the root node corresponds to a phoneme, and each node records the probability value of the corresponding phoneme; Determining candidate nodes for each level starting from the first level, and taking the connections between candidate nodes in all adjacent levels as the final path, where the candidate nodes for each level are determined by the following method: For the first level, sorting the probability values included in each node in the first level, and taking the preset number of nodes with the largest probability values as the candidate nodes for the first level; For non-first levels, multiplying the probability results corresponding to each candidate path from the root node to the previous level by the probability values included in each node in the current level, and taking the preset number of nodes with the largest products as the candidate nodes for the current level; where each candidate path from the root node to the previous level includes one candidate node in each level from the first level to the previous level, and the probability value corresponding to each candidate path is the product of the probability values included in the corresponding candidate nodes.
[0010] In yet another possible implementation, obtaining speech data, performing noise reduction on the speech data to obtain denoised speech data; Performing speech enhancement on the denoised speech data through a pre-trained generative adversarial network to obtain speech-enhanced speech data; Performing endpoint detection on the speech-enhanced speech data to remove non-speech data and silent data in the speech-enhanced speech data to obtain effective speech data; Performing frame segmentation and windowing on the effective speech data to obtain target speech data.
[0011] According to the second aspect of the embodiments of the present application, there is provided a speech recognition device, which includes: An acquisition module, configured to acquire target speech data in response to receiving a speech acquisition instruction; An input module, configured to input the target speech data into a locally deployed feature extraction model to obtain acoustic features output by the feature extraction model; the acoustic features are used to obtain the target text of the target speech data; where the feature extraction model includes a convolutional layer, multiple first modules, and a feedforward neural network, the number of first modules is less than a first preset threshold, and the number of neurons in the hidden layer of the feedforward neural network is less than a second preset threshold; Input the target voice data into a locally deployed feature extraction model to obtain the acoustic features output by the feature extraction model, including: Input the target voice data into a convolutional layer for convolutional processing to obtain the first features of the target voice data output by the convolutional layer; Input the first features into multiple first modules respectively. Each first module determines the correlation between each eigenvalue in the first features based on its own parameters. For each eigenvalue in the first features, the eigenvalue is used to fuse each eigenvalue according to the correlation between the eigenvalue and each eigenvalue, and the fused eigenvalue corresponding to the eigenvalue is obtained; Input the second features into a feedforward neural network for non-linear representation to obtain the third features output by the feedforward neural network, and obtain the acoustic features according to the third features; wherein, each eigenvalue in the second features is the fused eigenvalue corresponding to each eigenvalue in the first features.
[0012] According to the third aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory. When the processor executes the program, the steps of the method provided in the first aspect are implemented.
[0013] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.
[0014] According to the fifth aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When a processor of a computer device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the computer device executes the steps of the method provided in the first aspect.
[0015] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are: The speech recognition method provided by the embodiments of the present application obtains target speech data by responding to a received speech collection instruction, inputs the target speech data into a feature extraction model deployed locally, and the feature extraction model outputs acoustic features. The acoustic features can be used to obtain the target text of the target speech data. Since the feature extraction model is composed of a convolutional layer, multiple first modules, and a feedforward neural network, in the process of obtaining acoustic features, the target speech data is input into the convolutional layer for convolutional processing to obtain the first features of the target speech data output by the convolutional layer, and then the first features are respectively input into multiple first modules. Each first module determines the correlation between the respective feature values in the first features based on its own parameters, performs fusion on each feature value using the corresponding fusion of each feature value in the first features to obtain the fusion feature value corresponding to the feature value, inputs the second features containing the fusion feature values corresponding to each feature value into the feedforward neural network for non-linear representation to obtain the third features output by the feedforward neural network, and obtains the acoustic features according to the third features. Since the feature extraction model controls the number of first modules to be less than a first preset threshold and controls the number of neurons in the hidden layer of the feedforward neural network to be less than a second preset threshold, it realizes the reduction of the model parameter quantity on the basis of maintaining the global attention mechanism of the feature extraction model, reduces the computational complexity, so that the acoustic feature extraction can be performed in real time on a low-power chip, and realizes the real-time operation of speech recognition on a local device with limited resources, solving the problems of limited computing resources and low recognition robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description in the embodiments of the present application.
[0017] Figure 1 It is a schematic flowchart of a speech recognition method provided by the embodiments of the present application; Figure 2 It is a schematic flowchart of a method for determining acoustic features in a speech recognition method provided by the embodiments of the present application; Figure 3 It is a schematic flowchart of a method for determining a target text in a speech recognition method provided by the embodiments of the present application; Figure 4 It is a schematic flowchart of a method for training a speech recognition model in a speech recognition method provided by the embodiments of the present application; Figure 5 It is a schematic flowchart of another method for training a speech recognition model in a speech recognition method provided by the embodiments of the present application; Figure 6 It is a schematic flowchart of another method for determining a target text in a speech recognition method provided by the embodiments of the present application; Figure 7 Flow chart of a method for obtaining a target voice in a voice recognition method provided by an embodiment of the present application; Figure 8 Flow chart of another method for determining a target text in a voice recognition method provided by an embodiment of the present application; Figure 9 Flow chart of a method for determining an instruction in a voice recognition method provided by an embodiment of the present application; Figure 10 Structural diagram of a voice recognition device provided by an embodiment of the present application; Figure 11 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0018] Embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0019] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "" and "the" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0020] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0021] The related technologies will be described below: Although significant progress has been made in existing offline voice recognition technologies, there are still the following problems: Contradiction between model complexity and computing resources: High-precision speech recognition models usually have high computational complexity. Deep learning models (such as Transformer and Conformer) have a large number of parameters and are difficult to run in real time on resource-constrained local devices (CPU / GPU).
[0022] Insufficient support for multiple languages and dialects: Existing offline speech recognition systems still lack support for multiple languages and dialects and are difficult to meet the needs of global applications.
[0023] Great impact of environmental noise and accent differences: In practical applications, the differences in environmental noise and user accents have a significant impact on the recognition accuracy. The robustness of existing offline systems in these aspects needs to be improved.
[0024] End-to-end latency requirement: It is necessary to achieve a millisecond-level response without cloud interaction, and the efficiency of the traditional acoustic-language model cascade architecture is insufficient.
[0025] In view of at least one of the above technical problems or areas for improvement in the related technologies, the present application proposes a speech recognition method. By responding to receiving a voice collection instruction, target voice data is obtained, and the target voice data is input into a feature extraction model deployed locally. The feature extraction model outputs acoustic features, and the acoustic features can be used to obtain the target text of the target voice data. Since the feature extraction model is composed of a convolutional layer, multiple first modules, and a feedforward neural network, in the process of obtaining acoustic features, the target voice data is input into the convolutional layer for convolutional processing to obtain the first feature of the target voice data output by the convolutional layer. Then, the first feature is respectively input into multiple first modules. Each first module determines the correlation between each eigenvalue in the first feature based on its own parameters, performs fusion on each eigenvalue using the eigenvalue corresponding to each eigenvalue in the first feature to obtain the fused eigenvalue corresponding to the eigenvalue, and inputs the second feature containing the fused eigenvalue corresponding to each eigenvalue into the feedforward neural network for non-linear representation to obtain the third feature output by the feedforward neural network. The acoustic features are obtained according to the third feature. Since the feature extraction model controls the number of first modules to be less than a first preset threshold and controls the number of neurons in the hidden layer of the feedforward neural network to be less than a second preset threshold, while maintaining the global attention mechanism of the feature extraction model, the number of model parameters is reduced, the computational complexity is reduced, so that acoustic feature extraction can be performed in real time on a low-power chip, and speech recognition can be run in real time on resource-constrained local devices, solving the problems of limited computing resources and low recognition robustness.
[0026] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0027] An embodiment of the present application provides a voice recognition method. As Figure 1 shown, the method includes: S101, in response to receiving a voice collection instruction, obtain target voice data.
[0028] In an embodiment of the present application, the voice recognition method is applied to the terminal of an intelligent device. The entire process of the voice recognition process can be executed on the terminal of the intelligent device without interacting with the cloud.
[0029] In an embodiment of the present application, the target voice data is the data obtained after the terminal collects the sound emitted by the user. Analyzing the target voice data can determine the intention of the user's speech.
[0030] In an embodiment of the present application, the voice collection instruction is used to instruct the terminal to collect voice information within a preset range. The voice collection instruction can be issued by the user through the terminal. For example, the user issues a voice collection instruction and starts speaking by triggering the recording button according to the application program installed on the local device for controlling the device. The terminal receives the voice collection instruction and obtains the target voice data.
[0031] In an embodiment of the present application, the issuance of the voice collection instruction can also be directly triggered by the sound emitted by the user. For example, if it is detected that there is sound within the preset range, the collection of the target voice data is automatically performed.
[0032] S102, input the target voice data into a locally deployed feature extraction model to obtain the acoustic features output by the feature extraction model.
[0033] In an embodiment of the present application, the acoustic features are used to obtain the target text of the target voice data. The acoustic features can be Mel Frequency Cepstral Coefficients or filter banks. The acoustic features of the target voice data are used to characterize the voice content of the target voice data or the features of the speaker. After obtaining the acoustic features, the target voice information can be further recognized according to the acoustic features to determine the target text of the target voice data.
[0034] In an embodiment of the present application, the locally deployed feature extraction model can be composed of a convolutional neural network or a recurrent neural network.
[0035] In an embodiment of the present application, the feature extraction model includes a convolutional layer, a plurality of first modules, and a feed-forward neural network. The number of first modules is less than a first preset threshold, and the number of neurons in the hidden layer of the feed-forward neural network is less than a second preset threshold.
[0036] In an embodiment of the present application, the feature extraction model includes a convolutional layer for performing a convolution operation on target speech data. The feature extraction model further includes a plurality of first modules for extracting features from important parts of the target speech data and establishing dependencies between features at different positions. The feature extraction model further includes capturing the non-linear structure in the target speech data to extract high-quality acoustic features.
[0037] In an embodiment of the present application, the plurality of first modules represent a multi-head self-attention layer. The number of first modules represents the number of heads of the multi-head self-attention. The number of first modules used to form the multi-head self-attention layer in the present application is less than a first preset threshold, that is, the self-attention layer formed in the present application is a self-attention layer with reduced number of self-attention heads, thereby realizing a reduction in model parameters, achieving memory savings and reduction of computing resources in an environment with limited resources.
[0038] In an embodiment of the present application, the feed-forward neural network formed in the present application obtains a feed-forward neural network with reduced hidden layer dimension by reducing the number of neurons in the hidden layer, that is, controlling the number of neurons in the hidden layer to be less than a second preset threshold, making the model structure simpler, reducing redundant computing and storage requirements, thereby realizing a reduction in model parameters and making the model more efficient.
[0039] In an embodiment of the present application, to reduce the number of model parameters, the first preset threshold is usually set to 8, and the second preset threshold is set to 512.
[0040] In an example, the number of attention heads is reduced from 8 to 4, and the hidden layer dimension is reduced from 512 to 256. While realizing a reduction in the number of model parameters, through multi-task joint training of tasks such as phoneme recognition, the accuracy loss is compensated, achieving a situation where while the number of model parameters is reduced by more than half, the high dependence of the end-to-end model on device computing power is also broken, achieving the purpose of real-time feature extraction on low-power chips. In an embodiment of the present application, the method for determining acoustic features is as Figure 2 shown, and the specific content is as follows: S201, input the target speech data into the convolutional layer for convolution processing to obtain the first feature of the target speech data output by the convolutional layer; S202. Input the first feature into multiple first modules respectively. Each first module determines the correlation between each eigenvalue in the first feature based on its own parameters. For each eigenvalue in the first feature, the eigenvalue is fused with each eigenvalue by using the correlation between the eigenvalue and each eigenvalue, and the fused eigenvalue corresponding to the eigenvalue is obtained; S203. Input the second feature into a feedforward neural network for non-linear representation, obtain the third feature output by the feedforward neural network, and obtain the acoustic feature according to the third feature.
[0041] In S201 of the embodiment of the present application, the first feature includes the frequency information and time domain information of the target speech data. The target speech data is input into the convolutional layer for convolutional processing. In the convolutional layer, multiple convolutional kernels are used to simultaneously extract different features of the target speech data. For example, some convolutional kernels may focus on feature extraction in the low-frequency part, other convolutional kernels may focus on the high-frequency part, or capture some special patterns (such as phoneme changes and tone changes in speech).
[0042] In S202 of the embodiment of the present application, the first feature is divided to obtain multiple eigenvalues. For each eigenvalue, self-attention processing is performed on the eigenvalue to obtain the first attention result corresponding to each eigenvalue. The first self-attention results of each eigenvalue are concatenated to obtain the fused eigenvalue corresponding to each eigenvalue.
[0043] In the embodiment of the present application, the first feature is split into multiple eigenvalues. Each eigenvalue determines the correlation between the eigenvalue and other eigenvalues based on an independent attention mechanism, enabling the feature extraction model to obtain different features of the target speech data from multiple dimensions. Each eigenvalue can focus on different parts of the information in the first feature to which the input belongs, different context relationships, or different eigenvalues with different distance dependencies. The eigenvalues can learn more general feature representations, which can also show good adaptability when facing new and different data, thereby extracting more effective features and improving the accuracy of speech recognition.
[0044] In S203 of the embodiment of the present application, each eigenvalue in the second feature is the fused eigenvalue corresponding to each eigenvalue in the first feature. The neurons in the feedforward neural network perform a linear transformation on the second feature, and then use an activation function for non-linear transformation to output the third feature. The third feature contains a more advanced representation of the target speech data. After obtaining the third feature, the third feature can be converted into specific acoustic features through a layer of linear transformation and a regression model.
[0045] The speech recognition method provided by the embodiment of the present application obtains target speech data by responding to a received speech acquisition instruction, inputs the target speech data into a feature extraction model deployed locally, and the feature extraction model outputs acoustic features. The acoustic features can be used to obtain the target text of the target speech data. Since the feature extraction model is composed of a convolutional layer, multiple first modules, and a feedforward neural network, in the process of obtaining acoustic features, the target speech data is input into the convolutional layer for convolutional processing to obtain the first feature of the target speech data output by the convolutional layer. Then, the first feature is respectively input into multiple first modules. Each first module determines the correlation between each eigenvalue in the first feature based on its own parameters, performs fusion on each eigenvalue using the corresponding eigenvalue in the first feature to obtain the fused eigenvalue corresponding to the eigenvalue, and inputs the second feature containing the fused eigenvalues corresponding to each eigenvalue into the feedforward neural network for non-linear representation to obtain the third feature output by the feedforward neural network. The acoustic features are obtained according to the third feature. Since the feature extraction model controls the number of first modules to be less than a first preset threshold and controls the number of neurons in the hidden layer of the feedforward neural network to be less than a second preset threshold, it realizes the reduction of the model parameter quantity while maintaining the global attention mechanism of the feature extraction model, reduces the computational complexity, so that acoustic feature extraction can be performed in real time on a low-power chip, and realizes real-time operation of speech recognition on a resource-constrained local device, solving the problems of limited computing resources and low recognition robustness.
[0046] Based on the above embodiments, as an optional embodiment, the feature extraction model is a sub-model of a speech recognition model deployed locally, and the speech recognition model further includes an acoustic decoding model and a language model as sub-models; In the embodiment of the present application, the method for obtaining the target text is as Figure 3 shown, and the specific content is as follows: S301, input the acoustic features into the acoustic decoding model to obtain the phoneme probability distribution of the target speech data output by the acoustic decoding model; S302, input the phoneme probability distribution into the language model to obtain the target text of the target speech data output by the language model.
[0047] In the embodiment of the present application, a phoneme is the smallest unit in speech, and the phoneme probability distribution is used to characterize the probability that each phoneme may appear at each moment of the target speech data. According to the phoneme probability, the most likely phoneme sequence of the target speech data can be inferred, so as to obtain the target text of the target speech data.
[0048] In the embodiment of the present application, a locally deployed speech recognition model is used to perform speech recognition on target speech data to determine the target text of the target speech data. The speech recognition model is composed of a feature extraction model, an acoustic decoding model, and a language model. The acoustic decoding model is used to determine the phoneme probability distribution of the target speech data according to acoustic features, and the language model is used to determine the target text of the target speech data according to the phoneme probability distribution.
[0049] In S301 of the embodiment of the present application, the acoustic features are input into the acoustic decoding model. The acoustic decoding model can be a deep neural network or a temporal classification model. The acoustic decoding model decodes the acoustic features, such as using the Viterbi algorithm to decode the acoustic features and calculate the probability of each phoneme appearing at each moment.
[0050] In S302 of the embodiment of the present application, the phoneme probability distribution is input into the language model. The language model can be an n-gram model or a neural network language model. The language model processes the phoneme probability distribution, and by maximizing the phoneme probability of each frame, a preliminary phoneme sequence is obtained. Then the language model predicts the most likely target text of the target speech data according to the phoneme sequence.
[0051] In the above solution, the acoustic features are processed by the acoustic decoding model to obtain the phoneme probability distribution, and the target text of the target speech data is inferred through the language model and the phoneme probability, which can improve the recognition accuracy. Especially when dealing with noise, accent changes, and context ambiguity, the roles of the acoustic decoding model and the language model are particularly important.
[0052] Based on the above embodiments, as an alternative embodiment, the training method of the speech recognition model is as Figure 4 shown, and the specific content is as follows: S401, construct a teacher model and a student model; S402, use the sample speech data as the training sample, and the sample text corresponding to the sample speech data as the training label to train the teacher model to obtain the trained teacher model; S403, use the sample speech data as the training sample, the sample text corresponding to the sample speech data as the training label, and the text probability distribution output by the trained teacher model based on the training sample as the soft label to train the student model to obtain the trained student model; the text probability distribution is used to represent the probability distribution of each character corresponding to each moment in the training sample; S404, perform pruning processing on the trained student model, and use the pruned student model as the speech recognition model.
[0053] In the embodiments of the present application, both the teacher model and the student model include a feature extraction model to be trained, an acoustic decoding model, and a language model, and the model parameters of the teacher model are more than those of the student model. That is to say, the teacher model is a model with a relatively complex structure and high computational requirements. The trained teacher model has a high-accuracy speech recognition effect. The student model is a smaller and simpler network. During the process of speech recognition, the student model consumes less computational resources. The teacher model is used to guide the training of the student model.
[0054] In S401 of the embodiments of the present application, a relatively complex teacher model with high computational resource requirements is constructed, and a smaller student model with lower computational resource requirements is constructed.
[0055] In S402 of the embodiments of the present application, a large amount of sample speech data is used as training samples, and the sample text corresponding to the sample speech data is used as training labels to train the teacher model, and a trained teacher model is obtained.
[0056] In S403 of the embodiments of the present application, the teacher model outputs a text probability distribution according to the input sample speech data, the student model outputs a text probability distribution according to the same sample speech data as the teacher model, the difference between the text probability distribution output by the student model and the text probability distribution output by the teacher model is calculated, the student model also outputs a predicted label according to the sample speech data, the difference between the predicted label and the training label is calculated, and according to the difference between the text probability distribution output by the student model and the text probability distribution output by the teacher model and the difference between the predicted label and the training label, the parameters of the student model are updated by calculating the gradient and using the gradient descent algorithm. After the training is completed, the student model is evaluated, and a validation set or a test set can be used to detect its performance.
[0057] In S404 of the embodiments of the present application, the trained student model is pruned, and the less important part in the model structure is subtracted to obtain a speech recognition model.
[0058] In the above solution, during the process of training the speech recognition model, the student model is guided and trained by the teacher model with a complex model structure and high speech recognition accuracy. While reducing the number of parameters, the accuracy can also be ensured not to be lost. By pruning and training the unimportant structures in the student model, the volume of the model is reduced so that the model can also run on low-resource local devices. On the basis of the above embodiments, as an optional embodiment, a method for training a speech recognition model is further provided, as Figure 5 shown, and the specific content is as follows: S501. Determine the importance of each neuron in the trained student model, and prune the neurons whose importance is less than a preset importance threshold. S502. Use the sample speech data as the training sample and the corresponding sample text of the sample speech data as the training label to train the pruned student model to obtain a first model. S503. Quantize the first model to obtain a second model. S504. Evaluate the computing accuracy of the second model. S505-1. If the evaluation result meets the preset conditions, use the second model as the speech recognition model. S505-2. If the evaluation result does not meet the preset conditions, use the sample speech data as the training sample and the corresponding sample text of the sample speech data as the training label to train until the evaluation result meets the preset conditions.
[0059] In S501 of the embodiment of the present application, the importance of each neuron in the student model is determined according to indicators such as the activation value, gradient information, and importance of the neuron, and the neurons with an importance less than the preset importance threshold are selected for pruning.
[0060] In the embodiment of the present application, the entire layer or sub-network of the student model can also be removed, and some channels in the convolutional layer of the student model can also be removed. For example, redundant convolutional kernels are removed based on channel importance scoring (such as L1-norm) so that the volume of the model is reduced by 40%.
[0061] In S502 of the embodiment of the present application, after removing the neurons with lower importance in the student model, the student model is trained again. Using the sample speech data as the training sample and the corresponding sample text of the sample speech data as the training label, the pruned student model is trained to adjust the weights of the remaining neurons to restore the performance of the student model and obtain a first model. During the process of training the pruned student model, a lower learning rate can be used for training to avoid over-adjusting the already learned weights. After training the pruned student model, the student model needs to be evaluated, such as evaluating the computing efficiency and accuracy. When the expected standard is met, stop training to obtain a first model.
[0062] In S503 of the embodiments of the present application, quantizing the model refers to the process of converting high-precision model parameters into low-precision model parameters, such as the process of converting 32-bit floating-point numbers into 8-bit integers. Determine the quantization bit width, that is, determine how many bits of integers the floating-point numbers are mapped to. According to each layer of the first model, determine the scaling factor used to characterize the size of the quantization step and the zero point used to characterize the point that maps the zero value to an integer during the quantization process. Quantize the first model according to the quantization bit width, scaling factor, and zero point to obtain the second model.
[0063] In S504 of the embodiments of the present application, evaluate whether the degree of decrease in the accuracy of the second model meets the expected target, that is, evaluate whether the loss of calculation precision is large. For example, determine the precision loss of the second model by comparing the difference in the accuracy rates of the first model and the second model on the same test set.
[0064] In S505-1 of the embodiments of the present application, if the evaluation result meets the preset conditions, directly use the second model as the speech recognition model.
[0065] In S505-2 of the embodiments of the present application, if the evaluation result does not meet the preset conditions, continue to train the second model to make the second model adapt to low-precision calculations until the preset conditions are met.
[0066] In the embodiments of the present application, the second model that does not meet the preset conditions can also be trained by means of quantization-aware training.
[0067] In the embodiments of the present application, the model can also be dynamically quantized during model inference, and the quantization parameters are adjusted according to the actual input data and the characteristics of the model.
[0068] In one example, technologies such as model quantization, pruning, and distillation are used to optimize the speech recognition model to reduce the computational complexity and storage space of the model, so that it can run efficiently on local devices.
[0069] In the above solution, by pruning the student model, the number of model parameters is reduced, thereby improving the calculation efficiency and reducing the memory consumption, while maintaining the performance of the model as much as possible. By quantizing the model, the storage space and computational overhead of the model are reduced, and it will not have a great impact on the accuracy of the model. Through precision evaluation, it is confirmed whether unacceptable precision loss occurs after quantization; by training the second model, the precision of the model is restored or improved, ensuring that the model still maintains good inference performance while reducing storage and computational overhead.
[0070] Based on the above embodiments, as an optional embodiment, a method for determining the target text is provided as Figure 6 shown, and the specific content is as follows: S601. Construct a prefix tree based on the phoneme probability distribution; S602. Starting from the first level, determine the candidate nodes at each level layer by layer, and take the connections between the candidate nodes of all adjacent levels as the final paths; S603. According to the probability values of the candidate nodes, determine the probability values of the first paths, and select the first path with the largest probability value as the target path; S604. Determine the speech sequence according to the target path, and determine the target text according to the phoneme sequence.
[0071] In the embodiments of the present application, the phoneme probability distribution is used to represent the probability of each phoneme at each moment in the target speech data. The i-th level of the prefix tree corresponds to the i-th moment in the target speech data, where i ∈ N, and N is the number of moments in the target speech data. Each node in each level of the prefix tree except the root node corresponds to a phoneme, and each node records the probability value of the corresponding phoneme.
[0072] In S601 of the embodiments of the present application, a prefix tree is constructed based on the phoneme probability distribution. Each node of the prefix tree stores a phoneme and the probability value of the occurrence of the phoneme at the moment corresponding to the layer where the node is located. The nodes of the prefix tree are connected by lines.
[0073] In S602 of the embodiments of the present application, the candidate nodes at each level are determined based on the probability values of the nodes at that level. The connections between the candidate nodes of adjacent levels together form paths. Since each level includes at least one candidate node, the connections between the candidate nodes of adjacent levels can form multiple first paths.
[0074] In S603 of the embodiments of the present application, according to the probability values of the candidate nodes of each first path, the probability values of each first path can be determined. The larger the probability value of the first path, the more it can represent the phoneme sorting of the target speech data. Therefore, select the candidate path with the largest probability value as the target path.
[0075] In S604 of the embodiments of the present application, according to the order from the root node to the leaf node, arrange the phonemes included in the target path in sequence to obtain a phoneme sequence, and determine the target text of the target speech data according to the phoneme sequence and the mapping relationship between phonemes and characters.
[0076] In the embodiments of the present application, for the first level, sort the probability values included in each node in the first level, and take the preset number of nodes with the largest probability values as the candidate nodes of the first level.
[0077] In the embodiment of the present application, since the prefix tree contains the probabilities of each phoneme, and the probability of some phonemes appearing at the current moment is relatively low. Therefore, in the subsequent search process, for the nodes at the first level, first sort the probability values included in each node, and only select the preset number of nodes with the largest probability values as candidate nodes, and start searching for the nodes at the next level from the candidate nodes, thereby reducing the time required for the search and avoiding ineffective searches.
[0078] In the embodiment of the present application, for non-first levels, the probability results corresponding to each candidate path from the root node to the previous level are multiplied by the probability values included in each node in the current level, and the preset number of nodes with the largest product are used as the candidate nodes in the current level.
[0079] In the embodiment of the present application, each candidate path from the root node to the previous level includes one candidate node in each level from the first level to the previous level, and the probability value corresponding to each candidate path is the product of the probability values included in the corresponding candidate nodes.
[0080] In the embodiment of the present application, the probability values corresponding to each candidate node from the root node to the previous level are determined based on the product of the probability values of the nodes on the connection line. Each probability value corresponds to a different connection line combination. For each candidate path, determine the product of the probability value of the candidate path and the probability values of each node in the current level, and use the preset number of nodes with the largest product as the candidate nodes in the current level. That is, after determining the candidate nodes in the current level for the connection line combination formed by the root node and each candidate node in the previous level, a new candidate path is obtained.
[0081] In the above solution, by constructing a prefix tree based on the phoneme probability distribution, the search for the phoneme sequence can be accelerated, and during the search process, only the candidate paths with higher probabilities are retained, which can greatly reduce the computational complexity during the search process. By restricting the preset number of candidate paths to be retained, the search speed of the phoneme sequence can be greatly improved under the condition of losing a large amount of accuracy. And with the support of the prefix tree, the tree structure is used to quickly exclude the sequences with smaller probabilities, ensuring high efficiency while avoiding the excessive computational cost brought by the search.
[0082] Based on the above embodiments, as an alternative embodiment, the method for obtaining the target voice is as Figure 7 shown, and the specific content is as follows: S701. Obtain voice data, perform noise reduction on the voice data to obtain the noise-reduced voice data; S702. Perform voice enhancement on the noise-reduced voice data through a pre-trained generative adversarial network to obtain the voice-enhanced voice data; S703 performs endpoint detection on the speech data after speech enhancement, removes the non-speech data and silent data in the speech data after speech enhancement, and obtains the effective speech data; S704 performs frame segmentation and windowing on the effective speech data to obtain the target speech data.
[0083] In S701 of the embodiments of the present application, a noise reduction algorithm is used to reduce the noise of the speech data in the frequency domain or the time domain, and post-processing is performed on the speech data after noise reduction, such as smoothing, removing residual noise, etc., to obtain the speech data after noise reduction. For example, the real-time noise suppression based on deep learning (improved version of RNNoise) is adopted to dynamically separate the speech from the background noise.
[0084] In the embodiments of the present application, the generative adversarial network consists of a generator and a discriminator. The goal of the generator is to generate samples as close as possible to the real data, while the goal of the discriminator is to distinguish between real data and generated data. During the process of training the generative adversarial network, the generator and the discriminator confront each other, so that the generator gradually learns to generate more and more realistic data. The speech data after speech enhancement has the characteristics of higher speech data quality, less noise influence, and clearer and more natural speech.
[0085] In S702 of the embodiments of the present application, the speech data after noise reduction is input into the generator of the generative adversarial network. The generator maps the speech data after noise reduction into clearer and more natural speech data, generates a more realistic speech waveform, and inputs the output of the generator into the discriminator. The discriminator determines whether the output speech data is close to the real speech data. When it is determined that the speech data output by the generator is close to the real speech data, the output of the generator is used as the speech data after speech enhancement.
[0086] In the embodiments of the present application, if the speech data output by the generator does not meet the preset output standard, the discriminator sends feedback to the generator to make the output of the generator be improved in the direction closer to the real speech data.
[0087] In S703 of the embodiments of the present application, endpoint detection is used to identify the start and end times of the speech signal in the speech data, so that the non-speech data and silent data in the speech data after speech enhancement can be removed, and the effective speech data can be obtained. For example, by combining short-time energy and the LSTM classifier, the start and end points of the speech are accurately located, thus avoiding invalid calculations in the subsequent speech recognition process.
[0088] In S704 of the embodiments of the present application, framing the speech data is to segment the continuous speech signal into small time segments with a certain overlap between the time segments, so as to ensure the continuity of the speech data and avoid the loss of speech data. Windowing the speech data is to perform a window function process on each frame signal to reduce the spectrum leakage phenomenon caused by framing. The role of the window function is to reduce the boundary effect of the signal by gradually reducing both ends of each frame signal to zero, so as to obtain the target speech data with complete and high-quality audio data.
[0089] In one example, the input speech data is framed in units of 10 ms, and noise suppression, gain adjustment, and speech enhancement are performed in parallel to output high-quality target speech data.
[0090] In the above solution, the collected speech data is subjected to noise reduction, speech enhancement, endpoint detection, as well as framing and windowing processing, so as to output high-quality target speech data, improve the accuracy of speech recognition, and at the same time accurately locate the start and end points of the speech to avoid invalid calculations. The recognition accuracy is still significantly improved under the differences of environmental noise and user accents.
[0091] Based on the above embodiments, as an optional embodiment, the method for determining the target text is as Figure 8 shown, and the specific content is as follows: S801, according to a preset dictionary, determine the character sequence corresponding to the phoneme sequence; S802, determine the context information of the target speech data, and determine the probability value that the character sequence expresses the correct semantics of the target speech data according to the context information; S803-1, if the probability value is greater than a preset correct threshold, then use the candidate word sequence as the target text of the target speech data; S803-2, if the probability value is not greater than the preset correct threshold, then determine the target text based on the context information, phoneme sequence, and dictionary.
[0092] In S801 of the embodiments of the present application, the dictionary includes each phoneme and a character having a mapping relationship with each phoneme. Determine the phonemes included in the phoneme sequence, find the characters corresponding to each phoneme from the preset dictionary, and arrange the characters corresponding to each phoneme in the order of arrangement of the phonemes in the phoneme sequence, so as to obtain the character sequence.
[0093] In the embodiments of the present application, an N-gram enhanced language model is constructed based on the TinyBERT architecture, so that the number of parameters of the language model is reduced to 10% of that of the traditional BERT.
[0094] In S802 of the embodiments of the present application, a language model can be constructed through the text of the device's historical speech recognition, and the semantics of the current character can be predicted based on historical information, or the probability of the character corresponding to the target speech data can be predicted using the characters before the time series of the target speech data, so as to determine the probability value that the character sequence expresses the correct semantics of the target speech data.
[0095] In S803-1 of the embodiments of the present application, if the probability value that the character sequence generated based on the phoneme sequence expresses the correct semantics of the target speech data is greater than the correct threshold, it indicates that the character sequence can be directly used as the target text of the target speech data.
[0096] In S803-2 of the embodiments of the present application, if the probability value that the character sequence generated based on the phoneme sequence expresses the correct semantics of the target speech data is not greater than the correct threshold, it indicates that the character sequence generated based on the phoneme sequence does not match the content that the target speech data wants to express, and it needs to be optimized. Therefore, the character sequence is optimized based on the context information, phoneme sequence, and dictionary to determine the target text.
[0097] In one example, multiple candidate character sequences corresponding to phonemes are determined, and the most likely character sequence is selected from the multiple candidate character sequences as the target text according to the known context information and dictionary. In addition, after selecting a character sequence from the multiple candidate character sequences, the candidate character sequence can also be optimized according to the context information and dictionary, and the optimized character sequence is used as the target text. For example, according to the context before and after the target speech data, the topic vocabulary related to the target speech data, and the vocabulary of the target speech data, based on morphological analysis or syntactic analysis, the character sequence is optimized. The optimization content can include: correcting misspellings, selecting more appropriate vocabulary, etc. For example, it is judged whether the characters in the character sequence are in the dictionary. If not, the character sequence is optimized by means of pinyin correction.
[0098] In the embodiments of the present application, the output phoneme sequence is jointly decoded with the language model, and the phoneme sequence, language model, and dictionary are fused through a weighted finite state transducer to generate the final target text.
[0099] In the embodiments of the present application, matrix operations are optimized for the ARM NEON instruction set to achieve a three-fold increase in the CPU inference speed.
[0100] The speech recognition method provided by the embodiments of the present application combines advanced deep learning models and model optimization techniques to achieve efficient and accurate speech recognition on local devices. During the process of training the speech recognition model, each sub-model is optimized to reduce the computational complexity and storage space of the model, enabling it to run efficiently on local devices.
[0101] In the embodiments of the present application, an advanced deep learning model and an end-to-end training method are adopted to improve the accuracy of speech recognition. All calculations are completed on local devices without network transmission, reducing the latency of speech recognition and meeting the requirements of real-time conversations. In addition, there is no need to upload the user's speech data to the cloud, protecting user privacy. All data is processed locally, achieving high-precision recognition effects on low-computing-power devices and providing a reliable technical foundation for scenarios such as the Internet of Things and mobile terminals.
[0102] In one example, when the terminal is a smart clothes dryer terminal, a method for determining instructions is also proposed, as Figure 9 shown, and the specific content is as follows: S901. Filter the character sequence based on the drying scenario rules of the smart clothes dryer to obtain the target text; In one example, if the smart clothes dryer does not support the ironing function and the two characters "ironing" exist in the character sequence, then the two characters "ironing" in the character sequence are removed. When the drying rod of the smart clothes dryer is at the highest point and the two characters "rising" exist in the character sequence, then the two characters "rising" in the character sequence are removed; when the environment where the smart clothes dryer is located is a high-temperature weather and the four characters "high-temperature drying" exist in the character sequence, then the four characters "high-temperature drying" in the character sequence are removed.
[0103] S902. Determine the target score of the target text based on the phoneme sequence and its corresponding first weight, context information and its corresponding second weight, and user portrait and its corresponding third weight.
[0104] In the embodiments of the present application, the sum of the first weight, the second weight, and the third weight is 1. For example, the first weight is 0.4, the second weight is 0.4, and the third weight is 0.2.
[0105] S903. If the target score is greater than the preset threshold, determine whether to execute the target instruction corresponding to the target text according to the security decision tree; In one example, if the target instruction corresponding to the target text is to quickly descend, when it is determined that there are no obstacles below the current drying rod, execute the target instruction; if it is determined that there are obstacles below, then perform a slow descent, such as reducing the descent speed of the drying rod to 2 cm / s.
[0106] S904. If the target score is not greater than the preset threshold, perform confirmation processing on the target instruction according to the preset confirmation rule.
[0107] In one example, if the target score is between 0.5 and 0.7, a confirmation page of a confirmation control for confirming the execution of the target instruction is pushed and displayed through the APP, and when the confirmation control is triggered, the intelligent clothes dryer is instructed to perform corresponding operations; if the target score is between 0.3 and 0.5, "Please say the confirmation instruction" is played three times through the speaker, and if the voice data corresponding to the "confirmation instruction" is collected within the preset time period, the intelligent clothes dryer is instructed to perform corresponding operations; if the target score is less than 0.3, "Please re - say the instruction" is played through the speaker, and the target voice data collection process is entered. S905. Update the knowledge base of the intelligent clothes dryer based on the execution content of the target instruction.
[0108] An embodiment of the present application provides a voice recognition device, as Figure 10 shown. The voice recognition device 100 may include: an acquisition module 1001 and an input module 1002. Specifically, the acquisition module 1001 is configured to obtain target voice data in response to receiving a voice acquisition instruction; The input module 1002 is configured to input the target voice data into a locally - deployed feature extraction model to obtain acoustic features output by the feature extraction model; the acoustic features are used to obtain the target text of the target voice data; Wherein, the feature extraction model includes a convolutional layer, a plurality of first modules, and a feed - forward neural network. The number of first modules is less than a first preset threshold, and the number of neurons in the hidden layer of the feed - forward neural network is less than a second preset threshold; Inputting the target voice data into the locally - deployed feature extraction model to obtain the acoustic features output by the feature extraction model includes: Inputting the target voice data into the convolutional layer for convolutional processing to obtain a first feature of the target voice data output by the convolutional layer; Inputting the first feature into a plurality of first modules respectively. Each first module determines the correlation between each feature value in the first feature based on its own parameters. For each feature value in the first feature, the feature value is fused with the correlation between each feature value to obtain a fused feature value corresponding to the feature value; Inputting the second feature into the feed - forward neural network for non - linear representation to obtain a third feature output by the feed - forward neural network, and obtaining the acoustic features according to the third feature; wherein, each feature value in the second feature is the fused feature value corresponding to each feature value in the first feature.
[0109] The voice recognition device provided by the embodiment of the present application obtains target voice data by responding to a received voice collection instruction, inputs the target voice data into a feature extraction model deployed locally. The feature extraction model outputs acoustic features, which can be used to obtain the target text of the target voice data. Since the feature extraction model is composed of a convolutional layer, multiple first modules, and a feedforward neural network, in the process of obtaining acoustic features, the target voice data is input into the convolutional layer for convolutional processing to obtain the first feature of the target voice data output by the convolutional layer. Then, the first feature is respectively input into multiple first modules. Each first module determines the correlation between each eigenvalue in the first feature based on its own parameters, performs fusion corresponding to each eigenvalue in the first feature for each eigenvalue, and obtains the fused eigenvalue corresponding to the eigenvalue. The second feature containing the fused eigenvalues corresponding to each eigenvalue is input into the feedforward neural network for non-linear representation to obtain the third feature output by the feedforward neural network. The acoustic features are obtained according to the third feature. Since the feature extraction model controls the number of first modules to be less than a first preset threshold and controls the number of neurons in the hidden layer of the feedforward neural network to be less than a second preset threshold, on the basis of maintaining the global attention mechanism of the feature extraction model, the reduction of the model parameter quantity is realized, the calculation complexity is reduced, so that the acoustic feature extraction can be carried out in real time on a low-power chip, and the voice recognition can be run in real time on a local device with limited resources, solving the problems of limited computing resources and low recognition robustness.
[0110] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described here again.
[0111] An electronic device (computer device / equipment / system) is provided in the embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the voice recognition method. Compared with the related technology, it can be realized that on the basis of maintaining the global attention mechanism of the feature extraction model, the reduction of the model parameter quantity is realized, the calculation complexity is reduced, so that the acoustic feature extraction can be carried out in real time on a low-power chip, and the voice recognition can be run in real time on a local device with limited resources, solving the problems of limited computing resources and low recognition robustness.
[0112] In an optional embodiment, an electronic device is provided, as Figure 11 shown Figure 11The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0113] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 4001 can also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0114] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0115] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0116] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0117] Among them, the electronic device package may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The shown electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0118] The embodiments of the present application provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it can implement the steps and corresponding content of the foregoing method embodiments. Compared with the prior art, it can achieve: on the basis of maintaining the global attention mechanism of the feature extraction model, the number of model parameters is reduced, the computational complexity is reduced, so that acoustic feature extraction can be performed in real time on a low-power chip, and speech recognition can be run in real time on a resource-constrained local device, solving the problems of limited computing resources and low recognition robustness.
[0119] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0120] The embodiments of the present application also provide a computer program product, including a computer program, which can implement the steps and corresponding content of the foregoing method embodiments when executed by a processor. Compared with the prior art, it can achieve: while maintaining the global attention mechanism of the feature extraction model, the number of model parameters is reduced, the computational complexity is lowered, so that acoustic feature extraction can be performed in real time on a low-power chip, and speech recognition can be run in real time on a resource-constrained local device, solving the problems of limited computing resources and low recognition robustness.
[0121] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in words.
[0122] It should be understood that although the flowcharts of the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0123] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. A voice recognition method, characterized in that, Applied to a terminal, the method includes: In response to receiving a voice collection instruction, obtain target voice data; Input the target voice data into a locally deployed feature extraction model to obtain acoustic features output by the feature extraction model; the acoustic features are used to obtain the target text of the target voice data; Wherein, the feature extraction model includes a convolutional layer, a plurality of first modules, and a feedforward neural network, the number of the first modules is less than a first preset threshold, and the number of neurons in the hidden layer of the feedforward neural network is less than a second preset threshold; The step of inputting the target voice data into the locally deployed feature extraction model to obtain the acoustic features output by the feature extraction model includes: Input the target voice data into the convolutional layer for convolutional processing to obtain a first feature of the target voice data output by the convolutional layer; Input the first feature into the plurality of first modules respectively. Each first module determines the correlation between each feature value in the first feature based on its own parameters. For each feature value in the first feature, use the correlation between the feature value and each feature value to perform fusion on each feature value corresponding to each feature value in the first feature to obtain a fusion feature value corresponding to the feature value; Input the second feature into the feedforward neural network for non-linear representation to obtain a third feature output by the feedforward neural network, and obtain the acoustic features according to the third feature; wherein, each feature value in the second feature is the fusion feature value corresponding to each feature value in the first feature.
2. The method according to claim 1, characterized in that, The feature extraction model is a sub-model of a locally deployed speech recognition model, and the speech recognition model further includes an acoustic decoding model and a language model as sub-models; After obtaining the acoustic features output by the feature extraction model, it further includes: Input the acoustic features into the acoustic decoding model to obtain a phoneme probability distribution of the target voice data output by the acoustic decoding model; Input the phoneme probability distribution into the language model to obtain the target text of the target voice data output by the language model.
3. The method according to claim 2, wherein The speech recognition model is trained in the following manner: Construct a teacher model and a student model; both the teacher model and the student model include a feature extraction model, an acoustic decoding model, and a language model to be trained, and the model parameters of the teacher model are more than those of the student model; Use the sample voice data as a training sample, and use the sample text corresponding to the sample voice data as a training label to train the teacher model to obtain a trained teacher model; Use the sample voice data as a training sample, use the sample text corresponding to the sample voice data as a training label, and use the text probability distribution output by the trained teacher model based on the training sample as a soft label to train the student model to obtain a trained student model; the text probability distribution is used to represent the probability distribution of each character corresponding to each moment in the training sample; Perform pruning processing on the trained student model, and use the pruned student model as the speech recognition model.
4. The method according to claim 3, wherein Pruning the trained student model and using the pruned student model as the speech recognition model includes: Determining the importance of each neuron in the trained student model, and pruning the neurons with importance less than a preset importance threshold; Using the sample speech data as the training sample and the corresponding sample text of the sample speech data as the training label to train the pruned student model to obtain a first model; Performing quantization processing on the first model to obtain a second model; Evaluating the calculation accuracy of the second model. If the evaluation result meets the preset conditions, using the second model as the speech recognition model; If the evaluation result does not meet the preset conditions, using the sample speech data as the training sample and the corresponding sample text of the sample speech data as the training label for training until the evaluation result meets the preset conditions.
5. The method according to claim 2, wherein The phoneme probability distribution is used to represent the probability of each phoneme at each moment in the target speech data; Inputting the phoneme probability distribution into the language model to obtain the target text of the target speech data output by the language model, including: Constructing a prefix tree based on the phoneme probability distribution; the i-th level of the prefix tree corresponds to the i-th moment in the target speech data, i ∈ N, N is the number of moments in the target speech data, and each node in each level of the prefix tree except the root node corresponds to a phoneme, and each node records the probability value of the corresponding phoneme; Determining the candidate nodes of each level layer by layer starting from the first level, and taking the connection between the candidate nodes of all adjacent levels as the final path, wherein the candidate nodes of each level are determined by the following method: For the first level, sorting the probability values included in each node in the first level, and taking the preset number of nodes with the largest probability values as the candidate nodes of the first level; For non-first levels, multiplying the probability results corresponding to each candidate path from the root node to the previous level by the probability values included in each node in the current level, and taking the preset number of nodes with the largest product as the candidate nodes of the current level; wherein each candidate path from the root node to the previous level includes one candidate node in each level from the first level to the previous level, and the probability value corresponding to each candidate path is the product of the probability values included in the corresponding candidate nodes.
6. The method according to claim 5, wherein Determining the target text according to the phoneme sequence, including: Determining the character sequence corresponding to the phoneme sequence according to a preset dictionary; the dictionary includes each phoneme and a character having a mapping relationship with each phoneme; Determining the context information of the target speech data, and determining the probability value of the character sequence expressing the correct semantics of the target speech data according to the context information; If the probability value is greater than a preset correct threshold, using the candidate word sequence as the target text of the target speech data; If the probability value is not greater than the preset correct threshold, determining the target text based on the context information, the phoneme sequence, and the dictionary.
7. The method according to claim 1, wherein Obtaining the target speech data includes: Obtaining speech data, and performing noise reduction on the speech data to obtain the noise-reduced speech data; Performing voice enhancement on the noise-reduced voice data through a pre-trained generative adversarial network to obtain voice data after voice enhancement; Performing endpoint detection on the voice data after voice enhancement, removing non-voice data and silent data in the voice data after voice enhancement to obtain valid voice data; Performing frame segmentation and windowing on the valid voice data to obtain target voice data.
8. A voice recognition device, characterized in that, Including: An acquisition module, configured to acquire target voice data in response to receiving a voice acquisition instruction; An input module, configured to input the target voice data into a locally deployed feature extraction model to obtain acoustic features output by the feature extraction model; the acoustic features are used to obtain the target text of the target voice data; Wherein, the feature extraction model includes a convolutional layer, a plurality of first modules, and a feedforward neural network, the number of the first modules is less than a first preset threshold, and the number of neurons in the hidden layer of the feedforward neural network is less than a second preset threshold; The inputting the target voice data into the locally deployed feature extraction model to obtain the acoustic features output by the feature extraction model includes: Inputting the target voice data into the convolutional layer for convolutional processing to obtain a first feature of the target voice data output by the convolutional layer; Inputting the first feature into the plurality of first modules respectively, each first module determines the correlation between each eigenvalue in the first feature based on its own parameters, and for each eigenvalue in the first feature, using the correlation between the eigenvalue and each eigenvalue, performing fusion on each eigenvalue corresponding to each eigenvalue in the first feature to obtain a fusion eigenvalue corresponding to the eigenvalue; Inputting a second feature into the feedforward neural network for non-linear representation to obtain a third feature output by the feedforward neural network, and obtaining the acoustic features according to the third feature; wherein, each eigenvalue in the second feature is a fusion eigenvalue corresponding to each eigenvalue in the first feature.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Speech recognition method and device, electronic equipment and computer readable storage medium
CN121393430A