Speech recognition method and device, electronic equipment and storage medium

By classifying the application scenarios of the speech recognition model and selecting matching feature parameters for model training, the existing speech recognition model has solved the problem of large amount of calculation and low recognition accuracy, and achieved more efficient and accurate speech recognition.

CN120071902APending Publication Date: 2025-05-30BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510269549.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing speech recognition model has a large amount of calculation and is difficult to effectively adapt to diverse application scenarios, resulting in low recognition accuracy and efficiency.

Method used

By classifying different application scenarios, selecting the second feature parameters matching the application scenario from a plurality of first feature parameters based on classification information, training the speech recognition model, reducing the amount of calculation during model training and avoiding interference from irrelevant feature parameters.

Benefits of technology

This reduces the amount of computation of the speech recognition model, improves the accuracy and efficiency of the recognition results, and allows the model to learn and utilize features that are valuable to the application scenarios more concentratedly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071902A_ABST
    Figure CN120071902A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method and device, electronic equipment and a storage medium, and belongs to the technical field of signal processing, and the method comprises the steps: selecting a second feature parameter from a plurality of first feature parameters based on the classification information of a target application scene; the first characteristic parameter is obtained based on a voice signal training sample; training a speech recognition model based on the second feature parameters; and recognizing a voice signal of a target application scene based on the voice recognition model. According to the speech recognition method and device, the electronic equipment and the storage medium provided by the invention, the calculation amount of the speech recognition model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of signal processing, and more specifically, relates to a speech recognition method and apparatus, an electronic device, and a storage medium. Background Art

[0002] Speech recognition technology makes human-computer interaction more natural and convenient. People can operate mobile phones, smart home devices, etc. through voice commands. For example, users only need to say "turn on the living room light" to control the smart home system to turn on the corresponding lights, without manual operation, greatly improving the interaction efficiency.

[0003] Speech recognition technology usually realizes by training a speech recognition model. In order to better adapt to diverse application scenarios, the speech recognition model needs to cover various types of speech feature extraction and processing mechanisms, which undoubtedly increases the complexity and computational amount of the model. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a speech recognition method and apparatus, an electronic device, and a storage medium to reduce the computational amount of the speech recognition model.

[0005] In the first aspect of the embodiments of the present disclosure, a speech recognition method is provided, including: Selecting second feature parameters from multiple first feature parameters based on the classification information of the target application scenario; the first feature parameters are obtained based on speech signal training samples; Training a speech recognition model based on the second feature parameters; Recognizing the speech signal of the target application scenario based on the speech recognition model.

[0006] In the second aspect of the embodiments of the present disclosure, a speech recognition apparatus is provided, including: A parameter screening module, configured to select second feature parameters from multiple first feature parameters based on the classification information of the target application scenario; the first feature parameters are obtained based on speech signal training samples; A model training module, configured to train a speech recognition model based on the second feature parameters; A speech recognition module, configured to recognize the speech signal of the target application scenario based on the speech recognition model.

[0007] In the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above speech recognition method are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the above-mentioned speech recognition method.

[0009] The beneficial effects of the speech recognition method, apparatus, electronic device, and storage medium provided by the embodiments of the present disclosure are as follows: The embodiments of the present disclosure consider that although different application scenarios have different focuses on speech recognition, there are also some common requirements. Therefore, different application scenarios are classified, and second feature parameters matching the application scenario are selected from multiple first feature parameters based on the classification information. Training the speech recognition model based on the second feature parameters can reduce the computational amount during the model training process; at the same time, it avoids the interference of irrelevant feature parameters to the model, enabling the model to concentrate on learning and utilizing the features valuable for the application scenario, thereby more accurately processing the speech signals in this scenario and improving the accuracy of the speech recognition results.

[0010] The embodiments of the present disclosure select the second feature parameters based on the classification information of the target application scenario and train the speech recognition model based on the second feature parameters, which can enable the speech recognition model to balance the requirements of computational amount and generality, and further reduce the computational amount of the speech recognition model and improve the recognition accuracy of the speech recognition model on the basis of meeting generality. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0012] Figure 1 It is a schematic flowchart of the speech recognition method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of the speech recognition apparatus provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of the electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0014] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.

[0015] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a speech recognition method provided by an embodiment of the present disclosure. The method includes: S101: Select a second feature parameter from multiple first feature parameters based on the classification information of the target application scenario; the first feature parameters are obtained based on speech signal training samples.

[0016] In this embodiment, the application scenarios may include intelligent companion robots, the field of security monitoring, meeting recording systems, intelligent Q&A systems, etc. Speech signal training samples under each application scenario can be collected, and the speech signal training samples are processed through specific feature extraction methods (such as extracting Mel Frequency Cepstral Coefficients MFCC, Linear Prediction Coefficients LPC, etc.) to obtain multiple first feature parameters. The first feature parameters characterize information such as the acoustic characteristics and vocal characteristics of the speech signal from different angles.

[0017] Different application scenarios have different focuses on speech recognition. For example, an intelligent companion robot needs to provide corresponding companion services according to the user's emotional state. When the user is in a low mood, the robot can play soothing music or tell jokes to improve the user's mood. By recognizing the speech emotion features, the robot can interact with the user more humanely and enhance the user experience.

[0018] In the field of security monitoring, for some places that require strict identity verification, such as bank vaults, high-security-level laboratories, etc., the speech recognition system can combine the speaker characteristics to determine whether the speaker is an authorized person. Only the speech of authorized staff can open the security door or access sensitive devices, thereby improving security.

[0019] In a meeting scenario, there may be multiple speakers participating in the discussion. Recognizing the speaker characteristics can help the system better classify and organize the meeting content. For example, the system can record the speeches of different speakers separately according to their identities, which is convenient for subsequent review and collation of meeting minutes, and improves the accuracy and efficiency of meeting records.

[0020] In an intelligent Q&A system, when a user asks a question to the intelligent Q&A system, the system needs to better understand the speech of different users in order to provide accurate answers.

[0021] Considering that there are some common requirements in emotional feature recognition, speaker identity feature recognition, and speech content recognition for each application scenario, this embodiment classifies each application scenario accordingly. For example, intelligent companion robots, intelligent customer service, psychological counseling, etc. can be classified as the first type of application scenario. In this type of application scenario, it is necessary to focus on emotion-related feature parameters, including intonation feature parameters, speech rate feature parameters, and volume feature parameters.

[0022] For the intonation feature, when the user speaks in a higher tone, it may indicate emotions such as excitement and doubt. These intonation information can be used as important emotional cues. Specifically, the intonation feature can be quantified by extracting parameters such as the fundamental frequency contour and its change rate.

[0023] For the speech rate feature, a faster speech rate may indicate anxiety or excitement, while a slower speech rate may be related to emotions such as contemplation and sadness. Specifically, the speech rate feature can be quantified by calculating the number of syllables per unit time.

[0024] For the volume feature, a larger volume may be related to emotions such as anger and excitement, while a smaller volume may indicate emotions such as shyness and frustration. Specifically, the volume feature can be extracted through the amplitude of the speech signal.

[0025] Correspondingly, two types of application scenarios, namely the security monitoring field and the meeting recording system, can be classified as the second type of application scenario. In this type of application scenario, it is necessary to focus on the voiceprint feature parameters related to user identity.

[0026] Among them, the voiceprint features include the fundamental frequency, formants (peaks in the speech spectrum), harmonic structure, etc. Each person's voiceprint is unique like a fingerprint. By extracting these voiceprint features, it can be used for speaker identity recognition. For example, the frequency position and bandwidth of the formants vary among different speakers, and these differences can be used as important bases for distinguishing speakers.

[0027] Correspondingly, application scenarios such as intelligent question-and-answer systems and speech translation can be classified as the third type of application scenario. In this type of application scenario, by analyzing information such as formants in the vocal tract features, it can help determine the vowels and consonants contained in the speech signal, thereby identifying the specific speech content.

[0028] The mapping relationship between various application scenario classifications and their corresponding second feature parameters can be constructed in advance. For a specific target application scenario, the corresponding second feature parameters can be obtained by looking up the above mapping relationship.

[0029] S102: Train a speech recognition model based on the second feature parameters.

[0030] In this embodiment, a speech recognition model can be trained based on a deep neural network - hidden Markov model (DNN - HMM), a convolutional neural network (CNN), a recurrent neural network (RNN), etc. Taking the second feature parameter as input data, during the training process, in conjunction with the corresponding annotation data (such as the text transcription corresponding to the speech, the speaker identity annotation, etc.), the parameters of the model are adjusted through an optimization algorithm (such as stochastic gradient descent, Adam, etc.) so that the model learns the mapping relationship between the second feature parameter and the expected output.

[0031] S103: Recognize the speech signal of the target application scenario based on the speech recognition model.

[0032] In this embodiment, when the actual speech signal in the target application scenario is received, the actual speech signal can be input into the trained speech recognition model. The speech recognition model first processes it according to the above steps S101 - S102, extracts the corresponding second feature parameter, and then performs recognition based on the second feature parameter to output the recognition result for the speech signal.

[0033] For example, in an intelligent question - answering system, what the model finally outputs is the text content after converting the speech, so that the system can further perform semantic understanding and dialogue response based on the text.

[0034] It can be concluded from the above that in this embodiment, considering that although different application scenarios have different emphases on speech recognition, there are also some common requirements. Therefore, different application scenarios are classified, and the second feature parameter matching the application scenario is selected from multiple first feature parameters based on the classification information. Training the speech recognition model based on the second feature parameter can reduce the computational amount during the model training process; at the same time, it avoids the interference of irrelevant feature parameters on the model, enabling the model to concentrate on learning and utilizing the features valuable for the application scenario, thereby more accurately processing the speech signal in this scenario and improving the accuracy of the speech recognition result.

[0035] In this embodiment, by selecting the second feature parameter based on the classification information of the target application scenario and training the speech recognition model based on the second feature parameter, the speech recognition model can balance the requirements of computational amount and generality. On the basis of meeting generality, it further reduces the computational amount of the speech recognition model and improves the recognition accuracy of the speech recognition model.

[0036] In an embodiment of the present disclosure, the training samples of the target application scenario include the first feature parameter and the corresponding output parameter. Selecting the second feature parameter from multiple first feature parameters based on the classification information of the target application scenario includes: Calculating the correlation between each of the multiple first feature parameters and the output parameter respectively; Select the second feature parameter from multiple first feature parameters based on the calculation results of the correlation.

[0037] In this embodiment, the Pearson correlation coefficient or the Spearman rank correlation coefficient can be used to calculate the correlation between multiple first feature parameters and the output parameter (such as the text content corresponding to the voice signal). Specifically, by collecting a large amount of training sample data, the first feature parameter values and the corresponding output parameter values in each sample are substituted into the correlation coefficient calculation formula for calculation, and the correlation numerical values between each first feature parameter and the output parameter are obtained.

[0038] If there is a strong correlation between a certain first feature parameter and the output parameter, its correlation coefficient will approach +1 or -1 (positive correlation or negative correlation), otherwise, the correlation coefficient is close to 0.

[0039] According to the calculation results of the correlation, the first feature parameter with the corresponding correlation numerical value greater than the first threshold can be selected as the second feature parameter. For example, it can be set to select the feature parameter with the absolute value of the correlation numerical value greater than 0.5 as the second feature parameter, so as to realize the selection of the second feature parameter from multiple first feature parameters. Among them, the first threshold is a preset constant, and those skilled in the art can flexibly design the specific value of the first threshold according to actual needs.

[0040] It can be concluded from the above that in this embodiment, by screening the second feature parameter with high correlation with the output parameter, the redundant features that contribute little to the output result or even may cause interference can be removed, making the data used for training the speech recognition model more concise and targeted, which helps to improve the efficiency and effect of model training, avoid the model being misled by irrelevant features during the learning process, and further improve the performance of the model in the target application scenario.

[0041] In an embodiment of the present disclosure, the speech recognition model includes a first acoustic model, a language model, and a first decoder. Recognizing the speech signal of the target application scenario based on the speech recognition model includes: Extract the acoustic feature vector from the speech signal based on the first acoustic model; Fuse the acoustic feature vector and the word embedding vector output by the language model in the hidden layer of the first decoder to obtain the speech recognition result of the target application scenario.

[0042] In this embodiment, considering that in the first type of application scenario (or the third type of application scenario), when the user asks a question to the intelligent companion robot (or the intelligent question answering system), the intelligent companion robot needs to understand the intention of the question based on the previous conversation content, so as to provide accurate reply information. If the context is not considered, the user's question may be misunderstood, resulting in inaccurate answers.

[0043] Therefore, for the first type of application scenario (or the third type of application scenario), it can be set that in addition to the necessary first acoustic model and the first decoder, the speech recognition model further includes a language model to assist the first acoustic model in more accurately determining a word sequence that conforms to language habits and semantic logic, thereby improving the accuracy of speech recognition.

[0044] The first acoustic model can be constructed using a neural network architecture, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), etc. By inputting the speech signal into the first acoustic model, the second feature parameters corresponding to the target application scenario can be obtained, and the acoustic feature vector is obtained by forming the second feature parameters into a vector form.

[0045] The language model can learn language rules such as grammar, vocabulary collocations, and semantic knowledge. It is trained based on a large amount of text corpora and can predict the probability of the occurrence of word sequences. The word embedding vector is a low-dimensional and distributed representation method of words by the language model, which maps words into a vector space so that words with similar semantics are closer in this vector space. For example, the distance between "apple" and "banana" in the word embedding vector space may be closer than that between "apple" and "car" because they belong to the same semantic category (both are fruits). The language model can output the corresponding word embedding vector according to the input context and other information, and the word embedding vector contains rich semantic information.

[0046] The acoustic feature vector and the word embedding vector output by the language model are fused in the hidden layer of the first decoder. Based on the fused vector information, the decoder uses a certain search algorithm (such as the Viterbi algorithm, beam search algorithm, etc.) to find the most likely word sequence as the result of speech recognition.

[0047] From the above, it can be concluded that in this embodiment, by fusing the acoustic feature vector and the word embedding vector output by the language model in the hidden layer of the first decoder, the output speech recognition result takes into account both the actual pronunciation of the speech and the rationality of the language, which is beneficial to improving the accuracy of the speech recognition result.

[0048] In an embodiment of the present disclosure, the speech recognition method further includes: fusing the acoustic feature vector and the word embedding vector output by the language model in the hidden layer of the first decoder, including: At each time step of the first decoder, calculate the first attention weight of the acoustic feature vector and the second attention weight of the word embedding vector; Perform weighted summation on the acoustic feature vector based on the first attention weight to obtain the fused acoustic feature vector; Weighted sum of word embedding vectors is performed based on the second attention weight to obtain fused word embedding vectors; Concatenate the fused acoustic feature vectors and the fused word embedding vectors to obtain fused vectors.

[0049] In this embodiment, if the acoustic feature vectors and word embedding vectors are simply concatenated and input to the decoder, redundant information may be introduced or the information may conflict with each other, making it impossible for the decoder to fully utilize this information.

[0050] To solve the above problems, in this embodiment, after weighted fusion of the acoustic feature vectors and word embedding vectors based on the attention mechanism, concatenation is performed. Irrelevant acoustic features and semantic information can be filtered according to the specific requirements of each time step, so that the final fused vectors contain important information at both the acoustic level and the language semantic level.

[0051] Specifically, at each time step of the first decoder, calculate the first attention weight and the second attention weight, so as to assign different weights to the acoustic feature vectors and word embedding vectors respectively, enabling the model to focus on the part of the input data that is most relevant to the current task.

[0052] Specifically, the first attention weight of the acoustic feature vectors can be calculated through the first formula, and the second attention weight of the word embedding vectors can be calculated through the second formula.

[0053] The first formula is:

[0054] Among them, represents the first attention weight, represents the acoustic feature vector corresponding to the i-th time step, represents the hidden state vector of the first decoder at time step t - 1, which contains the context information in the process of generating words before, represents the transposed vector of, represents the randomly initialized weight matrix.

[0055] In the above first formula, the first attention weight of the acoustic feature vectors is obtained through the dot product operation combined with the non-linear activation function softmax.

[0056] The second formula is:

[0057] Among them, represents the second attention weight, represents the word embedding vector corresponding to the j-th word, Denote the hidden state vector of the first decoder at time step t-1, which contains the context information in the process of generating previous words. Denote the transposed vector of Denote the randomly initialized weight matrix.

[0058] In the above second formula, the second attention weight of the word embedding vector is obtained by combining the dot product operation with the non-linear activation function softmax.

[0059] Then, based on the obtained first attention weight, the acoustic feature vectors at M time steps can be weighted and summed to obtain the fused acoustic feature vector. The fusion process can adopt the following third formula:

[0060] Where Denote the fused acoustic feature vector, , Denote the number of frames of the speech signal, Denote the frame length of the speech signal, Denote the length of the time step.

[0061] In the above third formula, each acoustic feature vector is multiplied by its corresponding first attention weight and then summed, so that the acoustic features at different time steps can be fused according to the attention weights, highlighting the acoustic feature part related to the word to be generated currently.

[0062] Based on the obtained second attention weight, the word embedding vectors can be weighted and summed to obtain the fused word embedding vector. The fusion process can adopt the following fourth formula:

[0063] Where Denote the fused word embedding vector, and N denotes the length of the word sequence.

[0064] In the above fourth formula, each word embedding vector is multiplied by its corresponding second attention weight and then summed, which can focus on the semantic and syntactic information related to the word to be generated currently.

[0065] Finally, the fused acoustic feature vector and the fused word embedding vector are concatenated to obtain the fused vector.

[0066] It can be concluded from the above that in this embodiment, by dynamically calculating the attention weights at each time step and performing weighted summation, the first decoder can focus on the parts of the acoustic feature vector and the word embedding vector that are most relevant to the word to be recognized currently, making the speech recognition result more conform to the actual speech content and language norms.

[0067] In one embodiment of the present disclosure, the speech recognition model includes a second acoustic model, an identity recognition model, and a second decoder. Recognizing the speech signal of the target application scenario based on the speech recognition model includes: Input the speech signal into the identity recognition model to obtain the identity recognition result of the target user; Based on the identity recognition result of the target user and the second acoustic model, obtain the first phoneme probability distribution; Input the first phoneme probability distribution into the second decoder to obtain the recognition result of the speech signal.

[0068] In this embodiment, considering that in the second type of application scenario, the user identity is crucial for the speech recognition result. Therefore, in the second type of application scenario, it can be set that in addition to the necessary second acoustic model and second decoder, the speech recognition model further includes an identity recognition model to assist the second acoustic model to more accurately output the first phoneme probability distribution, thereby improving the accuracy of speech recognition.

[0069] Specifically, the identity recognition model mainly distinguishes different users based on the unique voiceprint features contained in the speech signal. Common voiceprint features include fundamental frequency, formant, harmonic structure, and pronunciation habits, etc. For example, due to different physiological factors such as the vocal cord structure and oral cavity shape of each person, there are differences in features such as fundamental frequency and formant during pronunciation. At the same time, different language habits and regional accents also constitute unique pronunciation habit features. The identity recognition model can extract these features and compare and analyze them with the voiceprint feature templates of known users stored in advance to determine whether the speaker corresponding to the current speech signal matches a certain known user, thereby obtaining the identity recognition result of the target user.

[0070] The second acoustic model can extract acoustic features from the input speech signal, learn the mapping relationship between the above features and phonemes, calculate the probability of each possible phoneme appearing in the current speech signal, and finally form a phoneme probability distribution. For example, for a piece of speech, the second acoustic model may output that the probability of the phoneme "a" is 0.2, the probability of the phoneme "o" is 0.15, etc., to represent the likelihood of each phoneme appearing. Since the phoneme is the smallest unit of speech, the speech recognition result can be determined based on the phoneme probability distribution.

[0071] Considering that different users may have different accents, speaking speeds, word - using habits, etc., which lead to changes in phoneme pronunciation and occurrence probabilities. To solve the above problems, in this embodiment, prior knowledge is provided for the second acoustic model based on the identity recognition result of the target user, so as to more accurately identify the phoneme combinations with accents. For example, for users with a strong accent, after the identity recognition model determines the dialect area to which the user belongs, the second acoustic model can better identify the phonemes and phoneme combinations unique to the dialect accordingly, thereby improving the accuracy of the first phoneme probability distribution.

[0072] After receiving the first phoneme probability distribution output from the second acoustic model, the second decoder can use a specific search algorithm (such as the Viterbi algorithm, beam search algorithm, etc.) to find the most likely phoneme sequence, and then generate the speech recognition result.

[0073] It can be concluded from the above that this embodiment takes into account the individual differences of different users in pronunciation habits, accents, speaking speeds, etc. By first determining the user identity through the identity recognition model, the second acoustic model can make targeted adjustments accordingly, effectively reducing the recognition errors caused by individual differences and improving the accuracy of speech recognition.

[0074] In an embodiment of the present disclosure, based on the identity recognition result of the target user and the second acoustic model, obtaining the first phoneme probability distribution includes: Input the speech signal into the identity recognition model to obtain the identity recognition result of the target user; Determine the region to which the target user belongs based on the identity recognition result of the target user; Adjust the second phoneme probability distribution output by the second acoustic model based on the region to which the target user belongs to obtain the first phoneme probability distribution.

[0075] In this embodiment, the identity information of each user includes which province, city or dialect area they are from, etc. When the specific identity of the target user is determined through the identity recognition model, the corresponding region information of the user can be directly determined according to the identity information.

[0076] On this basis, the second phoneme probability distribution output by the second acoustic model can be adjusted according to the speech characteristics of the region. The adjustment process may include: Compare each phoneme in the second phoneme probability distribution with the phoneme set of the target region to determine the target phoneme; the target region is the region to which the target user belongs; Increase the probability corresponding to the target phoneme by the first step length to obtain the first phoneme probability distribution.

[0077] In this embodiment, a phoneme set of common dialect regions can be pre-constructed. For example, Cantonese has a complete phoneme system, which contains some phonemes that do not exist in Mandarin, such as the "ng" (velar nasal) initial in Cantonese, and the entering tone phonemes, which do not exist in the phoneme set of Mandarin.

[0078] If the region to which the target user belongs is a certain dialect region, each phoneme in the second phoneme probability distribution can be compared with the phoneme set of this dialect region. The identical phonemes are the phonemes belonging to the dialect region among the phonemes in the second phoneme probability distribution, that is, the target phonemes. On the basis of the probability corresponding to the target phonemes, a first step length is increased to improve the probability of the target phonemes, so as to achieve accurate recognition of the dialect region speech signal. Among them, the first step length is a preset constant, such as 0.1, 0.2, etc.

[0079] It can be concluded from the above that after adjusting the output of the second acoustic model according to the region to which the target user belongs in this embodiment, the acoustic model can conform to the pronunciation habits of these regions to identify phonemes, avoid recognition errors caused by pronunciation habit differences, and improve the recognition accuracy of the voices of users in different regions.

[0080] In an embodiment of the present disclosure, the speech recognition model includes a second acoustic model, an identity recognition model, and a second decoder. Recognizing the speech signal of the target application scenario based on the speech recognition model includes: Inputting the speech signal into the second acoustic model to obtain a second phoneme probability distribution; Inputting the speech signal into the identity recognition model to obtain an identity recognition result of the target user; Determining the region to which the target user belongs based on the identity recognition result of the target user; Determining the phoneme set corresponding to the region to which the target user belongs as the initial search range of the second decoder, and inputting the second phoneme probability distribution into the second decoder to obtain a recognition result of the speech signal.

[0081] In this embodiment, after determining the region to which the target user belongs, determining the phoneme set corresponding to the region to which the target user belongs as the initial search range of the second decoder can enable the second decoder to preferentially screen within this specific phoneme range related to the user region, reduce the search space, improve the search efficiency, and at the same time make the search more focused on the phoneme combinations that conform to the actual situation of the regional speech.

[0082] It can be concluded from the above that in this embodiment, determining the phoneme set corresponding to the region to which the target user belongs as the initial search range of the second decoder can enable the decoder to make full use of the regional speech characteristics, quickly and accurately screen out the phoneme sequence that conforms to the actual speech, reduce the calculation amount of searching within unnecessary phoneme ranges, and improve the speed and accuracy of speech recognition.

[0083] A voice recognition method corresponding to the above embodiments Figure 2 is a structural block diagram of a voice recognition device provided by an embodiment of the present disclosure. For ease of description, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 , the voice recognition device 20 includes: a parameter screening module 21, a model training module 22, and a voice recognition module 23. Among them, the parameter screening module 21 is used to select second feature parameters from multiple first feature parameters based on the classification information of the target application scenario; the first feature parameters are obtained based on voice signal training samples; The model training module 22 is used to train a voice recognition model based on the second feature parameters; The voice recognition module 23 is used to recognize the voice signal of the target application scenario based on the voice recognition model.

[0084] In an embodiment of the present disclosure, the parameter screening module 21 is specifically used for: Calculating the correlation between multiple first feature parameters and output parameters respectively; Selecting second feature parameters from multiple first feature parameters based on the calculation results of the correlation.

[0085] In an embodiment of the present disclosure, the voice recognition model includes a first acoustic model, a language model, and a first decoder. The voice recognition module 23 is specifically used for: Extracting acoustic feature vectors from the voice signal based on the first acoustic model; the acoustic feature vectors are vectors composed of second feature parameters; Fusing the acoustic feature vectors and the word embedding vectors output by the language model in the hidden layer of the first decoder to obtain the voice recognition result of the target application scenario.

[0086] In an embodiment of the present disclosure, the voice recognition module 23 is specifically further used for: Calculating the first attention weight of the acoustic feature vectors and the second attention weight of the word embedding vectors at each time step of the first decoder; Performing weighted summation on the acoustic feature vectors based on the first attention weight to obtain fused acoustic feature vectors; Performing weighted summation on the word embedding vectors based on the second attention weight to obtain fused word embedding vectors; Concatenating the fused acoustic feature vectors and the fused word embedding vectors to obtain a fused vector.

[0087] In an embodiment of the present disclosure, the voice recognition model includes a second acoustic model, an identity recognition model, and a second decoder. The voice recognition module 23 is specifically used for: Input the voice signal into the identity recognition model to obtain the identity recognition result of the target user; Based on the identity recognition result of the target user and the second acoustic model, obtain the first phoneme probability distribution; Input the first phoneme probability distribution into the second decoder to obtain the recognition result of the voice signal.

[0088] In an embodiment of the present disclosure, the voice recognition module 23 is further specifically configured to: Input the voice signal into the identity recognition model to obtain the identity recognition result of the target user; Determine the region to which the target user belongs based on the identity recognition result of the target user; Adjust the second phoneme probability distribution output by the second acoustic model based on the region to which the target user belongs to obtain the first phoneme probability distribution.

[0089] In an embodiment of the present disclosure, the voice recognition model includes a second acoustic model, an identity recognition model, and a second decoder. The voice recognition module 23 is specifically configured to: Input the voice signal into the second acoustic model to obtain the second phoneme probability distribution; Input the voice signal into the identity recognition model to obtain the identity recognition result of the target user; Determine the region to which the target user belongs based on the identity recognition result of the target user; Determine the phoneme set corresponding to the region to which the target user belongs as the initial search range of the second decoder, and input the second phoneme probability distribution into the second decoder to obtain the recognition result of the voice signal.

[0090] See Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 complete mutual communication through the communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the modules 21 to 23 shown.

[0091] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0093] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0094] In specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the voice recognition method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0095] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiment are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0096] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0097] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0098] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0099] In several embodiments provided by this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces or units, or can also be an electrical, mechanical or other form of connection.

[0100] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0101] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0102] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A speech recognition method, characterized in that: include: Selecting a second characteristic parameter from a plurality of first characteristic parameters based on classification information of a target application scenario; Said The first feature parameter is obtained based on a speech signal training sample; Training a speech recognition model based on the second feature parameter; The speech signal of the target application scenario is recognized based on the speech recognition model.

2. The speech recognition method according to claim 1, characterized in that: The training sample of the target application scenario includes a first characteristic parameter and a corresponding output parameter, and the selecting a second characteristic parameter from a plurality of first characteristic parameters based on the classification information of the target application scenario includes: respectively calculating correlations between a plurality of first characteristic parameters and the output parameter; A second characteristic parameter is selected from the plurality of first characteristic parameters based on the calculation result of the correlation.

3. The speech recognition method according to claim 1, wherein: The speech recognition model includes a first acoustic model, a language model and a first decoder, and the recognition of a speech signal of a target application scenario based on the speech recognition model includes: Extracting an acoustic feature vector from the speech signal based on the first acoustic model; the acoustic feature vector is a vector composed of the second feature parameters; The acoustic feature vector and the word embedding vector output by the language model are fused in the hidden layer of the first decoder to obtain a speech recognition result of the target application scenario.

4. The speech recognition method according to claim 3, characterized in that: The step of fusing the acoustic feature vector and the word embedding vector output by the language model in a hidden layer of the first decoder comprises: At each time step of the first decoder, calculating a first attention weight of the acoustic feature vector and a second attention weight of the word embedding vector; Performing weighted summation on the acoustic feature vectors based on the first attention weights to obtain a fused acoustic feature vector; Performing weighted summation on the word embedding vectors based on the second attention weight to obtain a fused word embedding vector; The fused acoustic feature vector and the fused word embedding vector are concatenated to obtain a fused vector.

5. The speech recognition method according to claim 1, wherein: The speech recognition model includes a second acoustic model, an identity recognition model and a second decoder, and the recognition of the speech signal of the target application scenario based on the speech recognition model includes: Inputting the voice signal into the identity recognition model to obtain an identity recognition result of the target user; Obtaining a first phoneme probability distribution based on the target user's identity recognition result and the second acoustic model; The first phoneme probability distribution is input into the second decoder to obtain a recognition result of a speech signal.

6. The speech recognition method according to claim 5, characterized in that: Obtaining a first phoneme probability distribution based on the target user's identity recognition result and the second acoustic model includes: Inputting the voice signal into the identity recognition model to obtain an identity recognition result of the target user; Determine the region to which the target user belongs based on the identity recognition result of the target user; The second phoneme probability distribution output by the second acoustic model is adjusted based on the region to which the target user belongs to obtain the first phoneme probability distribution.

7. The speech recognition method according to claim 1, characterized in that: The speech recognition model includes a second acoustic model, an identity recognition model and a second decoder, and the recognition of the speech signal of the target application scenario based on the speech recognition model includes: Inputting the speech signal into the second acoustic model to obtain a second phoneme probability distribution; Inputting the voice signal into the identity recognition model to obtain an identity recognition result of the target user; Determine the region to which the target user belongs based on the identity recognition result of the target user; The phoneme set corresponding to the area to which the target user belongs is determined as the initial search range of the second decoder, and the second phoneme probability distribution is input into the second decoder to obtain a recognition result of the speech signal.

8. A speech recognition device, characterized in that: include: A parameter screening module, used for selecting a second characteristic parameter from a plurality of first characteristic parameters based on classification information of a target application scenario; The first characteristic parameter is obtained based on a speech signal training sample; A model training module, used for training a speech recognition model based on the second feature parameter; The speech recognition module is used to recognize the speech signal of the target application scenario based on the speech recognition model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.