Speech Recognition Method, Apparatus, Device, Readable Storage Medium and Computer Program
The target user is determined through the voiceprint recognition model and the speech decoding is performed using a language model trained based on the target user's historical text. The problem of low speech recognition accuracy in the prior art is solved and a higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111459909.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-02
AI Technical Summary
The use of common language models in speech recognition in prior art leads to a low accuracy of recognition results.
The voiceprint recognition model is called to determine the voiceprint recognition model and the speech decoding is performed based on the language model trained on the target user's historical text data to improve the recognition accuracy.
Through targeted recognition, the accuracy of speech recognition is significantly improved.
Smart Images

Figure CN114171014B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a voice recognition method, apparatus, device, readable storage medium, and computer program. Background Art
[0002] With the rapid development of science and technology, more and more intelligent products are applied to people's lives, and most intelligent products need to implement interactive functions through voice recognition technology. The so-called voice recognition technology is a technology that uses a computer to convert voice into text.
[0003] In related technologies, when performing voice recognition, generally, a segment of voice data of a user is first obtained, and then a general language model is called to participate in voice decoding, so as to obtain the text converted from the segment of voice data.
[0004] However, the method of using a general language model to recognize the voice data of different users has a low accuracy of the recognition result. Summary of the Invention
[0005] In view of this, embodiments of the present application provide a voice recognition method, apparatus, device, readable storage medium, and computer program, which improve the accuracy of voice recognition.
[0006] On the one hand, embodiments of the present application provide a voice recognition method, which includes:
[0007] Obtain voice data;
[0008] Call a voiceprint recognition model to process the voice data and a voice feature set to determine a target user matching the voice data, where multiple users' historical voice features are stored in the voice feature set;
[0009] In the process of decoding the voice data, call a target language model matching the target user to process the voice data to obtain a target text corresponding to the voice data, where the target language model is trained based on the historical text data of the target user;
[0010] Output the target text corresponding to the voice data.
[0011] Optionally, calling a voiceprint recognition model to process the voice data and a voice feature set to determine a target user matching the voice data includes:
[0012] Extract features from the voice data to obtain voice features;
[0013] Invoke the voiceprint recognition model to process the voice feature and the historical voice features of the first user in the voice feature set to determine the likelihood between the voice feature and the historical voice features of the first user;
[0014] In response to the likelihood being greater than or equal to the likelihood threshold, determine the first user as the target user matching the voice data.
[0015] Optionally, the method further includes:
[0016] If no target user is matched from the voice feature set, add the voice data as the voice data of a new user to the historical voice data set.
[0017] Optionally, the method further includes:
[0018] When the number of new users in the historical voice data set is greater than the first number and the voice data volume of each new user is greater than the second number, perform feature extraction on multiple voice data of each new user to obtain multiple voice features of each new user;
[0019] Take the average of the multiple voice features of each new user to obtain the average feature of each new user, and add the average feature as the voice feature of each new user to the voice feature set.
[0020] Optionally, the training process of the voiceprint recognition model includes:
[0021] Obtain a sample voice feature set, which is obtained based on the historical voice data of multiple users;
[0022] Based on the sample voice feature set, train the voiceprint recognition model with the user source as the supervision.
[0023] Optionally, based on the sample voice feature set, training the voiceprint recognition model with the user source as the supervision includes:
[0024] In the i-th iteration process, obtain a voice feature pair from the sample voice feature set, the voice feature pair includes a first sample voice feature and a second sample voice feature, where i is a positive integer;
[0025] Invoke the voiceprint recognition model to process the first sample voice feature and the second voice sample feature to determine the likelihood between the first sample voice feature and the second sample voice feature;
[0026] If it is determined that the i-th iteration meets the training stop condition based on the likelihood and the sample label of the voice feature pair, determine the voiceprint recognition model as the trained voiceprint recognition model, and the sample label is used to indicate whether the sample voice features in the voice feature pair come from the same user;
[0027] If it is determined that the i-th iteration does not meet the training stop condition based on the likelihood and the sample label of the voice feature pair, the model parameters of the voiceprint recognition model are adjusted, and the (i + 1)-th iteration process is performed based on the adjusted voiceprint recognition model.
[0028] On the one hand, an embodiment of the present application provides a voice recognition device, which includes:
[0029] A first acquisition module, configured to acquire voice data;
[0030] A first determination module, configured to call a voiceprint recognition model to process the voice data and a voice feature set to determine a target user matching the voice data, where multiple historical voice features of multiple users are stored in the voice feature set;
[0031] A call module, configured to call a target language model matching the target user to process the voice data during the decoding of the voice data, so as to obtain a target text corresponding to the voice data, where the target language model is trained based on historical text data of the target user;
[0032] An output module, configured to output the target text corresponding to the voice data.
[0033] Optionally, the first determination module is configured to:
[0034] Extract features from the voice data to obtain voice features;
[0035] Call a voiceprint recognition model to process the voice features and the historical voice features of the first user in the voice feature set to determine the likelihood between the voice features and the historical voice features of the first user;
[0036] In response to the likelihood being greater than or equal to a likelihood threshold, determine the first user as the target user matching the voice data.
[0037] Optionally, the device further includes:
[0038] An addition module, configured to add the voice data as voice data of a new user to the historical voice data set if no target user is matched from the voice feature set.
[0039] Optionally, the device further includes an update module, and the update module is configured to:
[0040] When the number of new users in the historical voice data set is greater than a first number, and the voice data volume of each new user is greater than a second number, extract features from multiple voice data of each new user to obtain multiple voice features of each new user;
[0041] Take the average of multiple voice features of each of the newly added users to obtain the average feature of each newly added user, and add the average feature as the voice feature of each newly added user to the voice feature set.
[0042] Optionally, the device further includes:
[0043] A second acquisition module, configured to acquire a sample voice feature set, which is obtained based on the historical voice data of multiple users;
[0044] A training module, configured to train a voiceprint recognition model based on the sample voice feature set with the user source as the supervision.
[0045] Optionally, the training module is configured to:
[0046] In the i-th iteration process, obtain a voice feature pair from the sample voice feature set, where the voice feature pair includes a first sample voice feature and a second sample voice feature, and i is a positive integer;
[0047] Call the voiceprint recognition model to process the first sample voice feature and the second voice sample feature to determine the likelihood between the first sample voice feature and the second sample voice feature;
[0048] If it is determined that the i-th iteration meets the training stop condition based on the likelihood and the sample label of the voice feature pair, then determine the voiceprint recognition model as the trained voiceprint recognition model, where the sample label is used to indicate whether the sample voice features in the voice feature pair come from the same user;
[0049] If it is determined that the i-th iteration does not meet the training stop condition based on the likelihood and the sample label of the voice feature pair, then adjust the model parameters of the voiceprint recognition model, and perform the (i + 1)-th iteration process based on the adjusted voiceprint recognition model.
[0050] On the one hand, an embodiment of the present application provides a computer device, which includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the voice recognition method.
[0051] On the one hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program is stored, and the computer program is loaded and executed by a processor to implement the voice recognition method.
[0052] On the one hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the speech recognition method.
[0053] The technical solution provided by the embodiment of the present application calls a voiceprint recognition model to process voice data and a voice data set, and can determine a target user matching the voice data, that is, first determine the identity of the speaker to which the voice data belongs. Then, based on the determined target user, a target language model matching the target user is called to process the voice data to obtain the target text of the voice data, and finally the target text corresponding to the voice data is output. Since the target language model is trained based on the historical text data of the target user, the target language model can more accurately reflect the speaking style of the target user. That is to say, using the target voice model, targeted recognition of voice data can be performed. Therefore, by adopting this speech recognition method, the accuracy of speech recognition can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 is a schematic diagram of an implementation environment of a speech recognition method provided by an embodiment of the present application;
[0056] Figure 2 is a flowchart of a speech recognition method provided by an embodiment of the present application;
[0057] Figure 3 is a flowchart of a speech recognition method provided by an embodiment of the present application;
[0058] Figure 4 is a flowchart of training a voiceprint recognition model in a speech recognition method provided by an embodiment of the present application;
[0059] Figure 5 is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application;
[0060] Figure 6 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0062] Unless otherwise defined, all technical terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art. In the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited. In the present application, the term "at least one" means one or more, and the meaning of "a plurality" means two or more.
[0063] Figure 1 is a schematic diagram of the implementation environment of a speech recognition method provided by an embodiment of the present application. Refer to Figure 1 In this implementation environment, a terminal 110 and a server 120 may be included.
[0064] The terminal 110 is connected to the server 120 through a wireless network or a wired network. Optionally, the terminal 110 is a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, etc., but is not limited thereto. The terminal 110 installs and runs an application program that supports speech recognition.
[0065] The server 120 is an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0066] Those skilled in the art can know that the number of the above terminals may be more or less. For example, there is only one of the above terminals, or there are dozens or hundreds of the above terminals, or a larger number. In this case, other terminals are also included in the above implementation environment. The embodiments of the present application do not limit the number and device types of the terminals.
[0067] In a possible implementation manner, the process of speech recognition provided by the embodiments of the present application can be triggered by the terminal when speech recognition is required. The following uses a speech recognition scenario as an example to introduce this application scenario:
[0068] The terminal can display a voice recognition option on the application interface. When the user wants to perform voice recognition on a piece of voice data, the user can click on the voice recognition option to trigger the terminal to send a voice recognition request to the server. After receiving the voice recognition request, the server will, in response to the voice recognition request, execute the steps of voice recognition provided in the embodiments of the present application to generate a target text and send the target text to the terminal for the user to use.
[0069] To make the technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0070] In the embodiments of the present application, the technical solutions provided in the embodiments of the present application can be implemented with the server or the terminal as the execution subject, or the technical methods provided in the present application can be implemented through the interaction between the terminal and the server. The embodiments of the present application do not make any limitations in this regard.
[0071] Figure 2 is a flowchart of a voice recognition method provided in the embodiments of the present application. Taking the server as the execution subject of this method as an example, see Figure 2 This method includes steps 201 to 204.
[0072] 201. The server obtains voice data.
[0073] Among them, the server pre-stores multiple pieces of voice data locally. When the user needs to perform voice recognition on one of the pieces of voice data, the server directly obtains the piece of voice data from local storage. Alternatively, when the user needs to perform voice recognition on voice data, the user sends a voice recognition instruction to the server through the terminal. The server, in response to the voice recognition instruction, obtains the voice data corresponding to the voice recognition instruction in real time.
[0074] 202. The server calls a voiceprint recognition model to process the voice data and the voice feature set to determine the target user matching the voice data. Multiple users' historical voice features are stored in the voice feature set.
[0075] Among them, two speech features are input into the voiceprint recognition model. One of the speech features belongs to a known user, and the other belongs to an unknown user. The voiceprint recognition model is used to determine whether the speech feature of the unknown user belongs to the same user as the speech feature of the known user. The voiceprint recognition model is a Probabilistic Linear Discriminant Analysis (PLDA) model. By inputting the two speech features into the PLDA model, the likelihood between the two speech features is obtained. The greater the likelihood, the higher the probability that the two speech features belong to the same user. The speech feature set is used to store the historical speech features of multiple users, and the historical speech features of each user are different. When determining the target user, a historical speech feature is sequentially called from the speech feature set, and based on the PLDA model, the speech data, and the historical speech feature, the target user matching the speech data is determined. The target user refers to the user who uttered the speech data.
[0076] 203. During the process of decoding the speech data by the server, the target language model matching the target user is called to process the speech data to obtain the target text corresponding to the speech data. Among them, the target language model is trained based on the historical text data of the target user.
[0077] Among them, since the target language model is trained based on the historical text data of the target user, the target language model can best reflect the speaking style and speaking habits of the target user. Based on the target text obtained after processing by the target language model, the content that the target user wants to express can be more accurately reflected. There are various ways to obtain the historical text data of the target user. For example, after manually annotating multiple historical speech data of the target user, the historical text data corresponding to each historical speech data is obtained.
[0078] It should be noted that the process of decoding the speech data includes: the process of performing speech recognition on the speech data using an acoustic model, a pronunciation dictionary, and a language model. This process includes the following multiple steps: calling the acoustic model, inputting the speech data into the acoustic model to obtain multiple groups of phonemes corresponding to the speech data; based on the multiple groups of phonemes, determining multiple candidate characters or candidate words corresponding to the multiple groups of phonemes in the pronunciation dictionary; calling the language model, inputting the multiple candidate characters or candidate words into the language model, and obtaining and outputting the unique text corresponding to the speech data. Among them, the server calling the target language model matching the target user in step 203 is the step of calling the language model during the speech recognition process.
[0079] 204. The server outputs the target text corresponding to the speech data.
[0080] The server outputs the target text corresponding to the voice data to the terminal. For example, the server outputs the target text corresponding to the voice data to the terminal. After receiving the target text, the terminal displays the target text on the interface of the application program, where the application program is a program installed on the terminal for providing voice recognition services; or, the server outputs the target text corresponding to the voice data to the terminal. After receiving the target text, the terminal converts the target text into audio, generates an instruction based on the audio, and then realizes the interaction between the terminal and other electronic devices based on the instruction.
[0081] In the technical solution provided by the embodiment of the present application, by invoking the voiceprint recognition model to process the voice data and the voice data set, the target user matching the voice data can be determined, that is, the identity of the speaker to which the voice data belongs is first determined. Then, based on the determined target user, the target language model matching the target user is invoked to process the voice data to obtain the target text of the voice data, and finally the target text corresponding to the voice data is output. Since the target language model is trained based on the historical text data of the target user, the target language model can more accurately reflect the speaking style of the target user. That is to say, by using the target voice model, the voice data can be specifically recognized. Therefore, by adopting this voice recognition method, the accuracy of voice recognition can be improved.
[0082] Figure 3 It is a flowchart of a voice recognition method provided by an embodiment of the present application. This embodiment is described by taking the execution entity as the server as an example. Refer to Figure 3 and this method includes steps 301 to 306.
[0083] 301. The server obtains voice data.
[0084] This step 301 is the same as step 201 above and will not be elaborated here.
[0085] 302. The server extracts features from the voice data to obtain voice features.
[0086] In some embodiments, the server invokes a Deep Neural Networks (DNN) model to extract features from the voice data to obtain the voice features corresponding to the voice data.
[0087] It should be noted that the voice features can also be obtained by invoking other models to process the voice data, which will not be elaborated here.
[0088] 303. The server invokes a voiceprint recognition model to process the voice features and the historical voice features of the first user in the voice feature set to determine the likelihood between the voice features and the historical voice features of the first user.
[0089] In some embodiments, the voice feature set stores not only a plurality of historical voice features, but also user information corresponding to each historical voice feature. For example, in the voice feature set, a user identity identifier is assigned to each historical voice feature. Among them, the historical voice feature of the first user is any user's historical voice feature in the voice feature set.
[0090] In some embodiments, the voice features of two voice data are input into the PLDA model, and the PLDA model outputs the likelihood between the two voice features to determine whether the two voice features belong to the same space. If they belong to the same space, it means that the two voice data belong to the same speaker. If they do not belong to the same space, it means that the two voice data do not belong to the same speaker. At the same time, the greater the likelihood, the greater the possibility that the two voice data belong to the same speaker.
[0091] In some embodiments, the likelihood is obtained by determining the log-likelihood ratio between two voice features. The formula for determining the log-likelihood ratio is as follows:
[0092]
[0093] Where η 1 , η 2 represent two voice features respectively, p(η 1 |H d ) and p(η 2 |H d ) represent the likelihood functions that the two voice features come from different spaces respectively, p(η 1 , η 2 |H s ) represents the likelihood function that the two voice features come from the same space, and score represents the log-likelihood ratio.
[0094] In some embodiments, the process for the server to call the voiceprint recognition model to process the voice feature and the historical voice feature of the first user in the voice feature set to determine the likelihood between the voice feature and the historical voice feature of the first user is as follows:
[0095] The server calls the voiceprint recognition model and sequentially inputs the voice feature and each historical voice feature of the first user into the voiceprint recognition model. Based on the above formula for solving the log-likelihood ratio, the log-likelihood ratio between the voice feature and each historical voice feature of the first user is obtained respectively.
[0096] 304. The server determines the first user as the target user matching the voice data in response to the likelihood being greater than or equal to the likelihood threshold.
[0097] In some embodiments, a likelihood threshold is preset. If the log-likelihood ratio determined based on the voice feature and the voice feature of the first user is greater than or equal to the likelihood threshold, it is considered that the first user and the speaker of the voice data are the same person.
[0098] Among them, there are many methods to determine the target user matching the voice data. The following are two possible ways:
[0099] First, if it is determined that multiple likelihoods are greater than or equal to the likelihood threshold, the first user corresponding to the maximum likelihood is determined as a target user matching the voice data.
[0100] Second, if it is determined that multiple likelihoods are greater than or equal to the likelihood threshold, the multiple likelihoods are sorted in descending order, and the first users corresponding to the first M likelihoods are determined as multiple target users matching the voice data, where M is a positive integer greater than 1.
[0101] In some embodiments, if the server fails to match a target user from the voice feature set, the voice data is added to the historical voice data set as the voice data of a new user. Among them, the historical voice data set stores the historical voice data of each user. The historical voice data set includes multiple users and multiple voice data of each user, and each voice data corresponds to the user information of the user to which it belongs.
[0102] It can be understood that if the log-likelihood ratio determined based on the voice feature and the historical voice features of each first user is less than the likelihood threshold, it means that there is no user corresponding to any historical voice feature in the historical feature voice feature set who is the same as the speaker of the voice data. That is to say, the speaker information of the voice data does not exist in the voice feature set.
[0103] In some embodiments, when the number of new users in the historical voice data set is greater than the first number, and the voice data volume of each new user is greater than the second number, the server extracts features from multiple voice data of each new user to obtain multiple voice features of each new user. The server takes the average of the multiple voice features of each new user to obtain the average feature of each new user, and adds the average feature to the voice feature set as the voice feature of each new user. Among them, the voice data volume is the number of voice data or the capacity of the voice data. For example, when the number of voice data of each new user is greater than 500, the server extracts features from multiple voice data of each new user respectively to obtain multiple voice features of each new user; or when the capacity of the voice data of each new user is greater than 10MB, the server extracts features from multiple voice data of each new user respectively to obtain multiple voice features of each new user.
[0104] It should be noted that the process of obtaining the original voice features in the voice feature set is the same as that of the above-mentioned newly added users, and will not be elaborated here.
[0105] In some embodiments, the server obtains the third quantity of voice data of each newly added user, extracts features from each piece of voice data, and obtains the third quantity of voice features of each newly added user, where the third quantity is less than the second quantity.
[0106] For example, when the number of newly added users is greater than 100 and the number of voice data of each newly added user is greater than 300, the server obtains 10 pieces of voice data of each newly added user, invokes the DNN model, extracts features from the 10 pieces of voice data of each newly added user, and obtains 10 voice features of each newly added user. The average value of the 10 voice features of each newly added user is taken respectively to obtain the voice feature of each newly added user, and the voice feature of each newly added user is added to the voice feature set. At the same time, each voice feature corresponds to a user information.
[0107] In some embodiments, the server invokes the DNN model to extract features from multiple pieces of voice data of each newly added user respectively, and obtains the voice features of each piece of voice data of the newly added user respectively. It can be understood that after the number of newly added users and the voice data of each newly added user reach a certain amount respectively, by invoking the DNN model at one time, extracting features from multiple pieces of voice data of each newly added user respectively, and then uniformly determining the voice features of each newly added user and storing them, not only the processing resources are saved, but also the efficiency is improved.
[0108] Steps 302 to 304 are an implementation manner in which the server invokes the voiceprint recognition model to process the voice data and the voice feature set to determine the target user matching the voice data. In some embodiments, the process of determining the target user can also be performed in other ways, which will not be elaborated here.
[0109] 305. During the process of decoding the voice data, the server invokes the target language model matching the target user to process the voice data to obtain the target text corresponding to the voice data, where the target language model is trained based on the historical text data of the target user.
[0110] The initial language model is trained based on the historical text data of each target user, where the initial language model is an initial N-gram language model and N is equal to 2 or 3, to obtain the trained target language model of each target user, and each target language model corresponds to the user information to which it belongs respectively.
[0111] In some embodiments, the server, based on the determined target user and user information, obtains a target language model that matches the target user, and then invokes the target language model to process the voice data to obtain the target text corresponding to the voice data.
[0112] According to step 304, the number of target users is a positive integer greater than or equal to 1. Since the number of target language models is the same as the number of target users, there are two ways to obtain the target text corresponding to the voice data as follows:
[0113] First, when the number of target users is 1, the server invokes a target language model to process the voice data, obtains a target text corresponding to the voice data, and outputs the target text to complete the entire speech recognition process.
[0114] Second, when the number of target users is M, the server invokes the target language models of each target user to process the voice data respectively, obtains M candidate texts, and each candidate text corresponds to an identification score. The candidate text with the highest identification score is determined as the target text and the target text is output. During the process of decoding the voice data, each group of phoneme information output by the acoustic model corresponds to a phoneme score. The larger the phoneme score, the greater the possibility that the group of phonemes matches the voice data; multiple candidate characters or candidate words determined based on the pronunciation dictionary respectively correspond to word scores. The larger the word score, the greater the possibility that the character or word matches the phoneme; the text output by the target language model corresponds to a text score. The larger the text score, the higher the possibility of the word combination in the text. The identification score is determined based on the phoneme score, word score, and text score, and the candidate text with the highest identification score among the M candidate texts is determined as the target text.
[0115] 306. The server outputs the target text corresponding to the voice data.
[0116] By using the above speech recognition method, targeted recognition of voice data can be performed, and the accuracy of speech recognition can be improved.
[0117] As Figure 4 shown, in some embodiments, the speech recognition method further includes a training process of the voiceprint recognition model, including the following steps 401 to step 402:
[0118] 401. The server obtains a sample voice feature set, which is obtained based on the historical voice data of multiple users.
[0119] In some embodiments, the server obtains the historical voice data of multiple users, extracts features from each piece of historical voice data to obtain the voice features of each piece of historical voice data, and stores them in the sample voice feature set. When it is necessary to train the voiceprint recognition model, the server obtains this sample voice feature set. The sample voice feature set includes multiple voice feature pairs, which are obtained by randomly grouping the sample voice features in the sample voice feature set. Each voice feature pair includes two sample voice features, and each voice feature pair corresponds to a sample label.
[0120] 402. The server trains the voiceprint recognition model based on this sample voice feature set, using the user source as the supervision.
[0121] In some embodiments, the training process includes multiple iterations. Next, taking the i-th iteration process as an example, an explanation is given:
[0122] In the i-th iteration process, a voice feature pair is obtained from the sample voice feature set. The voice feature pair includes a first sample voice feature and a second sample voice feature, where i is a positive integer.
[0123] The voiceprint recognition model is called to process the first sample voice feature and the second sample voice feature to determine the likelihood between the first sample voice feature and the second sample voice feature.
[0124] If it is determined that the i-th iteration meets the training stop condition based on the likelihood and the sample label of the voice feature pair, then the voiceprint recognition model is determined to be the trained voiceprint recognition model. The sample label is used to indicate whether the sample voice features in the voice feature pair come from the same user.
[0125] If it is determined that the i-th iteration does not meet the training stop condition based on the likelihood and the sample label of the voice feature pair, then the model parameters of the voiceprint recognition model are adjusted, and the (i + 1)-th iteration process is performed based on the adjusted voiceprint recognition model.
[0126] Among them, when the sample voice features in the voice sample pair come from the same user, the sample label corresponding to the voice sample pair is 1; when the sample voice features in the voice sample pair do not come from the same user, the sample label corresponding to the voice sample pair is 0.
[0127] In some embodiments, the i-th iteration meeting the training condition means that the loss value determined based on the loss function tends to be stable.
[0128] The likelihood and the sample label are input into the loss function to obtain the corresponding loss value. The parameters of the voiceprint recognition model are adjusted based on this loss value until the loss value obtained from the training tends to be stable, and then the voiceprint recognition model is determined to be the trained voiceprint recognition model.
[0129] Among them, the loss value tending to be stable indicates that the loss function converges. The likelihood between the first sample voice feature and the second sample voice feature is also obtained by solving the log-likelihood ratio between the first sample voice feature and the second sample voice feature. The formula for determining the log-likelihood ratio is the same as the formula for solving the log-likelihood ratio in step 303, which will not be elaborated here.
[0130] In some embodiments, the user source refers to whether the users are the same. The larger the log-likelihood ratio, the greater the possibility that the first sample voice feature and the second sample voice feature belong to the same user; the smaller the log-likelihood ratio, the smaller the possibility that the first sample voice feature and the second sample voice feature belong to the same user.
[0131] For example, when i = 3, the training process is as follows:
[0132] In the 3rd iteration process, the first sample voice feature and the second sample voice feature are obtained.
[0133] The first sample voice feature and the second sample voice feature are input into the voiceprint recognition model after adjusting the model parameters through the 2nd iteration, and the log-likelihood ratio between the first sample voice feature and the second sample voice feature is obtained.
[0134] The log-likelihood ratio and the sample label are input into the loss function to obtain the corresponding loss value. If the loss value tends to be stable, the training is stopped; if the loss value does not tend to be stable, the model parameters of the voiceprint recognition model are adjusted based on the loss value, and the 4th iteration process is performed based on the adjusted voiceprint recognition model.
[0135] The technical solution provided by the embodiment of the present application calls the voiceprint recognition model to process the voice data and the voice data set, and can determine the target user matching the voice data, that is, first determine the identity of the speaker to which the voice data belongs. Then, based on the determined target user, the target language model matching the target user is called to process the voice data to obtain the target text of the voice data, and finally the target text corresponding to the voice data is output. Since the target language model is trained based on the historical text data of the target user, the target language model can more accurately reflect the speaking style of the target user, that is, using the target voice model, the voice data can be specifically recognized. Therefore, adopting this voice recognition method can improve the accuracy of voice recognition.
[0136] As Figure 5 shown, the embodiment of the present application also provides a voice recognition device, and the device includes:
[0137] The first acquisition module 501 is used to acquire voice data.
[0138] The first determination module 502 is configured to call a voiceprint recognition model to process the voice data and the voice feature set to determine a target user matching the voice data, where historical voice features of multiple users are stored in the voice feature set.
[0139] The calling module 503 is configured to, during the process of decoding the voice data, call a target language model matching the target user to process the voice data to obtain a target text corresponding to the voice data, where the target language model is trained based on historical text data of the target user.
[0140] The output module 504 is configured to output the target text corresponding to the voice data.
[0141] In some embodiments, the first determination module 502 is configured to:
[0142] Extract features from the voice data to obtain voice features.
[0143] Call a voiceprint recognition model to process the voice features and the historical voice features of the first user in the voice feature set to determine the likelihood between the voice features and the historical voice features of the first user.
[0144] In response to the likelihood being greater than or equal to a likelihood threshold, determine the first user as the target user matching the voice data.
[0145] In some embodiments, the apparatus further includes:
[0146] The adding module is configured to, if no target user is matched from the voice feature set, add the voice data as voice data of a new user to the historical voice data set.
[0147] In some embodiments, the apparatus further includes an updating module, and the updating module is configured to:
[0148] When the number of new users in the historical voice data set is greater than a first number, and the voice data volume of each new user is greater than a second number, extract features from multiple voice data of each new user to obtain multiple voice features of each new user.
[0149] Take the average of the multiple voice features of each new user to obtain an average feature of each new user, and add the average feature as the voice feature of each new user to the voice feature set.
[0150] In some embodiments, the apparatus further includes
[0151] The second acquisition module is configured to acquire a sample voice feature set, where the sample voice feature set is obtained based on historical voice data of multiple users.
[0152] A training module, configured to train a voiceprint recognition model based on the sample voice feature set with the user source as the supervision.
[0153] In some embodiments, the training module is configured to:
[0154] In the i-th iteration process, obtain a voice feature pair from the sample voice feature set, where the voice feature pair includes a first sample voice feature and a second sample voice feature, and i is a positive integer.
[0155] Call the voiceprint recognition model to process the first sample voice feature and the second sample voice feature to determine the likelihood between the first sample voice feature and the second sample voice feature.
[0156] If it is determined that the i-th iteration meets the training stop condition based on the likelihood and the sample label of the voice feature pair, then determine the voiceprint recognition model as the trained voiceprint recognition model, where the sample label is used to indicate whether the sample voice features in the voice feature pair are from the same user.
[0157] If it is determined that the i-th iteration does not meet the training stop condition based on the likelihood and the sample label of the voice feature pair, then adjust the model parameters of the voiceprint recognition model and perform the (i + 1)-th iteration process based on the adjusted voiceprint recognition model.
[0158] The technical solution provided by the embodiments of the present application calls the voiceprint recognition model to process voice data and a voice data set, and can determine a target user matching the voice data, that is, first determine the identity of the speaker to whom the voice data belongs. Then, based on the determined target user, call a target language model matching the target user to process the voice data to obtain the target text of the voice data, and finally output the target text corresponding to the voice data. Since the target language model is trained based on the historical text data of the target user, the target language model can more accurately reflect the speaking style of the target user. That is to say, using the target voice model, targeted recognition of the voice data can be performed. Therefore, by adopting this voice recognition method, the accuracy of voice recognition can be improved.
[0159] Figure 6It is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 600 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 601 and one or more memories 602. Among them, at least one computer program is stored in the one or more memories 602, and the at least one computer program is loaded and executed by the one or more processors 601 to implement the methods provided by the above-mentioned various method embodiments. Of course, the computer device 600 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input and output. The computer device 600 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0160] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program, and the above computer program can be executed by a processor to complete the speech recognition method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0161] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. The processor of the computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above speech recognition method.
[0162] In some embodiments, the computer program involved in the embodiments of the present application may be deployed to be executed on a single computer device, or on multiple computer devices located at one location. Or, it may be executed on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network may form a blockchain system.
[0163] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk, or an optical disc, etc.
[0164] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the present application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary.
[0165] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A speech recognition method, characterized in that, the method includes: obtaining speech data; invoking a voiceprint recognition model to process the speech data and a speech feature set to determine a target user matching the speech data, wherein multiple historical speech features of multiple users are stored in the speech feature set; invoking an acoustic model, inputting the speech data into the acoustic model to obtain multiple groups of phonemes corresponding to the speech data; based on the multiple groups of phonemes, determining multiple candidate characters or candidate words corresponding to the multiple groups of phonemes respectively in a pronunciation dictionary; invoking a target language model, inputting the multiple candidate characters or candidate words into the target language model matching the target user to obtain and output a target text corresponding to the speech data, wherein the target language model is trained based on the historical text data of the target user; when the number of target users is M, the server invokes the target language models of each target user to process the speech data respectively to obtain M candidate texts, each candidate text corresponds to an identification score, and the candidate text with the highest identification score is determined as the target text, the identification score is determined based on a phoneme score, a word / phrase score, and a text score, each group of phoneme information output by the acoustic model corresponds to the phoneme score respectively; multiple candidate characters or candidate words determined based on the pronunciation dictionary correspond to the word / phrase score respectively; the text output by the target language model corresponds to the text score; outputting the target text corresponding to the speech data; the method further includes: if the target user is not matched from the speech feature set, adding the speech data as the speech data of a new user to the historical speech data set; when the number of the new users in the historical speech data set is greater than a first number, and the speech data volume of each new user is greater than a second number, performing feature extraction on multiple speech data of each new user to obtain multiple speech features of each new user.
2. The speech recognition method according to claim 1, characterized in that, the invoking of the voiceprint recognition model to process the speech data and the speech feature set to determine a target user matching the speech data includes: performing feature extraction on the speech data to obtain a speech feature; invoking the voiceprint recognition model to process the speech feature and the historical speech feature of the first user in the speech feature set to determine the likelihood between the speech feature and the historical speech feature of the first user; in response to the likelihood being greater than or equal to a likelihood threshold, determining the first user as the target user matching the speech data.
3. The speech recognition method according to any one of claims 1-2, characterized in that, the training process of the voiceprint recognition model includes: obtaining a sample speech feature set, the sample speech feature set being obtained based on the historical speech data of the multiple users; training the voiceprint recognition model based on the sample speech feature set with the user source as the supervision.
4. The speech recognition method according to claim 3, characterized in that, Training the voiceprint recognition model based on the sample voice feature set with the user source as the supervision includes: In the i-th iteration process, obtain voice feature pairs from the sample voice feature set, where the voice feature pairs include a first sample voice feature and a second sample voice feature, and i is a positive integer; Call the voiceprint recognition model to process the first sample voice feature and the second sample voice feature to determine the likelihood between the first sample voice feature and the second sample voice feature; If it is determined that the i-th iteration meets the training stop condition based on the likelihood and the sample label of the voice feature pair, then determine the voiceprint recognition model as the trained voiceprint recognition model, where the sample label is used to indicate whether the sample voice features in the voice feature pair come from the same user; If it is determined that the i-th iteration does not meet the training stop condition based on the likelihood and the sample label of the voice feature pair, then adjust the model parameters of the voiceprint recognition model and perform the (i + 1)-th iteration process based on the adjusted voiceprint recognition model.
5. A voice recognition device Characterized in that The device includes: A first acquisition module for acquiring voice data; A first determination module for calling a voiceprint recognition model to process the voice data and the voice feature set to determine a target user matching the voice data, where the voice feature set stores historical voice features of multiple users; A call module for calling an acoustic model, inputting the voice data into the acoustic model to obtain multiple groups of phonemes corresponding to the voice data; based on the multiple groups of phonemes, determining multiple candidate characters or candidate words respectively corresponding to the multiple groups of phonemes in a pronunciation dictionary; calling a target language model, inputting the multiple candidate characters or candidate words into the target language model matching the target user to obtain and output a target text corresponding to the voice data, where the target language model is trained based on the historical text data of the target user; when the number of target users is M, the server calls the target language models of each target user to process the voice data respectively to obtain M candidate texts, each candidate text corresponding to an identification score, and determining the candidate text with the highest identification score as the target text, where the identification score is determined based on a phoneme score, a word score, and a text score; each group of phoneme information output by the acoustic model corresponds to the phoneme score; the multiple candidate characters or candidate words determined based on the pronunciation dictionary respectively correspond to the word score; the text output by the target language model corresponds to the text score; An output module for outputting the target text corresponding to the voice data; The device further includes: An addition module for adding the voice data as the voice data of a new user to the historical voice data set if the target user is not matched from the voice feature set; The update module is configured to: when the number of new users in the historical voice dataset is greater than a first number, and the voice data volume of each new user is greater than a second number, perform feature extraction on multiple voice data of each new user to obtain multiple voice features of each new user.
6. A computer device, characterized in that the computer device includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the voice recognition method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the voice recognition method according to any one of claims 1 to 4.
8. A computer program product, characterized in that the computer program product includes program code, the program code is stored in a computer-readable storage medium, a processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code so that the computer device executes the voice recognition method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Voice processing method and device, terminal and storage medium
CN111951790A
Voiceprint recognition method and device, storage medium and computer equipment
CN112259106A
Voice synthesis model training method and device, storage medium and electronic equipment
CN112309365A