Training of a language model for speech recognition, speech recognition method and device

By performing domain classification and weight evaluation on text data in the field of speech recognition, generating training statement sets and training language models, the problem of data imbalance is solved and the analysis performance of language models in the field of data sparseness is improved.

CN114299920BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111021975.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-01
Publication Date
2025-06-10
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively alleviate the problem of data imbalance in multiple fields in the field of speech recognition, resulting in poor analysis performance of language models in the field of data sparseness.

Method used

By performing domain classification processing on statements in the text dataset, evaluating the weight of each statement set, determining the target statement set, calculating the number of samples, assigning the sampling probability, performing statement extraction, generating a training statement set, and training the language model.

Benefits of technology

It effectively alleviates the problem of data imbalance in multiple fields, improves the analysis performance of language models in the field of data sparseness, and does not require the input of additional features and domain information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299920B_ABST
    Figure CN114299920B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for training a language model for speech recognition and a speech recognition method, which relate to the technical fields of artificial intelligence and map vehicle connection. The method includes: performing domain classification processing on the sentences in a text dataset to obtain at least one sentence set; performing weight evaluation on each sentence set to determine a target sentence set that meets a predetermined importance condition based on the weights of each sentence set; performing calculation processing based on the number of sentences and weights corresponding to the target sentence set to obtain a sampling number; performing sampling probability allocation processing according to the sampling number and the weights of the sentence sets to obtain the sampling probabilities of the sentences in each sentence set; extracting sentences from each sentence set according to the corresponding sampling probabilities to generate a training sentence set; and training a language model based on the training sentence set. The present application improves the analysis performance of the language model for speech recognition in data-sparse domains, and no additional features and domain information need to be input when the language model performs analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method and device for training a language model for speech recognition and a speech recognition method. Background Art

[0002] In fields such as speech recognition, a language model is usually used to analyze the occurrence probability of a sentence, that is, to estimate the occurrence probability of a piece of text. At present, in related technologies, there are methods for training a language model based on features and methods based on models.

[0003] The method based on features requires training an additional feature extraction model. The performance of the language model depends to a large extent on the performance of the feature extraction model, and the overall computational complexity is relatively high. The method based on the model adds a domain module to the model structure and updates it through data in the corresponding domain. When performing analysis, domain information needs to be input, but usually the domain information cannot be obtained. Moreover, in related technologies, it is difficult to effectively alleviate the problem of uneven data in multiple domains, resulting in poor analysis performance of the language model for data-sparse domains. Summary of the Invention

[0004] An embodiment of the present application provides a training solution for a language model for speech recognition, which can effectively alleviate the problem of uneven data in multiple domains, improve the analysis performance of the language model for speech recognition for data-sparse domains, and the language model does not need to input additional features and domain information when performing analysis.

[0005] The embodiments of the present application provide the following technical solutions:

[0006] According to an embodiment of the present application, a method for training a language model for speech recognition includes: performing domain classification processing on sentences in a text dataset to obtain sentence sets in at least one domain; performing weight evaluation on each of the sentence sets to determine a target sentence set that meets a predetermined importance condition based on the weights of each of the sentence sets; performing calculation processing based on the number of sentences and weights corresponding to the target sentence set to obtain the sampling number of sentences for training the language model; performing sampling probability allocation processing according to the sampling number and the weights of each of the sentence sets to obtain the sampling probability of sentences in each of the sentence sets; extracting sentences from each of the sentence sets according to the corresponding sampling probability to generate a training sentence set; and training the language model based on the training sentence set to obtain a trained language model.

[0007] According to an embodiment of the present application, a training device for a language model for speech recognition includes: a classification module configured to perform domain classification processing on statements in a text dataset to obtain statement sets for at least one domain; an evaluation module configured to perform weight evaluation on each of the statement sets to determine a target statement set that meets a predetermined importance condition based on the weights of each of the statement sets; a calculation module configured to perform calculation processing based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training the language model; an allocation module configured to perform sampling probability allocation processing according to the sampling number and the weights of each of the statement sets to obtain the sampling probability of statements in each of the statement sets; an extraction module configured to extract statements from each of the statement sets according to the corresponding sampling probability to generate a training statement set; and a training module configured to train the language model based on the training statement set to obtain a trained language model.

[0008] In some embodiments of the present application, the evaluation module includes: a grammar model training unit configured to train a target grammar model based on the statement sets for each domain to obtain a domain grammar model corresponding to each domain; a correlation analysis unit configured to perform correlation occurrence probability analysis on each word in a validation dataset by using each of the domain grammar models to obtain the correlation occurrence probability output by each of the domain grammar models; a difference processing unit configured to perform expectation maximization weight interpolation processing based on the correlation occurrence probability output by each of the domain grammar models to obtain the weights of each of the statement sets; and a target determination unit configured to determine a target statement set that meets a predetermined importance condition based on the weights.

[0009] In some embodiments of the present application, the target determination unit is configured to: determine a predetermined number of the statement sets with the largest weights as candidate statement sets; and determine the statement set with the largest number of included statements from the candidate statement sets as the target statement set.

[0010] In some embodiments of the present application, the target determination unit is configured to: determine at least one of the statement sets with weights greater than a predetermined threshold as candidate statement sets; and determine the statement set with the largest number of included statements from the candidate statement sets as the target statement set.

[0011] In some embodiments of the present application, the sum of the weights of all the statement sets is equal to one; the calculation module includes: an integer ratio calculation unit configured to calculate the integer ratio of the number of statements and the weights corresponding to the target statement set; and a sampling number determination unit configured to determine the integer ratio as the sampling number of statements collected from all the statement sets for training the language model.

[0012] In some embodiments of the present application, the allocation module includes: a weight ratio calculation unit, configured to calculate the weight ratio of the weight of each of the statement sets to the sum of the weights of all the statement sets, to obtain the weight ratio corresponding to each statement set; a statement sampling number calculation unit, configured to calculate the product of the weight ratio corresponding to each of the statement sets and the number of samplings, as the statement sampling number corresponding to each of the statement sets; a sampling probability determination unit, configured to perform a ratio calculation based on the statement sampling number corresponding to each of the statement sets and the number of statements, to determine the sampling probability of the statements in each of the statement sets.

[0013] In some embodiments of the present application, the sampling probability determination unit is configured to: for each of the statement sets, when the statement sampling number corresponding to the statement set is less than the number of statements, calculate the ratio of the statement sampling number corresponding to the statement set to the number of statements; when the statement sampling number corresponding to the statement set is greater than or equal to the number of statements, perform statement replication and expansion processing on the statement set, and calculate the ratio of the statement sampling number corresponding to the statement set to the number of statements after expansion; use the ratio corresponding to each of the statement sets as the sampling probability of the statements in each of the statement sets.

[0014] In some embodiments of the present application, the extraction module includes: a first extraction unit, configured to, for each of the statement sets, when the statement sampling number corresponding to the statement set is less than the number of statements, extract statements from the statement set according to the corresponding sampling probability; a second extraction unit, configured to, when the statement sampling number corresponding to the statement set is greater than or equal to the number of statements, extract statements from the statement set after expansion corresponding to the statement set according to the corresponding sampling probability; an aggregation unit, configured to determine the set of statements extracted from all the statement sets as the training statement set.

[0015] In some embodiments of the present application, the training module includes: a prediction unit, configured to analyze the statement occurrence probability of the statements in the training statement set by using the language model, to obtain the predicted statement occurrence probability corresponding to the statements in the training statement set; a cross-entropy calculation unit, configured to calculate the cross-entropy of the language model on the training statement set based on the predicted statement occurrence probability; an update unit, configured to update the parameters in the language model by stochastic gradient descent, so that the cross-entropy is less than a predetermined threshold, to generate the trained language model, and the trained language model is used for statement occurrence probability analysis.

[0016] According to an embodiment of the present application, a speech recognition method includes: performing speech recognition based on speech data of a target speech to obtain at least one candidate recognition text corresponding to the target speech; using a trained language model to perform sentence occurrence probability analysis on the at least one candidate recognition text to obtain a text score representing the sentence occurrence probability, where the trained language model is trained according to the aforementioned language model training method; and determining the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text.

[0017] According to an embodiment of the present application, a speech recognition device includes: a first-pass decoding module configured to perform speech recognition based on speech data of a target speech to obtain at least one candidate recognition text corresponding to the target speech; a second-pass decoding module configured to use a trained language model to perform sentence occurrence probability analysis on the at least one candidate recognition text to obtain a text score representing the sentence occurrence probability, where the trained language model is trained according to the aforementioned language model training method; and a recognition module configured to determine the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text.

[0018] In some embodiments of the present application, each candidate recognition text corresponds to an acoustic score and a language score. The acoustic score represents the occurrence probability of the target speech given the candidate recognition text, and the language score represents the occurrence probability of the word sequence corresponding to the candidate recognition text. The recognition module 730 is configured to: perform weighted summation on the acoustic score, the language score, and the text score corresponding to each candidate recognition text to obtain an accuracy score corresponding to each candidate recognition text; and determine the candidate recognition text with the maximum accuracy score as the speech recognition result of the target speech.

[0019] In some embodiments of the present application, the first-pass decoding module 710 is configured to: perform acoustic decoding processing on the speech data of the target speech to obtain at least one phoneme sequence corresponding to the target speech and the acoustic score corresponding to the phoneme sequence; and perform language decoding processing on each phoneme sequence to obtain at least one candidate recognition text corresponding to each phoneme sequence and the language score corresponding to the candidate recognition text.

[0020] According to another embodiment of the present application, a computer-readable storage medium stores a computer program, which, when executed by a processor of a computer, causes the computer to execute the method described in the embodiments of the present application.

[0021] According to another embodiment of the present application, an electronic device includes: a memory storing a computer program; and a processor configured to read the computer program stored in the memory to execute the method described in the embodiments of the present application.

[0022] According to another embodiment of the present application, a computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations described in the embodiments of the present application.

[0023] In the embodiments of the present application, domain classification processing is performed on the statements in the text dataset to obtain statement sets in at least one domain; weight evaluation is performed on each statement set to determine a target statement set that meets a predetermined importance condition based on the weights of each statement set; calculation processing is performed based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training a language model; sampling probability allocation processing is performed according to the sampling number and the weights of each statement set to obtain the sampling probability of the statements in each statement set; statements are extracted from each statement set according to the corresponding sampling probability to generate a training statement set; the language model is trained based on the training statement set to obtain a trained language model.

[0024] In this way, by performing domain classification processing on the statements in the text dataset to obtain at least one statement set, and then performing a series of instance samplings based on evaluating the weights of each statement set to obtain a training statement set for training the language model, the problem of data imbalance in multiple domains can be effectively alleviated, the analysis performance of the language model for speech recognition in data-sparse domains can be improved, and no additional features and domain information need to be input when the trained language model performs analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 A schematic diagram of a system to which the embodiments of the present application can be applied is shown.

[0027] Figure 2 A flowchart of a method for training a language model according to an embodiment of the present application is shown.

[0028] Figure 3 A flowchart of a method for determining a target statement set according to an embodiment of the present application is shown.

[0029] Figure 4Shows a flowchart of a sampling probability determination training method according to an embodiment of the present application.

[0030] Figure 5 Shows a flowchart of a speech recognition method according to an embodiment of the present application.

[0031] Figure 6 Shows a structural diagram of a speech recognition system applying an embodiment of the present application in a scenario.

[0032] Figure 7 Shows based on Figure 6 of the speech recognition system for speech recognition flowchart.

[0033] Figure 8 Shows a block diagram of a training device for a language model according to an embodiment of the present application.

[0034] Figure 9 Shows a block diagram of a speech recognition device according to an embodiment of the present application.

[0035] Figure 10 Shows a block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0037] Figure 1 Shows a schematic diagram of a system 100 to which the embodiments of the present application can be applied. As Figure 1 shown, the system 100 may include a server 101 and a terminal 102. The server 101 and the terminal 102 may be directly or indirectly connected through a wireless communication method, and the present application does not make special limitations here.

[0038] The server 101 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0039] The terminal 102 may be any device, and the terminal 102 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, VR / AR devices, smart watches, and computers, etc.

[0040] Among them, the server 101 can perform language model training, and the terminal 102 can perform speech recognition based on the language model trained by the server 101.

[0041] In one implementation of this example, the server 101 can perform domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain; perform weight evaluation on each of the statement sets to determine a target statement set that meets the predetermined importance condition based on the weights of each of the statement sets; perform ratio calculation processing based on the number of statements and weights of the target statement set to obtain the sampling number of statements for training the language model; perform sampling probability allocation processing according to the sampling number and the weights of each of the statement sets to obtain the sampling probability of the statements in each of the statement sets; extract statements from each of the statement sets according to the corresponding sampling probability to generate a training statement set; train the language model based on the training statement set to obtain a trained language model, and the trained language model is used for statement occurrence probability analysis.

[0042] In one implementation, the terminal 102 can perform speech recognition based on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech; use the trained language model to perform statement occurrence probability analysis on the at least one candidate recognition text to obtain a text score representing each of the candidate recognition texts, and the trained language model is obtained by the server 101 according to the foregoing embodiments; determine the speech recognition result of the target speech based on the text scores corresponding to each of the candidate recognition texts.

[0043] Figure 2 Schematically shows a flowchart of a method for training a language model according to an embodiment of the present application. The execution subject of the method for training the language model can be any terminal, such as Figure 1 the server 101 or the terminal 102 shown, etc.

[0044] As Figure 2 shown, the method for training the language model for speech recognition can include steps S210 to S260.

[0045] Step S210: Perform domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain; Step S220: Perform weight evaluation on each statement set to determine a target statement set that meets the predetermined importance condition based on the weight of each statement set; Step S230: Perform calculation processing based on the number of statements and the weight corresponding to the target statement set to obtain the sampling number of statements for training the language model; Step S240: Perform sampling probability allocation processing according to the sampling number and the weight of each statement set to obtain the sampling probability of the statements in each statement set; Step S250: Extract statements from each statement set according to the corresponding sampling probability to generate a training statement set; Step S260: Train the language model based on the training statement set to obtain a trained language model.

[0046] The following describes the specific process of each step when training a language model.

[0047] In step S210, perform domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain.

[0048] In the implementation manner of this example, the text dataset is a set of pre-collected text data. The text dataset may include at least one text data, and each text data may include at least one statement.

[0049] According to the collection sources of the text data in the text dataset, the text data can be classified by domain to obtain a sub-text data set corresponding to each domain. For example, the text data in the text dataset is classified by domain according to 6 collection sources: navigation, music, video, news, novel, and chat, to obtain 6 sub-text data sets for domains.

[0050] Among them, each sub-text data set is a statement set, that is, the set of statements in all text data within the sub-text data set forms a statement set.

[0051] In step S220, perform weight evaluation on each statement set to determine a target statement set that meets the predetermined importance condition based on the weight of each statement set.

[0052] In the implementation manner of this example, performing weight evaluation on each statement set can determine the weight of each statement set. The weight can represent the importance of each statement set, and the higher the weight, the higher the importance of the statement set. The predetermined importance condition is a predetermined weight evaluation condition. For example, the statement set with the largest number of statements among the predetermined number of the largest weights, etc. According to the predetermined importance condition, the target statement set can be selected from all statement sets.

[0053] Among them, the methods for evaluating the weights of the statement sets may include: the method of evaluating based on the training target grammar model, or the method of evaluating based on calculating the occurrence frequencies of words in the statement sets, etc.

[0054] The method of evaluating based on calculating the occurrence frequencies of words in the statement sets, for example, calculating the occurrence frequency of each word in the statement set in the statement set. For example, a certain word may appear 5 times. Then, calculate the average frequency of the frequencies of all words in the statement set, and normalize the average frequencies corresponding to all statement sets to obtain the weight corresponding to each statement set.

[0055] The method of evaluating based on the training target grammar model is as described in steps S221 to S224 in the following embodiments.

[0056] In one embodiment, refer to Figure 3 , step S220, perform weight evaluation on each statement set to determine the target statement set that meets the predetermined importance condition based on the weight of each statement set, including:

[0057] Step S221, train the target grammar model based on the statement sets in each field to obtain the domain grammar model corresponding to each field; Step S222, perform an associated occurrence probability analysis on each word in the validation data set using each domain grammar model to obtain the associated occurrence probability output by each domain grammar model; Step S223, perform an expectation maximization weight interpolation process based on the associated occurrence probability output by each domain grammar model to obtain the weight of each statement set; Step S224, determine the target statement set that meets the predetermined importance condition based on the weight.

[0058] The target grammar model is the grammar model, such as the N-Gram (Nth-order Markov chain) model. The grammar model performs an associated occurrence probability analysis on each word in the text data according to the context. The associated occurrence probability is the occurrence probability of the word calculated according to the context of the word in the text data.

[0059] Training the target grammar model based on the statement sets in each field can obtain the domain grammar model corresponding to each field, such as 6 domain grammar models corresponding to 6 fields of navigation, music, video, news, novels, and chatting. Among them, during the training process, the target grammar model performs an associated occurrence probability analysis on each word in the text data in the statement set according to the context. For example, when the N-gram model analyzes the associated occurrence probability of the Nth word in the text data, it is determined by the probabilities of the first N-1 words (i.e., the context) in the text. In an example of the present application, the target grammar model is a 3-gram model, that is, N is equal to 3.

[0060] The validation data set is a set of preset validation text data, and the validation text data set may include validation text data of at least one field.

[0061] For each word in the validation dataset, the associated occurrence probability is analyzed using each domain grammar model, and the associated occurrence probabilities output by each domain grammar model are obtained. Each domain grammar model outputs an associated occurrence probability for each word in the validation dataset. For example, if the validation dataset includes M words, each domain grammar model outputs M associated occurrence probabilities.

[0062] Based on the associated occurrence probabilities output by each domain grammar model, expectation-maximization weight interpolation processing is performed, that is, the associated occurrence probabilities output by each domain grammar model are subjected to weight interpolation processing through the expectation-maximization algorithm.

[0063] Specifically, first, the sum of the probabilities of the M associated occurrence probabilities output by each domain grammar model can be calculated. For example, when there are D domain grammar models, D probability sums P 1 to P D can be calculated; then, through the expectation-maximization algorithm, weight interpolation is performed on the probability sum corresponding to each domain grammar model. For example, D probability sums P 1 to P D are calculated, and weight interpolation is performed on P 1 to P D by the expectation-maximization algorithm. Solving for the maximum of W 1 *P 1 +W 2 *P 2 +...+W D *P D can be used to obtain W 1 to W D , where W 1 to W D are the weights of each statement set, and the sum of the weights of all statement sets is equal to 1, that is, W 1 +W 2 +...+W D = 1.

[0064] In this way, the calculated weights can accurately represent the importance of the statement sets, accurately determine the statement sets that meet the predetermined importance conditions as the target statement sets, and can improve the effectiveness of the present application in alleviating multi-domain data imbalance.

[0065] In one embodiment, step S224 of determining the target statement set that meets the predetermined importance conditions based on the weights includes: determining a predetermined number of statement sets with the largest weights as candidate statement sets; and determining the statement set with the largest number of included statements from the candidate statement sets as the target statement set.

[0066] For example, the weights include 0.2, 0.3, 0.15, and the predetermined number can be set according to requirements. For example, if the predetermined number is 2, then the predetermined number of statement sets with the largest weights, that is, the statement sets with weights of 0.2 and 0.3 (i.e., the top 2 statement sets in terms of weight ranking), and the neighborhood corresponding to the candidate statement sets determined in this way is the core neighborhood. Then, further determine the statement set with the largest number of statements from the candidate statement sets as the most core target statement set, which can further improve the effectiveness of alleviating multi-domain data imbalance in this application.

[0067] In one embodiment, in step S224, determining the target statement set that meets the predetermined importance condition based on the weights includes: determining at least one statement set with a weight greater than the predetermined threshold as the candidate statement set; and determining the statement set with the largest number of statements from the candidate statement sets as the target statement set.

[0068] For example, the weights include 0.2, 0.3, 0.15, and the predetermined threshold can be set according to requirements. For example, if the predetermined threshold is 0.2, then at least one statement set with a weight greater than the predetermined threshold, that is, the statement sets with weights of 0.2 and 0.3, and the neighborhood corresponding to the candidate statement sets determined in this way is the core neighborhood. Then, further determine the statement set with the largest number of statements from the candidate statement sets as the most core target statement set, which can also further improve the effectiveness of alleviating multi-domain data imbalance in this application.

[0069] In step S230, calculation processing is performed based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training the language model.

[0070] In the implementation manner of this example, the number of statements in the target statement set is the number of statements included in the target statement set, and the sampling number is the number of statements collected from all statement sets for training the language model.

[0071] Ratio calculation processing is performed based on the number of statements and weights of the target statement set to obtain the sampling number of statements for training the language model. In one example, it can be directly dividing the number of statements in the target statement set by the weight to generate an integer ratio, and then determining the integer ratio as the sampling number of statements for training the language model. In this way, the sampling number can be efficiently and extremely effectively determined; in another example, it can also be dividing the difference between the number of statements in the target statement set and the predetermined standard value by the weight to generate an integer ratio, and then determining the integer ratio as the sampling number of statements for training the language model. In this way, the sampling number can be conservatively determined.

[0072] In one embodiment, the sum of the weights of all statement sets is equal to one; in step S230, based on the number of statements and the weights of the target statement set, calculation processing is performed to obtain the sampling number of statements for training the language model, including: calculating the integer ratio of the number of statements and the weights corresponding to the target statement set; and determining the integer ratio as the sampling number of statements collected from all statement sets for training the language model.

[0073] The sum of the weights of all statement sets is equal to one, that is, the weight corresponding to the target statement set is less than one. By dividing the number of statements corresponding to the target statement set by the weight to obtain an integer ratio, the number of statements corresponding to the target statement set can be directly scaled up to the integer ratio, and the effective sampling number of statements collected from all statement sets for training the language model can be efficiently determined.

[0074] In step S240, according to the sampling number and the weights of each statement set, sampling probability distribution processing is performed to obtain the sampling probability of statements in each statement set.

[0075] In the implementation manner of this example, the sampling number is the total number of statements to be collected from all statement sets. The weight of the statement set can reflect the importance of the statement set. Furthermore, the statement sampling number can be correspondingly allocated according to the weight of each statement set, and then the sampling probability can be determined according to the ratio of the statement sampling number corresponding to each statement set to the number of statements in the statement set.

[0076] In one embodiment, refer to Figure 4 , step S240, according to the sampling number and the weights of each statement set, sampling probability distribution processing is performed to obtain the sampling probability of statements in each statement set, including:

[0077] Step S241, calculating the weight ratio of the weight of each statement set to the sum of the weights of all statement sets to obtain the weight ratio corresponding to each statement set; step S242, calculating the product of the weight ratio corresponding to each statement set and the sampling number as the statement sampling number corresponding to each statement set; step S243, performing a ratio calculation based on the statement sampling number corresponding to each statement set and the number of statements to determine the sampling probability of statements in each statement set.

[0078] If there are a total of D statement sets, the weights of the D statement sets include W 1 to W D , and the sum of the weights of all statement sets, that is, Y = W 1 + W 2 +... + W D . The weight of statement set D1 is W 1 , then the weight ratio corresponding to statement set D1 is W 1 / Y, and so on to obtain the weight ratio corresponding to each statement set. In some examples, the sum of weights Y = 1. In this case, the weight ratio of the statement set is the weight of the statement set itself.

[0079] The number of samples is the total number of statements to be collected from all statement sets. Calculate the product of the weight ratio corresponding to each statement set and the number of samples. For example, the number of samples is N, and the statement set D 1 (the first statement set) has a weight ratio of W 1 / Y. At this time, for the statement set D 1 the product of the corresponding weight ratio and the number of samples is d 1 = N * W 1 / Y, and d 1 is the number of statement samples to be collected from the statement set D 1 . And so on, for the statement set D D (the D-th statement set), the product of the corresponding weight ratio and the number of samples is d D = N * W D / Y. In this way, W 1 : W 2 :...: W D = d 1 : d 2 :...: d D , and d 1 + d 2 +...+ d D = N.

[0080] Finally, the number of statements corresponding to the statement set is the total sum of the number of all statements in the statement set. Based on the ratio calculation of the statement sample number corresponding to each statement set and the number of statements, it can be the number ratio obtained by dividing the statement sample number corresponding to the statement set by the number of statements. Based on the number ratio, determine the sampling probability of the statements in the statement set.

[0081] In one embodiment, step S243, based on the ratio calculation of the statement sample number corresponding to each of the statement sets and the number of statements to determine the sampling probability of the statements in each statement set, includes:

[0082] For each statement set, when the statement sample number corresponding to the statement set is less than the number of statements, calculate the ratio of the number of the statement sample number corresponding to the statement set to the number of statements; when the statement sample number corresponding to the statement set is greater than or equal to the number of statements, perform statement replication and expansion processing on the statement set, and calculate the ratio of the statement sample number corresponding to the statement set to the number of statements after expansion; use the ratio corresponding to each statement set as the sampling probability of the statements in each statement set.

[0083] When the number of statement samples corresponding to a statement set is less than the number of statements, directly use the ratio of the number of statement samples corresponding to the statement set divided by the number of statements as the sampling probability. For example, for statement set D i The number of statements corresponding to (where i is the i-th statement set) is s i , if the number of statement samples d i is less than s i , then for statement set D i The number of statement samples d corresponding to it i is divided by the number of statements s 1 to obtain the ratio d i / s 1 , and at this time d i / s i is used as the sampling probability of the statements in statement set D i .

[0084] When the number of statement samples corresponding to a statement set is greater than or equal to the number of statements, perform statement replication and expansion processing on the statement set. For example, for statement set D i The number of statements corresponding to (where i is the i-th statement set) is s i , if the number of statement samples d i is greater than or equal to s i , then replicate the statements in statement set D i to expand statement set D i to obtain an expanded statement set. The number of expanded statements in the expanded statement set is m*s i , such that m*s i is greater than d i .

[0085] Calculate the ratio of the number of statement samples corresponding to the statement set to the number of expanded statements, that is, the number of statement samples d corresponding to statement set D i is divided by the number of expanded statements m*s i to obtain the ratio d i / (m*s i ), and at this time d i / (m*s i ) is used as the sampling probability of the statements in statement set D i . i

[0086] In one embodiment, performing statement replication and expansion processing on a statement set includes: rounding up the ratio of the number of statement samples corresponding to the statement set to the number of statements to obtain an expansion multiple; performing replication and expansion processing on the statements in the statement set according to the expansion multiple.

[0087] The number of statements in statement set Di is s i , the number of statement samples is d i , and the ratio of the number of statement samples to the number of statements is d​i / s i , the number ratio d i / s i Round up to get After rounding up, the expansion multiple m is obtained. Then, the statements in the statement set Di are copied m times, and the number of statements in the expanded statement set after expansion is m * s i , so that m * s i is greater than d i .

[0088] In step S250, statements are extracted from each statement set according to the corresponding sampling probability to generate a training statement set.

[0089] In the implementation of this example, for the statement set D i , in one example, the statement set D i The corresponding sampling probability can be d i / s i , at this time, directly extract statements from the statement set D i according to the sampling probability d i / s i to obtain the statement set c i extracted from the statement set D i . In one example, the statement set D i The corresponding sampling probability can be d i / (m * s i ), at this time, it is necessary to expand the statement set D i to obtain the expanded statement set and extract statements according to the sampling probability d i / (m * s i ) to obtain the statement set c i extracted from the statement set D i .

[0090] Finally, the total set I of all the statement sets c 1 、c 2 、...、c i 、...、c D extracted from the D statement sets is the training statement set.

[0091] In one embodiment, step S250, extracting statements from each statement set according to the corresponding sampling probability to generate a training statement set, includes:

[0092] For each statement set, when the number of statement samples corresponding to the statement set is less than the number of statements, statements are extracted from the statement set according to the corresponding sampling probability; when the number of statement samples corresponding to the statement set is greater than or equal to the number of statements, statements are extracted from the expanded statement set corresponding to the statement set according to the corresponding sampling probability; the set of statements extracted from all statement sets is determined as the training statement set.

[0093] For statement set D i , when the number of statement samples corresponding to statement set D i is less than the number of statements, the sampling probability corresponding to statement set D i is d i / s i , and at this time, statements are directly extracted from statement set D i according to the sampling probability d i / s i to obtain the set of statements c i extracted from statement set D. i .

[0094] For statement set D i , when the number of statement samples corresponding to statement set D i is greater than or equal to the number of statements, the sampling probability corresponding to statement set D i is d i / (m * s i ), and at this time, statements are extracted from the expanded statement set obtained by expanding statement set D i according to the sampling probability d i / (m * s i ) to obtain the set of statements c i extracted from statement set D. i .

[0095] Finally, the total set I of all sets of statements c 1 、c 2 、...、c i 、...、c D extracted from D statement sets is the training statement set.

[0096] In step S260, the language model is trained based on the training statement set to obtain the trained language model.

[0097] In the implementation of this example, the language model is a neural network language model, and the language model can adopt a recurrent neural network (RNN) structure; it can also adopt the structure of a deep neural network (DNN), a long short-term memory network (LSTM), a convolutional neural network (CNN), an attention encoder-decoder (Transformer) model, or other neural network structures.

[0098] The language model can be used to perform sentence occurrence probability analysis. The sentence occurrence probability is the probability that a sentence appears as a sentence according to the collocation method of the words in the sentence (which can include the context and word correctness, etc.). Training the language model based on the training sentence set means training the language model to perform sentence occurrence probability analysis on the sentences in the training sentence set, and adjusting the parameters in the language model according to the analysis results until the performance such as the analysis accuracy of the language model meets the requirements to obtain the trained language model. Based on the trained language model, the sentence occurrence probability analysis can be performed on the sentence to be analyzed.

[0099] In one embodiment, step S260, training the language model based on the training sentence set to obtain a trained language model, includes:

[0100] Using the language model to perform sentence occurrence probability analysis on the sentences in the training sentence set to obtain the predicted sentence occurrence probabilities corresponding to the sentences in the training sentence set; calculating the cross-entropy of the language model on the training sentence set based on the predicted sentence occurrence probabilities; updating the parameters in the language model by stochastic gradient descent so that the cross-entropy is less than a predetermined threshold to generate a trained language model, and the trained language model is used to perform sentence occurrence probability analysis.

[0101] Using the language model to perform sentence occurrence probability analysis on the sentences in the training sentence set can enable the language model to analyze and predict the predicted sentence occurrence probability of each sentence. Each sentence can correspond to a true sentence occurrence probability. The distance between the predicted sentence occurrence probability and the true sentence occurrence probability can be calculated through the cross-entropy loss function. This distance is the cross-entropy of the language model on the training sentence set, and the cross-entropy can reflect the difficulty of the language model in analyzing and recognizing sentences.

[0102] The parameters of the language model can be optimized using the Stochastic Gradient Descent (SGD) algorithm, reducing the cross-entropy of the language model on the training dataset such that the cross-entropy is less than a predetermined threshold. At this point, it reflects that the language model can easily and accurately analyze and recognize sentences, and then a trained language model is obtained. Generally, the SGD algorithm and various improved SGD algorithms, such as gradient optimization algorithms like Adagrad, Adadelta, Adam, RMSProp, etc., can be used to update and optimize the parameters of the language model.

[0103] In this way, based on steps S210 to S260, by performing domain classification processing on the sentences in the text dataset, at least one sentence set is obtained. Then, based on evaluating the weights of each sentence set, a series of instance samplings are performed to obtain a training sentence set for training the language model, which can effectively alleviate the problem of data imbalance in multiple domains, improve the analysis performance of the language model for data-sparse domains, and no additional features and domain information need to be input when the trained language model performs analysis.

[0104] Figure 5 The flowchart of a speech recognition method according to an embodiment of the present application is schematically shown. The execution subject of this speech recognition method can be any terminal, such as Figure 1 the server 101 or the terminal 102 shown, etc.

[0105] As Figure 5 shown, this speech recognition method may include steps S310 to S330.

[0106] Step S310: Perform speech recognition on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech; Step S320: Use the trained language model to perform sentence occurrence probability analysis on at least one candidate recognition text to obtain a text score representing each candidate recognition text, where the trained language model is trained according to the language model training method in any of the foregoing embodiments; Step S330: Determine the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text.

[0107] The following describes the specific processes of each step when performing speech recognition.

[0108] In step S310, speech recognition is performed on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech.

[0109] In the implementation of this example, the target voice can be a voice in any field of at least one field. For example, the query voices in the fields of music, point of interest (POI), weather, news, general knowledge, etc. in the in-vehicle intelligent voice recognition system. The voice data of the target voice can be voice waveform data.

[0110] Speech recognition is to extract acoustic features from language data to obtain acoustic features, decode the acoustic features to obtain a phoneme sequence, and decode the phoneme sequence to obtain at least one possible candidate recognition text of the target voice. The candidate recognition text is the word sequence that the phoneme sequence may correspond to. Furthermore, the voice data of the target voice is decoded once through speech recognition to obtain at least one possible candidate recognition text.

[0111] Among them, speech recognition can evaluate an acoustic score and a language score for each candidate recognition text. The acoustic score can represent the occurrence probability of the audio corresponding to the target voice when a given candidate recognition text is provided, and the language score can represent the occurrence probability of the word sequence corresponding to the candidate recognition text.

[0112] Among them, when performing speech recognition based on the voice data of the target voice, a hybrid speech recognition model composed of an acoustic model (such as a Hidden Markov Model (HMM)) and a language model (such as a Deep Neural Network Model (DNN)), as well as end-to-end speech recognition models such as RNN-T, Encoder-Decoder, LAS, etc. can be used for speech recognition.

[0113] In one embodiment, step S310, performing speech recognition based on the voice data of the target voice to obtain at least one candidate recognition text corresponding to the target voice, includes: performing acoustic decoding processing on the voice data of the target voice to obtain at least one phoneme sequence corresponding to the target voice and the acoustic score corresponding to the phoneme sequence; performing language decoding processing on each phoneme sequence to obtain at least one candidate recognition text corresponding to each phoneme sequence and the language score corresponding to the candidate recognition text.

[0114] A pre-trained speech recognition model can be used to perform speech recognition on the voice data of the target voice. First, acoustic features are extracted from the language data to obtain acoustic features; then, the speech model can include an acoustic model (such as a Hidden Markov Model (HMM)) to perform acoustic decoding on the acoustic features to obtain a phoneme sequence and the acoustic score corresponding to the phoneme sequence; at the same time, the speech model can also include a language model (such as a Deep Neural Network Model (DNN)). The language model can perform language decoding on the phoneme sequence to obtain at least one word sequence that the phoneme sequence may correspond to and the language score corresponding to the word sequence. Each word sequence is a candidate recognition text.

[0115] Furthermore, each candidate recognition text may correspond to an acoustic score and a language score. The acoustic score may represent the occurrence probability of the audio corresponding to the target speech when the given candidate recognition text is provided, and the language score may represent the occurrence probability of the word sequence corresponding to the candidate recognition text.

[0116] In step S320, using the trained language model, perform sentence occurrence probability analysis on at least one candidate recognition text to obtain a text score representing the sentence occurrence probability. The trained language model is trained according to the training method of the language model for speech recognition in any of the foregoing embodiments.

[0117] In the implementation manner of this example, the trained language model is trained according to the training method of the language model in any of the foregoing embodiments, has multi-domain analysis capabilities and good analysis performance in the data sparse domain, and can further perform accurate sentence occurrence probability analysis on each candidate recognition text, perform two-pass decoding on the target speech, and obtain a text score representing the sentence occurrence probability corresponding to each candidate recognition text.

[0118] In step S330, determine the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text.

[0119] In the implementation manner of this example, based on the text score corresponding to each candidate recognition text, the candidate recognition text with the highest text score can be determined as the speech recognition result, or the speech recognition result can be comprehensively determined by combining the acoustic score, language score, and text score of the candidate recognition text obtained by speech recognition.

[0120] The speech recognition result is comprehensively determined by combining the acoustic score, language score, and text score of the candidate recognition text obtained by speech recognition, as described in the following embodiments.

[0121] In one embodiment, each candidate recognition text corresponds to an acoustic score and a language score. The acoustic score represents the occurrence probability of the target speech when the given candidate recognition text is provided, and the language score represents the occurrence probability of the word sequence corresponding to the candidate recognition text; step S330, determining the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text, includes:

[0122] Perform weighted summation on the acoustic score, language score, and text score corresponding to each candidate recognition text to obtain an accuracy score corresponding to each candidate recognition text; determine the candidate recognition text with the maximum accuracy score as the speech recognition result of the target speech.

[0123] For example, through speech recognition, N candidate recognition texts h 1 to h N corresponding to the target speech are obtained. The acoustic score and language score corresponding to each candidate recognition text may be (a1 , l 1 ), (a 2 , l 2 ) to (a N , l N ), where a i can be the acoustic score of the i-th candidate recognition text h i , and l i can be the language score of the i-th candidate recognition text. The text score of each candidate recognition text can be x 1 , x 2 ,..., x N , and x i can be the text score of the i-th candidate recognition text h i .

[0124] Perform a weighted sum of the acoustic score, language score, and text score corresponding to each candidate recognition text. The weighted sum can be performed according to a predetermined weighting coefficient. For example, for the candidate recognition text h i , the corresponding acoustic score a i , language score l i , and text score h i , the weighted sum is Q i = αa i + βl i + γx i . Q i is the accuracy score corresponding to the candidate recognition text h i . Among them, the magnitude relationship of the first weighting coefficient α, the second weighting coefficient β, and the third weighting coefficient γ can be 0 < α, β, γ < 1, and α + β + γ = 1. In one implementation, α = 0.5, β = 0.25, and γ = 0.25.

[0125] Finally, determine the candidate recognition text with the maximum accuracy score as the speech recognition result of the target speech, which can further effectively improve the accuracy of the speech recognition result.

[0126] In this way, based on steps S310 to S330, on the basis of one-pass decoding, perform two-pass decoding through the trained language model, and determine the speech recognition result based on the text score, which can effectively improve the accuracy of the speech recognition result, especially can effectively improve the recognition accuracy of speech recognition in the data-sparse field.

[0127] According to the method described in the above embodiments, the following will further illustrate with examples in application scenarios. The flow of the training method of the language model for speech recognition according to the embodiments of the present application in this scenario, the meanings of the nouns in this scenario are the same as those in the foregoing embodiments, and specific descriptions can be referred to in the foregoing embodiments. This flow can specifically include steps 1 to 6:

[0128] Step 1: Perform domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain. Specifically, classify according to the collection sources of the statements in the text dataset to obtain statement sets for at least one domain.

[0129] Step 2: Evaluate the weights of each statement set to determine the target statement set that meets the predetermined importance conditions based on the weights of each statement set. Step 2 includes Steps 2.1 to 2.3:

[0130] Step 2.1: Count the number of statements included in each statement set to obtain the number of statements corresponding to each statement set.

[0131] Step 2.2: Train the target grammar model based on the statement sets in each domain to obtain the domain grammar model corresponding to each domain;

[0132] Step 2.3: Perform an analysis of the associated occurrence probability of each word in the validation dataset using each of the domain grammar models to obtain the associated occurrence probabilities output by each domain grammar model; perform an expectation-maximization weight interpolation process based on the associated occurrence probabilities output by each domain grammar model to obtain the weights of each statement set, and the sum of the weights of all statement sets is equal to one; determine the target statement set that meets the predetermined importance conditions based on the weights.

[0133] Among them, determining the target statement set that meets the predetermined importance conditions based on the weights may include: determining the predetermined number of statement sets with the largest weights as candidate statement sets; determining the statement set with the largest number of included statements from the candidate statement sets as the target statement set. Or, determining at least one statement set with a weight greater than the predetermined threshold as a candidate statement set; determining the statement set with the largest number of included statements from the candidate statement sets as the target statement set.

[0134] For example, if there are D statement sets, and the numbers of statements corresponding to the D statement sets are s 1 、s i 、...、s D respectively, determine the predetermined number K of statement sets with the largest weights or K statement sets with weights greater than the predetermined threshold as candidate statement sets. The statement set with the largest number of included statements is the statement set corresponding to max{s 1 ,s 2 ,...,s k}.

[0135] Step 3: Perform calculation processing based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training the language model. Specifically: Refer to the formula Calculate the number of statements corresponding to the target statement set max{s 1 ,s2 , ..., s k} and weights The integer ratio; determine the integer ratio as the sampling number N of the statements collected from all statement sets for training the language model.

[0136] Step 4, perform sampling probability distribution processing according to the sampling number and the weight of each statement set to obtain the sampling probability of the statements in each statement set. Specifically, referring to the formula Calculate the weight ratio of the weight of each statement set to the sum of the weights of all statement sets to obtain the weight ratio corresponding to each statement set; calculate the product of the weight ratio corresponding to each statement set and the sampling number as the statement sampling number corresponding to each statement set; for each statement set, when the statement set D i The corresponding statement sampling number d i Is less than the number of statements s i , calculate the statement sampling number d i Corresponding to the statement set and the number of statements s i The ratio r i = d i / s i ; when the statement set D i The corresponding statement sampling number d i Is greater than or equal to the number of statements s i , perform statement replication and expansion processing on the statement set D i , and calculate the statement sampling number d i Corresponding to the statement set and the number of statements after expansion m*s i The ratio r i = d i / (m*s i ); use the ratio r i Corresponding to each statement set as the sampling probability of the statements in each statement set.

[0137] Among them, the statement replication and expansion processing for the statement set D i Includes: rounding up the ratio of the statement sampling number corresponding to the statement set Di to the number of statements to obtain the expansion multiple; replicating and expanding the statements in the statement set according to the expansion multiple. Specifically, the ratio of the statement sampling number to the number of statements is d i / s i , and rounding up the ratio d i / s i Is Rounding up to obtain the expansion multiple m, and then replicating the statements in the statement set Di m times to obtain the number of statements after expansion in the statement set after expansion as m*s i , so that m*s i Is greater than d i .

[0138] Step 5: Extract sentences from each sentence set according to the corresponding sampling probability to generate a training sentence set. Specifically, for each sentence set, when the number of sentence samplings corresponding to the sentence set is less than the number of sentences, extract sentences from the sentence set according to the corresponding sampling probability; when the number of sentence samplings corresponding to the sentence set is greater than or equal to the number of sentences, extract sentences from the expanded sentence set corresponding to the sentence set according to the corresponding sampling probability; determine the set of sentences extracted from all sentence sets as the training sentence set.

[0139] Step 6: Train the language model based on the training sentence set to obtain a trained language model. Specifically, construct a language model; use the language model to analyze the sentence occurrence probability of the sentences in the training sentence set to obtain the predicted sentence occurrence probability corresponding to the sentences in the training sentence set; calculate the cross-entropy of the language model on the training sentence set based on the predicted sentence occurrence probability; update the parameters in the language model by stochastic gradient descent so that the cross-entropy is less than a predetermined threshold to generate a trained language model.

[0140] Among them, the language model is a neural network language model, and the language model can adopt a recurrent neural network (RNN) structure; it can also adopt a deep neural network (DNN), a long short-term memory network (LSTM), a convolutional neural network (CNN), an attention encoder-decoder (Transformer) model structure or other neural network structures.

[0141] After training to generate a trained language model, the trained language model can be used for speech recognition. Figure 6 Fig. shows a structural diagram of a speech recognition system applying an embodiment of the present application in a scenario. Figure 7 Fig. shows Figure 6 a flowchart of speech recognition based on the speech recognition system. The meanings of the nouns in this scenario are the same as those in the foregoing embodiments, and specific descriptions can be referred to in the foregoing embodiments.

[0142] Refer to Figure 6 , in this scenario, the speech recognition system 400 may include a speech data acquisition module 410, a first-pass decoding module 420, a second-pass decoding module 430, and an identification module 440. The speech recognition process based on this speech recognition system may specifically include steps S510 to S540:

[0143] Step S510: Based on the voice data acquisition module 410, acquire the voice data of the target voice.

[0144] The voice recognition system is, for example, an in-vehicle intelligent voice recognition system. The target voice can be a voice in any field among at least one field input into the voice recognition system. For example, the query voice in the in-vehicle intelligent voice recognition system. The voice data of the target voice can be voice waveform data.

[0145] Step S520: Based on the one-pass decoding module 420, perform a first decoding process on the voice data to obtain at least one candidate recognition text. Specifically, it includes: performing voice recognition on the voice data of the target voice to obtain at least one candidate recognition text corresponding to the target voice. Specifically: performing acoustic decoding on the voice data of the target voice to obtain at least one phoneme sequence corresponding to the target voice and the acoustic score corresponding to the phoneme sequence; performing language decoding on each phoneme sequence to obtain at least one candidate recognition text corresponding to each phoneme sequence and the language score corresponding to the candidate recognition text. Each candidate recognition text corresponds to an acoustic score and a language score. The acoustic score represents the occurrence probability of the target voice when a given candidate recognition text is provided, and the language score represents the occurrence probability of the word sequence corresponding to the candidate recognition text.

[0146] Step S530: Based on the two-pass decoding module 430, perform a second decoding process on the candidate recognition text using the trained language model. Specifically, it includes: using the trained language model to analyze the sentence occurrence probability of at least one candidate recognition text to obtain a text score representing the sentence occurrence probability. The trained language model is trained according to the aforementioned training method of the language model.

[0147] Step S540: Based on the recognition module 440, obtain the voice recognition result. Specifically, it includes: determining the voice recognition result of the target voice based on the text score corresponding to each candidate recognition text. Specifically, perform a weighted sum of the acoustic score, language score, and text score corresponding to each candidate recognition text to obtain the accuracy score corresponding to each candidate recognition text; determine the candidate recognition text with the maximum accuracy score as the voice recognition result of the target voice.

[0148] Specifically, if N candidate recognition texts h 1 to h N corresponding to the target voice are obtained through voice recognition. The acoustic score and language score corresponding to each candidate recognition text can be (a 1 , l 1 ), (a 2 , l 2 ) to (a N , l N ), where a i can be the i-th candidate recognition text hi The acoustic score, l i can be the language score of the i-th candidate recognition text. The text score of each candidate recognition text can be x 1 , x 2 , ..., x N , x i can be the text score of the i-th candidate recognition text h i .

[0149] Performing a weighted sum on the acoustic score, language score, and text score corresponding to each candidate recognition text, the weighted sum can be performed according to a predetermined weighting coefficient. For example, for the candidate recognition text h i , the corresponding acoustic score a i , language score l i , and text score h i , the weighted sum is Q i = αa i + βl i + γx i , Q i is the accuracy score corresponding to the candidate recognition text h i . Among them, the magnitude relationship of the first weighting coefficient α, the second weighting coefficient β, and the third weighting coefficient γ can be 0 < α, β, γ < 1, and α + β + γ = 1. In one implementation, α = 0.5, β = 0.25, γ = 0.25.

[0150] In this way, the accuracy of the speech recognition result is effectively improved, and particularly, the recognition accuracy rate of speech recognition in the data sparse field can be effectively improved.

[0151] Furthermore, the following table gives a detailed comparison of the word error rates of the language recognition results after adopting the training method of the language model proposed in this application and the conventional language model training method in this scenario. Among them, the text dataset contains 4 fields: navigation, music, news, and chatting. Among them, navigation and music are the core fields. The text data in the music field is 5GB, and the text data in the navigation field is 40GB. The music field is relatively data sparse compared to the navigation field. Test set 1 and test set 2 are user data collected from the intelligent in-vehicle voice service. Test set 1 only contains data in the navigation field, and test set 2 only contains data in the music field. The lower the word error rate, the higher the accuracy rate of the speech recognition result. The test results show that by applying the embodiments of this application, in the music field with the data sparse problem, the speech recognition performance is significantly improved, and the word error rate has a relative decrease of 9%.

[0152] Conventional method Method of the present application Test set 1 (navigation field) 2.73% 2.53% Test set 2 (music field) 6.53% 5.95%

[0153] To facilitate the better implementation of the training method of the language model provided in the embodiments of the present application, the embodiments of the present application also provide a training device for a language model based on the above-mentioned training method of the language model. The meanings of the nouns are the same as those in the above-mentioned training method of the language model, and the specific implementation details can be referred to the descriptions in the method embodiments.

[0154] Figure 8 The block diagram of a training device for a language model for speech recognition according to an embodiment of the present application is shown.

[0155] As Figure 8 shown, the training device 600 for a language model for speech recognition may include a classification module 610, an evaluation module 620, a calculation module 630, an allocation module 640, an extraction module 650, and a training module 660.

[0156] The classification module 610 may be used to perform domain classification processing on the statements in the text dataset to obtain statement sets in at least one domain; the evaluation module 620 may be used to perform weight evaluation on each of the statement sets to determine a target statement set that meets the predetermined importance condition based on the weights of each of the statement sets; the calculation module 630 may be used to perform calculation processing based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training the language model; the allocation module 640 may be used to perform sampling probability allocation processing according to the sampling number and the weights of each of the statement sets to obtain the sampling probability of the statements in each of the statement sets; the extraction module 650 may be used to extract statements from each of the statement sets according to the corresponding sampling probability to generate a training statement set; the training module 660 may be used to train the language model based on the training statement set to obtain a trained language model.

[0157] In some embodiments of the present application, the evaluation module 620 includes: a grammar model training unit, configured to train a target grammar model based on the statement sets in each domain to obtain a domain grammar model corresponding to each domain; an association analysis unit, configured to perform an associated occurrence probability analysis on each word in the verification dataset by using each of the domain grammar models to obtain the associated occurrence probability output by each of the domain grammar models; a difference processing unit, configured to perform an expectation maximization weight interpolation process based on the associated occurrence probabilities output by each of the domain grammar models to obtain the weights of each of the statement sets; a target determination unit, configured to determine a target statement set that meets the predetermined importance condition based on the weights.

[0158] In some embodiments of the present application, the target determination unit is configured to: determine a predetermined number of the statement sets with the largest weights as candidate statement sets; and determine the statement set with the largest number of included statements from the candidate statement sets as the target statement set.

[0159] In some embodiments of the present application, the target determination unit is configured to: determine at least one of the statement sets whose weights are greater than a predetermined threshold as a candidate statement set; determine the statement set with the largest number of statements included from the candidate statement sets as the target statement set.

[0160] In some embodiments of the present application, the sum of the weights of all the statement sets is equal to one; the calculation module 630 includes: an integer ratio calculation unit, configured to calculate the integer ratio of the number of statements corresponding to the target statement set and the weight; a sampling number determination unit, configured to determine the integer ratio as the sampling number of statements collected from all the statement sets for training the language model.

[0161] In some embodiments of the present application, the allocation module 640 includes: a weight ratio calculation unit, configured to calculate the weight ratio of the weight of each statement set to the sum of the weights of all the statement sets, to obtain the weight ratio corresponding to each statement set; a statement sampling number calculation unit, configured to calculate the product of the weight ratio corresponding to each statement set and the sampling number as the statement sampling number corresponding to each statement set; a sampling probability determination unit, configured to perform a ratio calculation based on the statement sampling number corresponding to each statement set and the number of statements to determine the sampling probability of the statements in each statement set.

[0162] In some embodiments of the present application, the sampling probability determination unit is configured to: for each statement set, when the statement sampling number corresponding to the statement set is less than the number of statements, calculate the ratio of the statement sampling number corresponding to the statement set to the number of statements; when the statement sampling number corresponding to the statement set is greater than or equal to the number of statements, perform statement replication and expansion processing on the statement set, and calculate the ratio of the statement sampling number corresponding to the statement set to the number of statements after expansion; use the ratio corresponding to each statement set as the sampling probability of the statements in each statement set.

[0163] In some embodiments of the present application, the extraction module 650 includes: a first extraction unit, configured to, for each statement set, when the statement sampling number corresponding to the statement set is less than the number of statements, extract statements from the statement set according to the corresponding sampling probability; a second extraction unit, configured to, when the statement sampling number corresponding to the statement set is greater than or equal to the number of statements, extract statements from the statement set after expansion corresponding to the statement set according to the corresponding sampling probability; a set unit, configured to determine the set of statements extracted from all the statement sets as the training statement set.

[0164] In some embodiments of the present application, the training module 660 includes: a prediction unit configured to analyze the occurrence probability of statements in the training statement set using the language model to obtain the predicted statement occurrence probability corresponding to the statements in the training statement set; a cross-entropy calculation unit configured to calculate the cross-entropy of the language model on the training statement set based on the predicted statement occurrence probability; and an update unit configured to update the parameters in the language model by stochastic gradient descent to make the cross-entropy less than a predetermined threshold, thereby generating the trained language model, which is used for statement occurrence probability analysis.

[0165] In this way, based on the training device 600 for a language model used in speech recognition, by performing domain classification processing on the statements in the text dataset, at least one statement set can be obtained. Then, based on evaluating the weights of each statement set, a series of instance samplings are performed to obtain a training statement set for training the language model, which can effectively alleviate the problem of data imbalance in multiple domains, improve the analysis performance of the language model for data-sparse domains, and no additional features and domain information need to be input when the trained language model performs analysis.

[0166] Figure 9 The block diagram of a speech recognition device according to an embodiment of the present application is shown.

[0167] As Figure 9 shown, the speech recognition device 700 may include a first-pass decoding module 710, a second-pass decoding module 720, and an identification module 730.

[0168] The first-pass decoding module 710 may be configured to perform speech recognition based on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech; the second-pass decoding module 720 may be configured to use the trained language model to analyze the occurrence probability of the at least one candidate recognition text to obtain a text score representing the occurrence probability of the statement, and the trained language model is trained according to any one of the foregoing methods; the identification module 730 may be configured to determine the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text.

[0169] In some embodiments of the present application, each candidate recognition text corresponds to an acoustic score and a language score. The acoustic score represents the occurrence probability of the target speech when a given candidate recognition text is provided, and the language score represents the occurrence probability of the word sequence corresponding to the candidate recognition text; the identification module 730 is configured to: perform a weighted sum of the acoustic score, the language score, and the text score corresponding to each candidate recognition text to obtain an accuracy score corresponding to each candidate recognition text; and determine the candidate recognition text with the maximum accuracy score as the speech recognition result of the target speech.

[0170] In some embodiments of the present application, the one-pass decoding module 710 is configured to: perform acoustic decoding processing on the speech data of the target speech to obtain at least one phoneme sequence corresponding to the target speech and the acoustic score corresponding to the phoneme sequence; perform language decoding processing on each phoneme sequence to obtain at least one candidate recognition text corresponding to each phoneme sequence and the language score corresponding to the candidate recognition text.

[0171] In this way, based on the speech recognition device 700, on the basis of one-pass decoding, two-pass decoding is performed through a trained language model, and the speech recognition result is determined based on text scoring, which can effectively improve the accuracy of the speech recognition result, especially can effectively improve the recognition accuracy of speech recognition in the field of data sparsity.

[0172] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0173] In addition, an embodiment of the present application further provides an electronic device, which can be a terminal or a server. For example, Figure 10 as shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:

[0174] The electronic device may include a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, an input unit 804 and other components. Those skilled in the art can understand that Figure 10 the structure of the electronic device shown in

[0175] The processor 801 is the control center of the electronic device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 802, and invoking the data stored in the memory 802, it executes various functions of the computer device and processes data, thereby performing an overall detection of the electronic device. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 801 either.

[0176] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the computer device. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0177] The electronic device further includes a power supply 803 for supplying power to each component. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 803 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0178] The electronic device may further include an input unit 804, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0179] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 801 in the electronic device will load the executable files corresponding to the processes of one or more computer programs into the memory 802 according to the following instructions, and the processor 801 will run the computer programs stored in the memory 802 to implement various functions. For example, the processor 801 can execute:

[0180] Perform domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain; perform weight evaluation on each of the statement sets to determine a target statement set that meets a predetermined importance condition based on the weights of each of the statement sets; perform calculation processing based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training a language model; perform sampling probability allocation processing based on the sampling number and the weights of each of the statement sets to obtain the sampling probability of statements in each of the statement sets; extract statements from each of the statement sets according to the corresponding sampling probability to generate a training statement set; train the language model based on the training statement set to obtain a trained language model.

[0181] In some embodiments, the processor 801 may execute:

[0182] Perform speech recognition on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech; use the trained language model to perform statement occurrence probability analysis on the at least one candidate recognition text to obtain a text score representing the statement occurrence probability, where the trained language model is trained according to any one of the foregoing methods; determine the speech recognition result of the target speech based on the text scores corresponding to each of the candidate recognition texts.

[0183] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the foregoing embodiments can be completed by a computer program or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0184] Therefore, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any one of the methods provided by the embodiments of the present application.

[0185] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), a magnetic disk, an optical disk, etc.

[0186] Since the computer program stored in the computer-readable storage medium can execute the steps in any one of the methods provided by the embodiments of the present application, the beneficial effects achievable by the methods provided by the embodiments of the present application can be realized. For details, see the foregoing embodiments and will not be repeated here.

[0187] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the foregoing embodiments of the present application.

[0188] Those skilled in the art will readily conceive of other implementations of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0189] It should be understood that the present application is not limited to the embodiments described above and shown in the drawings, but various modifications and changes can be made without departing from its scope.

Claims

1. A training method for a language model for speech recognition, characterized in that, it includes: Performing domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain; Performing weight evaluation on each of the statement sets to determine a target statement set that meets a predetermined importance condition based on the weights of each of the statement sets; Performing calculation processing based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training the language model; Performing sampling probability allocation processing according to the sampling number and the weights of each of the statement sets to obtain the sampling probability of the statements in each of the statement sets, including: calculating the weight ratio of the weight of each of the statement sets to the sum of the weights of all the statement sets to obtain the weight ratio corresponding to each statement set; calculating the product of the weight ratio corresponding to each of the statement sets and the sampling number as the statement sampling number corresponding to each of the statement sets; performing ratio calculation based on the statement sampling number corresponding to each of the statement sets and the number of statements to determine the sampling probability of the statements in each of the statement sets; Extracting statements from each of the statement sets according to the corresponding sampling probability to generate a training statement set; Training the language model based on the training statement set to obtain a trained language model.

2. The method according to claim 1, characterized in that, The performing weight evaluation on each of the statement sets to determine a target statement set that meets a predetermined importance condition based on the weights of each of the statement sets includes: Training a target grammar model based on the statement sets in each domain to obtain a domain grammar model corresponding to each domain; Performing associated occurrence probability analysis on each word in the validation dataset using each of the domain grammar models to obtain the associated occurrence probability output by each of the domain grammar models; Performing expectation maximization weight interpolation processing based on the associated occurrence probability output by each of the domain grammar models to obtain the weight of each of the statement sets; Determining a target statement set that meets a predetermined importance condition based on the weights.

3. The method according to claim 2, characterized in that, The determining a target statement set that meets a predetermined importance condition based on the weights includes: Determining a predetermined number of the statement sets with the largest weights as candidate statement sets, or determining at least one of the statement sets with weights greater than a predetermined threshold as candidate statement sets; Determining the statement set with the largest number of included statements from the candidate statement sets as the target statement set.

4. The method according to claim 1, characterized in that, The sum of the weights of all the statement sets is equal to one; The performing calculation processing based on the number of statements and weights corresponding to the target statement set to obtain the sampling number of statements for training the language model includes: Calculating the integer ratio of the number of statements and weights corresponding to the target statement set; Determining the integer ratio as the sampling number of statements collected from all the statement sets for training the language model.

5. The method according to claim 1, characterized in that, Performing ratio calculation based on the number of sampled statements and the number of statements corresponding to each statement set to determine the sampling probability of statements in each statement set, including: For each statement set, when the number of sampled statements corresponding to the statement set is less than the number of statements, calculating the ratio of the number of sampled statements corresponding to the statement set to the number of statements; When the number of sampled statements corresponding to the statement set is greater than or equal to the number of statements, performing statement replication and expansion processing on the statement set, and calculating the ratio of the number of sampled statements corresponding to the statement set to the number of statements after expansion; Taking the ratio corresponding to each statement set as the sampling probability of statements in each statement set.

6. The method according to claim 5, wherein, The extracting statements from each statement set according to the corresponding sampling probability to generate a training statement set includes: For each statement set, when the number of sampled statements corresponding to the statement set is less than the number of statements, extracting statements from the statement set according to the corresponding sampling probability; When the number of sampled statements corresponding to the statement set is greater than or equal to the number of statements, extracting statements from the expanded statement set corresponding to the statement set according to the corresponding sampling probability; Determining the set of statements extracted from all statement sets as the training statement set.

7. The method according to claim 1, wherein, The training the language model based on the training statement set to obtain a trained language model includes: Using the language model to perform statement occurrence probability analysis on the statements in the training statement set to obtain the predicted statement occurrence probability corresponding to the statements in the training statement set; Calculating the cross-entropy of the language model on the training statement set based on the predicted statement occurrence probability; Updating the parameters in the language model by stochastic gradient descent so that the cross-entropy is less than a predetermined threshold to generate the trained language model, and the trained language model is used for statement occurrence probability analysis.

8. A speech recognition method, wherein, including: Performing speech recognition on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech; Using the trained language model to perform statement occurrence probability analysis on the at least one candidate recognition text to obtain a text score representing the statement occurrence probability, and the trained language model is trained according to the method according to any one of claims 1 to 7; Determining the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text.

9. The method according to claim 8, wherein, Each candidate recognition text corresponds to an acoustic score and a language score, the acoustic score represents the occurrence probability of the target speech when a given candidate recognition text is provided, and the language score represents the occurrence probability of the word sequence corresponding to the candidate recognition text; The determining the speech recognition result of the target speech based on the text score corresponding to each candidate recognition text includes: Perform a weighted sum of the acoustic score, the language score, and the text score corresponding to each of the candidate recognition texts to obtain an accuracy score corresponding to each of the candidate recognition texts; Determine the candidate recognition text with the maximum accuracy score as the speech recognition result of the target speech.

10. The method according to claim 9, wherein, performing speech recognition on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech, including: performing acoustic decoding processing on the speech data of the target speech to obtain at least one phoneme sequence corresponding to the target speech and the acoustic score corresponding to the phoneme sequence; performing language decoding processing on each of the phoneme sequences to obtain at least one candidate recognition text corresponding to each of the phoneme sequences and the language score corresponding to the candidate recognition text.

11. A training device for a language model for speech recognition, wherein, comprising: a classification module, configured to perform domain classification processing on the statements in the text dataset to obtain statement sets for at least one domain; an evaluation module, configured to perform weight evaluation on each of the statement sets to determine a target statement set that meets a predetermined importance condition based on the weights of each of the statement sets; a calculation module, configured to perform calculation processing based on the number of statements and the weights corresponding to the target statement set to obtain the sampling number of statements for training the language model; an allocation module, configured to perform sampling probability allocation processing according to the sampling number and the weights of each of the statement sets to obtain the sampling probability of the statements in each of the statement sets, including: calculating the weight ratio of the weight of each of the statement sets to the sum of the weights of all the statement sets to obtain the weight ratio corresponding to each of the statement sets; calculating the product of the weight ratio corresponding to each of the statement sets and the sampling number as the statement sampling number corresponding to each of the statement sets; performing a ratio calculation based on the statement sampling number corresponding to each of the statement sets and the number of statements to determine the sampling probability of the statements in each of the statement sets; an extraction module, configured to extract statements from each of the statement sets according to the corresponding sampling probability to generate a training statement set; a training module, configured to train the language model based on the training statement set to obtain a trained language model.

12. A speech recognition device, wherein, comprising: a first-pass decoding module, configured to perform speech recognition on the speech data of the target speech to obtain at least one candidate recognition text corresponding to the target speech; a second-pass decoding module, configured to use the trained language model to analyze the statement occurrence probability of the at least one candidate recognition text to obtain a text score representing the statement occurrence probability, and the trained language model is trained according to the method according to any one of claims 1 to 7; a recognition module, configured to determine the speech recognition result of the target speech based on the text score corresponding to each of the candidate recognition texts.

13. A computer-readable storage medium, wherein, A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that the computer program product includes a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.

15. An electronic device, characterized in that it includes: a memory storing a computer program; a processor reading the computer program stored in the memory to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Text classification method and device based on deep learning, server and storage medium

    CN112329836A