Speech recognition method, device, electronic device and storage medium

By iterating the student model through parameter, using the differences between the domain sample speech and the general speech recognition model, the problem of low speech recognition accuracy in specific domain scenarios is solved, and higher recognition accuracy and faster model convergence is achieved.

CN114708852BActive Publication Date: 2025-05-13HEFEI IFLY DIGITAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210255584.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-05-13
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The prior art has low speech recognition accuracy in specific fields.

Method used

By using the difference between the label recognition text and the first recognized text based on the domain sample speech, and the difference between the first recognized text and the second recognized text, the student model is iterated parameterally to obtain the speech recognition model. This model can not only learn specific language expressions in the field sample pronunciation, but also learn common language expressions from the teacher model.

Benefits of technology

This improves the accuracy of speech recognition in specific domain scenarios, reduces the demand for the speech collection volume of domain samples, and speeds up the convergence speed of student models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708852B_ABST
    Figure CN114708852B_ABST
Patent Text Reader

Abstract

The present invention provides a speech recognition method, device, electronic device and storage medium, the method comprising: inputting speech features of speech to be recognized into a speech recognition model to obtain recognition text output by the speech recognition model; the speech recognition model is obtained by performing parameter iteration on a student model based on the difference between a label recognition text of a domain sample speech and a first recognition text, and the difference between the first recognition text and the second recognition text; the first recognition text is determined by the student model based on the speech features of the domain sample speech, the second recognition text is determined by the teacher model based on the speech features of the domain sample speech, and the teacher model is trained based on general sample speech and its label recognition text. The speech recognition method, device, electronic device and storage medium provided by the present invention can accurately perform speech recognition in specific domain scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech recognition method, device, electronic equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, speech recognition technology has been widely used in various fields such as education, entertainment, medical care, and transportation.

[0003] At present, speech recognition models are mostly obtained by collecting a large amount of corpus data in general scenarios, and speech recognition is performed based on the speech recognition model. However, when the speech recognition model is applied to specific domain scenarios, the speech recognition accuracy is low. Summary of the invention

[0004] The present invention provides a speech recognition method, device, electronic device and storage medium, which are used to solve the defect of low speech recognition accuracy in specific field scenarios in the prior art.

[0005] The present invention provides a speech recognition method, comprising:

[0006] Determine the speech to be recognized;

[0007] Inputting the speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model;

[0008] The speech recognition model is obtained by iterating parameters of a student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text;

[0009] The first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech. The teacher model is trained based on general sample speech and its label recognition text.

[0010] According to a speech recognition method provided by the present invention, the training step of the speech recognition model includes:

[0011] Perturbing the speech features of the domain sample speech, and inputting the perturbed speech features of the domain sample speech into the student model to obtain the first recognition text output by the student model;

[0012] Inputting the speech features of the domain sample speech into the teacher model to obtain the second recognition text output by the teacher model;

[0013] Based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text, iterate the parameters of the student model to obtain the speech recognition model;

[0014] The initialization parameters of the student model are iteratively obtained based on the general sample speech and its label recognition text.

[0015] According to a speech recognition method provided by the present invention, the difference between the label recognition text and the first recognition text based on the domain sample speech, and the difference between the first recognition text and the second recognition text, iterate the parameters of the student model to obtain the speech recognition model, including:

[0016] Determining a first loss value based on a difference between the labeled recognition text of the domain sample speech and the first recognition text;

[0017] determining a second loss value based on a difference between the first recognized text and the second recognized text;

[0018] Based on the first loss value and the second loss value, parameters of the student model are iterated to obtain the speech recognition model.

[0019] According to a speech recognition method provided by the present invention, the step of determining the label recognition text of the domain sample speech includes:

[0020] Inputting the speech features of the domain sample speech into the student model to obtain a first label recognition text output by the student model;

[0021] Inputting the speech features of the domain sample speech into a general speech recognition model to obtain a second label recognition text output by the general speech recognition model;

[0022] Determining a label recognition text of the domain sample speech from the first label recognition text and the second label recognition text based on the speech duration of the domain sample speech;

[0023] The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

[0024] According to a speech recognition method provided by the present invention, the step of determining the label recognition text of the domain sample speech from the first label recognition text and the second label recognition text based on the speech duration of the domain sample speech comprises:

[0025] Determining the number of characters per unit time length of the domain sample speech based on the speech time length of the domain sample speech and the number of characters of the first label recognition text;

[0026] If the number of characters per unit time length of the domain sample speech is less than the character threshold, the first label recognition text is used as the label recognition text; if not, the second label recognition text is used as the label recognition text of the domain sample speech.

[0027] According to a speech recognition method provided by the present invention, the speech features of the speech to be recognized are input into a speech recognition model to obtain a recognition text output by the speech recognition model, and then the method further includes:

[0028] Based on the number of characters in the recognized text and the speech duration of the speech to be recognized, the recognized text is corrected to obtain a corrected text.

[0029] According to a speech recognition method provided by the present invention, the recognized text is corrected based on the number of characters in the recognized text and the speech duration of the speech to be recognized to obtain the corrected text, including:

[0030] Determining the number of characters per unit time length of the speech to be recognized based on the number of characters in the recognized text and the speech time length of the speech to be recognized;

[0031] If the number of characters per unit time length of the speech to be recognized is greater than or equal to the character threshold, the speech features of the speech to be recognized are input into a universal speech recognition model to obtain a universal recognition text output by the universal speech recognition model, and the universal recognition text is used as the correction text;

[0032] If the number of characters per unit time length of the speech to be recognized is less than the character threshold, the recognized text is used as the corrected text;

[0033] The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

[0034] The present invention also provides a speech recognition device, comprising:

[0035] A speech determination unit, used to determine the speech to be recognized;

[0036] A speech recognition unit, used for inputting the speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model;

[0037] The speech recognition model is obtained by iterating parameters of a student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text;

[0038] The first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech. The teacher model is trained based on general sample speech and its label recognition text.

[0039] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned speech recognition methods is implemented.

[0040] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the speech recognition method described in any one of the above is implemented.

[0041] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the speech recognition method described above is implemented.

[0042] The speech recognition method, device, electronic device and storage medium provided by the present invention perform parameter iteration on a student model based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text, so that the student model can not only learn the specific language expressions in the domain sample speech based on the difference between the label recognition text of the domain sample speech and the first recognition text, but also learn the common language expressions in the domain sample speech from the teacher model based on the difference between the first recognition text and the second recognition text, thereby enabling the trained speech recognition model to accurately perform speech recognition in domain scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 It is one of the flow charts of the speech recognition method provided by the present invention;

[0045] Figure 2This is one of the flow charts of the speech recognition model training method provided by the present invention;

[0046] Figure 3 is a flowchart of an implementation of step 230 in the speech recognition model training method provided by the present invention;

[0047] Figure 4 This is the second flow chart of the speech recognition model training method provided by the present invention;

[0048] Figure 5 It is a flow chart of the method for determining the label recognition text of the field sample speech provided by the present invention;

[0049] Figure 6 It is a flowchart of an implementation method of step 530 in the method for determining the label recognition text of the domain sample speech provided by the present invention;

[0050] Figure 7 It is a flow chart of the method for determining the corrected text provided by the present invention;

[0051] Figure 8 This is the second flow chart of the speech recognition method provided by the present invention;

[0052] Fig. 9 It is a structural schematic diagram of the speech recognition device provided by the present invention;

[0053] Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] With the rapid development of artificial intelligence technology, speech recognition technology has been widely used in various fields such as education, entertainment, medical care, and transportation. At present, speech recognition models are mostly obtained by collecting a large amount of corpus data in common scenarios, and speech recognition is performed based on the speech recognition model. The speech recognition model has good recognition effect when applied to common scenarios.

[0056] However, since the speech to be recognized corresponding to specific domain scenarios and general scenarios has different degrees of differences in channels, topics, speakers, environmental noise, etc., when the speech recognition model trained by traditional methods is applied to specific scenarios, the recognition effect is poor.

[0057] To this end, the present invention provides a speech recognition method. Figure 1 It is one of the flow charts of the speech recognition method provided by the present invention, such as Figure 1 As shown, the method comprises the following steps:

[0058] Step 110: Determine the speech to be recognized.

[0059] Specifically, the speech to be recognized is the speech data that needs to be recognized. The speech to be recognized can be the speech data recorded in real time by the user through the electronic device. The electronic device here can be a smart phone, a tablet computer, or a smart appliance such as a stereo, a television, and an air conditioner. After obtaining the speech to be recognized, the electronic device can also amplify and reduce noise of the speech to be recognized. In addition, the speech to be recognized can also be the stored or received speech data, which is not specifically limited in the embodiment of the present invention.

[0060] Step 120: input the speech features of the speech to be recognized into the speech recognition model to obtain the recognition text output by the speech recognition model;

[0061] The speech recognition model is obtained by iterating parameters of the student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text;

[0062] The first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech. The teacher model is trained based on general sample speech and its label recognition text.

[0063] Specifically, the general sample speech is the speech collected in the general scene, and its label recognition text is the label corresponding to the general sample speech, so the teacher model trained based on the general sample speech and its label recognition text can be understood as a speech recognition model suitable for general scenes in the traditional method. The domain sample speech is the speech collected in the domain scene, and its label recognition text is the label corresponding to the domain sample speech. Among them, there are different degrees of differences between the speech in the domain scene and the speech in the general scene in terms of channel, topic, speaker, environmental noise, etc. For example, the general scene can be a general life scene, and the domain scene can be a specific industry scene, such as the domain scene can be a medical scene.

[0064] After the speech to be recognized is determined, the speech features of the speech to be recognized may be extracted. The speech features of the speech to be recognized may be extracted by a feature extraction algorithm, such as by extracting the speech features of the speech to be recognized based on Fourier transform.

[0065] After obtaining the speech features of the speech to be recognized, the speech features of the speech to be recognized are input into the speech recognition model, and the speech recognition model performs speech recognition on the speech features of the speech to be recognized to obtain the recognized text. The speech recognition model is obtained by iterating the parameters of the student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text.

[0066] Here, the first recognition text is determined by the student model based on the speech features of the domain sample speech, so that the difference between the label recognition text of the domain sample speech and the first recognition text is used to characterize the model performance of the student model in the domain scenario. The smaller the difference, the better the recognition effect of the student model in the domain scenario. The second recognition text is determined by the teacher model based on the speech features of the domain sample speech, so that the difference between the first recognition text and the second recognition text is used to characterize the performance difference between the student model and the teacher model when applied in the domain scenario. The smaller the difference between the first recognition text and the second recognition text, the smaller the performance difference between the student model and the teacher model in the domain scenario, which means that the more knowledge the student model learns from the teacher model, the better the performance.

[0067] It should be noted that if the initial model of the speech recognition model in the traditional method is trained based solely on domain sample speech and its labeled recognition text, in order to enable the trained model to have a better recognition effect, a large amount of domain sample speech is required. However, the sample speech in the domain scenario is usually not easy to obtain, that is, it is difficult to obtain a sufficient amount of domain sample speech.

[0068] Because the speech recognition model in the traditional method has good recognition effect in general scenarios, that is, the speech recognition model in the traditional method has the function of recognizing speech in general scenarios, but the recognition effect is poor for certain specific words, specific sentences, etc. in domain scenarios.

[0069] In this regard, the embodiment of the present invention is based on the fact that the speech recognition model (i.e., the teacher model) in the traditional method has a certain speech recognition ability, and then combines the domain sample speech and its label recognition text to iterate the parameters of the student model, so that the student model can not only learn the language expressions of specific words, specific sentences, etc. in the domain sample speech based on the difference between the label recognition text of the domain sample speech and the first recognition text, but also can learn the language expressions of common words, common sentences, etc. in the domain sample speech from the teacher model based on the difference between the first recognition text and the second recognition text, so that the trained speech recognition model can accurately perform speech recognition in the domain scenario.

[0070] In addition, since the teacher model is trained based on general sample speech and its label recognition text, and general sample speech is relatively easy to obtain, the initial model of the teacher model can be trained based on sufficient sample speech, so that the teacher model can accurately recognize speech in general scenarios, that is, the teacher model can accurately recognize general words, general sentences, etc. in speech, and the speech in the domain scenario includes both general words and general sentences, as well as specific words and specific sentences. The embodiment of the present invention trains the student model in combination with the teacher model that can accurately recognize general words and general sentences in speech, so that there is no need to incrementally obtain the corresponding domain sample speech for training in terms of general words, general sentences, etc. in the domain scenario, that is, not only the amount of domain sample speech collected is reduced, but also the convergence speed of the student model is accelerated.

[0071] The speech recognition method provided by the embodiment of the present invention performs parameter iteration on the student model based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text, so that the student model can not only learn the specific language expressions in the domain sample speech based on the difference between the label recognition text of the domain sample speech and the first recognition text, but also learn the common language expressions in the domain sample speech from the teacher model based on the difference between the first recognition text and the second recognition text, thereby enabling the trained speech recognition model to accurately perform speech recognition in domain scenarios.

[0072] Based on the above embodiments, Figure 2 It is one of the flow charts of the speech recognition model training method provided by the present invention, such as Figure 2 As shown, the training steps of the speech recognition model include:

[0073] Step 210, perturb the speech features of the domain sample speech, and input the perturbed speech features of the domain sample speech into the student model to obtain the first recognition text output by the student model; the initialization parameters of the student model are iteratively obtained based on the general sample speech and its label recognition text.

[0074] Step 220: input the speech features of the domain sample speech into the teacher model to obtain a second recognition text output by the teacher model;

[0075] Step 230: Based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text, the student model is iterated on parameters to obtain a speech recognition model.

[0076] Specifically, the initialization parameters of the student model are iteratively obtained based on the general sample speech and its label recognition text, that is, it can be understood that the initialization model of the student model and the teacher model are both speech recognition models in the traditional method.

[0077] After determining the speech features of the domain sample speech, the speech features of the domain sample speech are disturbed, and the disturbed domain sample speech features are input into the student model, and the student model performs speech recognition based on the disturbed domain sample speech features to obtain a first recognition text. At the same time, the speech features of the domain sample speech are input into the teacher model, and the teacher model performs speech recognition based on the domain sample speech features to obtain a second recognition text.

[0078] Then, based on the difference between the label recognition text of the domain sample speech and the first recognition text, the student model learns the language expressions of specific words, specific sentences, etc. in the domain sample speech, and based on the difference between the first recognition text and the second recognition text, the student model learns the language expressions of common words, common sentences, etc. in the domain sample speech from the teacher model, thereby enabling the trained speech recognition model to accurately perform speech recognition in domain scenarios.

[0079] Among them, perturbing the speech features of the domain sample speech can achieve the expansion of the domain sample speech, so that the student model can better learn the language expression of speech in the domain scenario. Optionally, when perturbing the speech features of the domain sample speech, the speech features of the domain sample speech can be masked to obtain mask features corresponding to the speech features to expand the domain sample speech.

[0080] It can be seen that the embodiment of the present invention iterates the parameters of the student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text, so that the obtained speech recognition model can accurately perform speech recognition in the domain scenario.

[0081] Based on any of the above embodiments, Figure 3 is a flow chart of an implementation of step 230 in the speech recognition model training method provided by the present invention, such as Figure 3 As shown, step 230 specifically includes:

[0082] Step 231, determining a first loss value based on the difference between the label recognition text of the domain sample speech and the first recognition text;

[0083] Step 232: determining a second loss value based on the difference between the first recognized text and the second recognized text;

[0084] Step 233: Based on the first loss value and the second loss value, perform parameter iteration on the student model to obtain a speech recognition model.

[0085] Specifically, the size of the first loss value is used to characterize the recognition effect of the student model in the domain scenario. The smaller the first loss value, the better the recognition effect of the student model in the domain scenario. The size of the second loss value is used to characterize the degree of performance difference between the student model and the teacher model in the domain scenario. The smaller the second loss value, the smaller the performance difference between the student model and the teacher model in the domain scenario, which means that the more knowledge the student model learns from the teacher model, the better the performance.

[0086] After determining the first loss value and the second loss value, the weights of the two can be added to obtain the loss value of the student model, and the parameters of the student model can be iterated based on the loss value to obtain the speech recognition model. Since the student model is obtained after the parameters are iterated based on the first loss value and the second loss value, the student model can learn specific expressions in the domain sample speech based on the first loss value, and can also learn general expressions in the domain sample speech from the teacher model based on the second loss value, so that the trained speech recognition model can accurately perform speech recognition in the domain scenario.

[0087] Based on any of the above embodiments, Figure 4 FIG. 2 is a flow chart of the speech recognition model training method provided by the present invention. Figure 4 As shown, the training method includes:

[0088] First, extract the speech features Fb of the domain sample speech, then perturb the speech features of the domain sample speech (such as mask processing tfmask), and input the perturbed speech features of the domain sample speech into the student model student to obtain the first recognition text predict_s output by the student model. At the same time, input the speech features Fb of the domain sample speech into the teacher model teacher to obtain the second recognition text predict_t output by the teacher model teacher. The initialization parameters of the student model student are iteratively obtained based on the general sample speech and its label recognition text.

[0089] Then, based on the difference between the label recognition text Pseudo label of the domain sample speech and the first recognition text predict_s, a first loss value is determined, and based on the difference between the first recognition text predict_s and the second recognition text predict_t, a second loss value is determined.

[0090] Then, the first loss value and the second loss value are weighted and added to obtain the loss value of the student model, and the parameters of the student model are iterated based on the loss value to obtain the speech recognition model. The loss value Loss of the student model can be determined based on the following formula:

[0091] Loss = λ×CE+(1-λ)×KLD

[0092] Among them, CE represents the first loss value, KLD represents the second loss value, and λ represents the weight, which is an adjustable parameter with a value range of (0,1).

[0093] Based on any of the above embodiments, Figure 5 FIG. 1 is a flow chart of a method for determining a label recognition text of a field sample speech provided by the present invention. Figure 5 As shown, the steps of determining the label recognition text of the domain sample speech include:

[0094] Step 510: input the speech features of the domain sample speech into the student model to obtain the first label recognition text output by the student model;

[0095] Step 520: Input the speech features of the domain sample speech into the general speech recognition model to obtain the second label recognition text output by the general speech recognition model;

[0096] Step 530: Determine the label recognition text of the domain sample speech from the first label recognition text and the second label recognition text based on the speech duration of the domain sample speech;

[0097] The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

[0098] It should be noted here that if manually annotated domain sample speech is used to obtain its corresponding label recognition text, language experts in specific fields are required to perform the annotation, which not only takes a long time but also has a high annotation cost.

[0099] In this regard, the embodiment of the present invention first inputs the speech features of the domain sample speech into the student model, and the student model performs speech recognition based on the speech features of the domain sample speech to obtain a first label recognition text. At the same time, the speech features of the domain sample speech are input into the general speech recognition model, and the general speech recognition model performs speech recognition based on the speech features of the domain sample speech to obtain a second label recognition text. Among them, the general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model. For example, the structure of the student model can be an end-to-end model structure, and the general speech recognition model can be an acoustic model + language model structure, wherein the acoustic model can adopt a triphone model structure, and the speech model can adopt an N-Gram model structure, and the embodiment of the present invention does not specifically limit this.

[0100] Due to the different structures of the student model and the teacher model, there may be differences between the first label recognition text and the second label recognition text output by the student model. In this case, it is necessary to filter out the text with higher accuracy from the first label recognition text and the second label recognition text as the label recognition text of the domain sample speech.

[0101] Since the student model is obtained through iterative training based on the domain sample speech and its label recognition text, that is, as the student model is continuously updated, its speech recognition accuracy in the domain scenario will become higher and higher, but the general speech recognition model is obtained through training based on the general sample speech and its label recognition text, that is, although the general speech recognition model can accurately recognize speech in general scenarios, its speech recognition accuracy in the domain scenario remains unchanged. In order to accurately obtain the label recognition text of the domain sample speech, when determining the label recognition text of the domain sample speech based on the student model and the general speech recognition model, the embodiment of the present invention gives priority to the first label recognition text output by the student model as the label recognition text of the domain sample speech. If the first label recognition text is inaccurate, the second label recognition text is used as the label recognition text of the domain sample speech.

[0102] Specifically, in some cases, the student model may have insertion errors when performing speech recognition on the domain sample speech. For example, the domain sample speech is "I love work", but the student model may have insertion errors in the recognition process, resulting in the first label recognition text being "I love love love love love love work", which obviously contains too many inserted words "love". Given that the number of characters a user speaks per unit time is usually upper bounded, the number of characters per unit time of the domain sample speech can be determined based on the speech time of the domain sample speech. If the number of characters per unit time is larger, it indicates that the probability of insertion errors in the first label recognition text output by the student model is higher, and the second label recognition text can be selected as the label recognition text for the domain sample speech; if the number of characters per unit time is smaller, it indicates that the probability of insertion errors in the first label recognition text output by the student model is lower, and the first label recognition text can be selected as the label recognition text for the domain sample speech.

[0103] It can be seen that the embodiment of the present invention can accurately identify insertion errors in the first label recognition text based on the speech duration of the domain sample speech, and then accurately determine the label recognition text of the domain sample speech from the first label recognition text and the second label recognition text.

[0104] Figure 6 is a flowchart of an implementation of step 530 in the method for determining the label recognition text of the domain sample speech provided by the present invention, such as Figure 6 As shown, step 530 includes:

[0105] Step 531: determining the number of characters per unit time length of the domain sample speech based on the speech time length of the domain sample speech and the number of characters of the first label recognition text;

[0106] Step 532: If the number of characters per unit time length of the domain sample speech is less than the character threshold, the first label recognition text is used as the label recognition text; if not, the second label recognition text is used as the label recognition text of the domain sample speech.

[0107] Specifically, the number of characters per unit time length of the domain sample speech refers to the number of characters per unit time length corresponding to the first label recognition text obtained when the student model recognizes the domain sample speech, which can be determined by the speech duration of the domain sample speech and the number of characters in the first label recognition text, such as the number of characters per unit time length of the domain sample speech = the number of characters in the first label recognition text / the speech duration of the domain sample speech.

[0108] If the number of characters per unit time length of the domain sample speech is less than the character threshold, it indicates that the probability of insertion errors in the first label recognition text outputted by the student model is low, and the first label recognition text can be selected as the label recognition text of the domain sample speech. If the number of characters per unit time length of the domain sample speech is greater than or equal to the character threshold, it indicates that the probability of insertion errors in the first label recognition text outputted by the student model is high, and the second label recognition text can be selected as the label recognition text of the domain sample speech.

[0109] It can be seen that the embodiment of the present invention can determine whether there is an insertion error in the first label recognition text based on the number of characters per unit time length and the character threshold of the domain sample speech, and then accurately determine the label recognition text of the domain sample speech from the first label recognition text and the second label recognition text.

[0110] Based on any of the above embodiments, step 120 inputs the speech features of the speech to be recognized into the speech recognition model to obtain the recognition text output by the speech recognition model, and then further includes:

[0111] Based on the number of characters in the recognized text and the speech duration of the speech to be recognized, the recognized text is corrected to obtain a corrected text.

[0112] Specifically, when the speech recognition model performs speech recognition on the domain sample speech, there may be insertion errors. For example, the speech to be recognized is "I love work", but the speech recognition model may have insertion errors in the recognition process, resulting in the obtained recognized text may be "I love love love love love love work", which obviously contains too many inserted words "love". Given that under normal circumstances there is an upper limit on the number of characters a user can speak per unit time, the number of characters per unit time of the speech to be recognized can be determined based on the number of characters in the recognized text and the speech duration of the speech to be recognized. If the number of characters per unit time is large, it means that the probability of insertion errors in the recognized text output by the speech model is higher, and it needs to be corrected to obtain the corrected text, so that the speech recognition result corresponding to the speech to be recognized can be obtained more accurately.

[0113] Based on any of the above embodiments, Figure 7 is a flow chart of the method for determining the corrected text provided by the present invention, such as Figure 7 As shown, based on the number of characters in the recognized text and the speech duration of the speech to be recognized, the recognized text is corrected to obtain a corrected text, including:

[0114] Step 710: Determine the number of characters per unit time length of the speech to be recognized based on the number of characters in the recognized text and the speech time length of the speech to be recognized;

[0115] Step 720: If the number of characters per unit time length of the speech to be recognized is greater than or equal to the character threshold, the speech features of the speech to be recognized are input into the general speech recognition model to obtain the general recognition text output by the general speech recognition model, and the general recognition text is used as the correction text;

[0116] If the number of characters per unit time length of the speech to be recognized is less than the character threshold, the recognized text is used as the corrected text;

[0117] The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

[0118] Specifically, the number of characters per unit time length of the speech to be recognized refers to the number of characters per unit time length corresponding to the recognized text obtained when the speech recognition model recognizes the speech to be recognized, which can be determined by the speech duration of the speech to be recognized and the number of characters in the recognized text, such as the number of characters per unit time length of the speech to be recognized = the number of characters in the recognized text / the speech duration of the speech to be recognized.

[0119] If the number of characters per unit time length of the speech to be recognized is less than the character threshold, it indicates that the probability of insertion errors in the recognized text output in the speech recognition model is low, and the recognized text can be selected as the correction text. If the number of characters per unit time length of the speech to be recognized is greater than or equal to the character threshold, it indicates that the probability of insertion errors in the recognized text output in the speech recognition model is high, and the general recognition text can be selected as the label recognition text of the domain sample speech. Among them, the character threshold can be set according to the actual situation, and the embodiment of the present invention does not make specific limitations on this.

[0120] In addition, the general speech recognition model is obtained by training based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model, that is, the structure of the general speech recognition model is different from that of the speech recognition model. For example, the structure of the student model can be an end-to-end model structure, and the general speech recognition model can be an acoustic model + language model structure, wherein the acoustic model can adopt a triphone model structure, and the speech model can adopt an N-Gram model structure, which is not specifically limited in the embodiment of the present invention.

[0121] It can be seen that the embodiment of the present invention can determine whether there are insertion errors in the recognized text based on the number of characters per unit time length of the speech to be recognized and the character threshold, and thus can accurately obtain the corrected text.

[0122] Based on any of the above embodiments, the present invention further provides a speech recognition method. Figure 8 FIG. 2 is a flow chart of the speech recognition method provided by the present invention. Figure 8 As shown, the method includes:

[0123] First, the speech features of the speech to be recognized are extracted. Then, the speech features of the speech to be recognized are respectively input into the speech recognition model and the general speech recognition model to obtain the recognition text output by the speech recognition model and the general recognition text output by the general speech recognition model.

[0124] Then, based on the number of characters in the recognized text and the speech duration of the speech to be recognized, the number of characters per unit duration of the speech to be recognized is determined. If the number of characters per unit duration is less than the character threshold, the recognized text is used as the recognition result of the speech to be recognized; if not, the general recognized text is used as the recognition result of the speech to be recognized. Among them, the structure of the speech recognition model is an end-to-end structure, and the structure of the general speech recognition model is an acoustic model + language model structure.

[0125] Specifically, the training steps of the speech recognition model include:

[0126] First, the speech features of the domain sample speech are disturbed, and the speech features of the disturbed domain sample speech are input into the student model to obtain the first recognition text output by the student model. At the same time, the speech features of the domain sample speech are input into the teacher model to obtain the second recognition text output by the teacher model. Then, based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text, the student model is iterated to obtain a speech recognition model. Among them, the initialization parameters of the student model are iteratively obtained based on the general sample speech and its label recognition text; the teacher model is trained based on the general sample speech and its label recognition text.

[0127] In addition, the steps of determining the label recognition text of the domain sample speech include:

[0128] First, the speech features of the domain sample speech are input into the student model to obtain the first label recognition text output by the student model. At the same time, the speech features of the domain sample speech are input into the general speech recognition model to obtain the second label recognition text output by the general speech recognition model. The general speech recognition model is trained based on the general sample speech and its label recognition text, and the general speech recognition model has a different structure from the student model.

[0129] If the number of characters per unit time length of the domain sample speech is less than the character threshold, the first label recognition text is used as the label recognition text; if not, the second label recognition text is used as the label recognition text of the domain sample speech. The number of characters per unit time length of the domain sample speech is determined based on the speech duration of the domain sample speech and the number of characters of the first label recognition text.

[0130] The speech recognition device provided by the present invention is described below. The speech recognition device described below and the speech recognition method described above can be referred to each other.

[0131] Based on any of the above embodiments, the present invention further provides a speech recognition device, Fig. 9 : is a schematic diagram of the structure of the speech recognition device provided by the present invention, such as Fig. 9 As shown, the device comprises:

[0132] A speech determination unit 910, configured to determine a speech to be recognized;

[0133] The speech recognition unit 920 is used to input the speech features of the speech to be recognized into the speech recognition model to obtain the recognition text output by the speech recognition model;

[0134] The speech recognition model is obtained by iterating parameters of the student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text;

[0135] The first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech. The teacher model is trained based on general sample speech and its label recognition text.

[0136] Based on any of the above embodiments, the device further includes:

[0137] A first recognition unit is used to perturb the speech features of the domain sample speech, and input the perturbed speech features of the domain sample speech into the student model to obtain the first recognition text output by the student model;

[0138] A second recognition unit, configured to input the speech features of the domain sample speech into the teacher model to obtain the second recognition text output by the teacher model;

[0139] A parameter iteration unit, configured to perform parameter iteration on the student model based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text, so as to obtain the speech recognition model;

[0140] The initialization parameters of the student model are iteratively obtained based on the general sample speech and its label recognition text.

[0141] Based on any of the above embodiments, the parameter iteration unit includes:

[0142] A first loss determination unit, configured to determine a first loss value based on a difference between the label recognition text of the domain sample speech and the first recognition text;

[0143] a second loss determining unit, configured to determine a second loss value based on a difference between the first recognized text and the second recognized text;

[0144] An iterative subunit is used to iterate the parameters of the student model based on the first loss value and the second loss value to obtain the speech recognition model.

[0145] Based on any of the above embodiments, the device further includes:

[0146] A first label recognition unit, used for inputting the speech features of the domain sample speech into the student model to obtain a first label recognition text output by the student model;

[0147] A second label recognition unit, used for inputting the speech features of the domain sample speech into a general speech recognition model to obtain a second label recognition text output by the general speech recognition model;

[0148] a label determination unit, configured to determine the label recognition text of the domain sample speech from the first label recognition text and the second label recognition text based on the speech duration of the domain sample speech;

[0149] The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

[0150] Based on any of the above embodiments, the label determination unit includes:

[0151] A first character number determination unit, configured to determine the number of characters per unit time length of the domain sample speech based on the speech time length of the domain sample speech and the number of characters of the first label recognition text;

[0152] The label determination subunit is used to use the first label recognition text as the label recognition text if the number of characters per unit time length of the domain sample speech is less than a character threshold; if not, use the second label recognition text as the label recognition text of the domain sample speech.

[0153] Based on any of the above embodiments, the device further includes:

[0154] The correction unit is used to input the speech features of the speech to be recognized into the speech recognition model to obtain the recognized text output by the speech recognition model, and then correct the recognized text based on the number of characters in the recognized text and the speech duration of the speech to be recognized to obtain the corrected text.

[0155] Based on any of the above embodiments, the correction unit includes:

[0156] A second character number determination unit, configured to determine the number of characters per unit time length of the speech to be recognized based on the number of characters in the recognized text and the speech time length of the speech to be recognized;

[0157] a correction text determination unit, configured to input the speech features of the speech to be recognized into a general speech recognition model if the number of characters per unit time length of the speech to be recognized is greater than or equal to a character threshold, obtain a general recognition text output by the general speech recognition model, and use the general recognition text as the correction text;

[0158] If the number of characters per unit time length of the speech to be recognized is less than the character threshold, the recognized text is used as the corrected text;

[0159] The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

[0160] Fig.10 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Fig.10 As shown, the electronic device may include: a processor 1010, a memory 1020, a communication interface 1030 and a communication bus 1040, wherein the processor 1010, the memory 1020 and the communication interface 1030 communicate with each other through the communication bus 1040. The processor 1010 may call the logic instructions in the memory 1020 to execute the speech recognition method, which includes: determining the speech to be recognized; inputting the speech features of the speech to be recognized into the speech recognition model to obtain the recognition text output by the speech recognition model; the speech recognition model is obtained by performing parameter iteration on the student model based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text; the first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech, and the teacher model is trained based on the general sample speech and its label recognition text.

[0161] In addition, the logic instructions in the above-mentioned memory 1020 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0162] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the speech recognition method provided by the above-mentioned methods, and the method includes: determining a speech to be recognized; inputting the speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model; the speech recognition model is obtained by performing parameter iteration on a student model based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text; the first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech, and the teacher model is trained based on the general sample speech and its label recognition text.

[0163] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the above-mentioned speech recognition methods, the methods comprising: determining a speech to be recognized; inputting speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model; the speech recognition model is obtained by performing parameter iteration on a student model based on the difference between a label recognition text of a domain sample speech and a first recognition text, and the difference between the first recognition text and the second recognition text; the first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech, and the teacher model is trained based on a general sample speech and its label recognition text.

[0164] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0165] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech recognition method, characterized in that: include: Determine the speech to be recognized; Inputting the speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model; The speech recognition model is obtained by iterating parameters of a student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text; The first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech. The teacher model is trained based on general sample speech and its label recognition text; the general sample speech is speech collected in a general scenario, and the domain sample speech is speech collected in a domain scenario.

2. The speech recognition method according to claim 1, characterized in that: The training steps of the speech recognition model include: perturbing the speech features of the domain sample speech, and inputting the perturbed speech features of the domain sample speech into the student model to obtain the first recognized text output by the student model; Inputting the speech features of the domain sample speech into the teacher model to obtain the second recognition text output by the teacher model; Based on the difference between the label recognition text of the domain sample speech and the first recognition text, and the difference between the first recognition text and the second recognition text, iterate the parameters of the student model to obtain the speech recognition model; The initialization parameters of the student model are iteratively obtained based on the general sample speech and its label recognition text.

3. The speech recognition method according to claim 2, characterized in that: The method of performing parameter iteration on the student model based on the difference between the label recognition text and the first recognition text and the difference between the first recognition text and the second recognition text to obtain the speech recognition model includes: Determining a first loss value based on a difference between the labeled recognition text of the domain sample speech and the first recognition text; determining a second loss value based on a difference between the first recognized text and the second recognized text; Based on the first loss value and the second loss value, parameters of the student model are iterated to obtain the speech recognition model.

4. The speech recognition method according to claim 1, characterized in that: The step of determining the label recognition text of the domain sample speech includes: Inputting the speech features of the domain sample speech into the student model to obtain a first label recognition text output by the student model; Inputting the speech features of the domain sample speech into a general speech recognition model to obtain a second label recognition text output by the general speech recognition model; Determining a label recognition text of the domain sample speech from the first label recognition text and the second label recognition text based on the speech duration of the domain sample speech; The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

5. The speech recognition method according to claim 4, characterized in that: The step of determining the label recognition text of the domain sample speech from the first label recognition text and the second label recognition text based on the speech duration of the domain sample speech includes: Determining the number of characters per unit time length of the domain sample speech based on the speech time length of the domain sample speech and the number of characters of the first label recognition text; If the number of characters per unit time length of the domain sample speech is less than the character threshold, the first label recognition text is used as the label recognition text; if not, the second label recognition text is used as the label recognition text of the domain sample speech.

6. The speech recognition method according to any one of claims 1 to 5, characterized in that: The step of inputting the speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model further includes: Based on the number of characters in the recognized text and the speech duration of the speech to be recognized, the recognized text is corrected to obtain a corrected text.

7. The speech recognition method according to claim 6, characterized in that: The step of correcting the recognized text based on the number of characters in the recognized text and the speech duration of the speech to be recognized to obtain a corrected text includes: Determining the number of characters per unit time length of the speech to be recognized based on the number of characters in the recognized text and the speech time length of the speech to be recognized; If the number of characters per unit time length of the speech to be recognized is greater than or equal to the character threshold, the speech features of the speech to be recognized are input into a universal speech recognition model to obtain a universal recognition text output by the universal speech recognition model, and the universal recognition text is used as the correction text; If the number of characters per unit time length of the speech to be recognized is less than the character threshold, the recognized text is used as the corrected text; The general speech recognition model is trained based on general sample speech and its label recognition text, and the structure of the general speech recognition model is different from that of the student model.

8. A speech recognition device, characterized in that: include: A speech determination unit, used to determine the speech to be recognized; A speech recognition unit, used for inputting the speech features of the speech to be recognized into a speech recognition model to obtain a recognition text output by the speech recognition model; The speech recognition model is obtained by iterating parameters of a student model based on the difference between the label recognition text and the first recognition text of the domain sample speech, and the difference between the first recognition text and the second recognition text; The first recognition text is determined by the student model based on the speech features of the domain sample speech, and the second recognition text is determined by the teacher model based on the speech features of the domain sample speech. The teacher model is trained based on general sample speech and its label recognition text; the general sample speech is speech collected in a general scenario, and the domain sample speech is speech collected in a domain scenario.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the speech recognition method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Speech recognition method and system, electronic equipment and storage medium

    CN111613212A

  • Voice recognition model training method and device, and voice recognition method and device

    CN111754985A

  • Model distillation method and device and storage medium

    CN114090727A