Stuttering speech recognition method and device, computer device and storage medium

By acquiring initial text information, utilizing a target stuttering predictor and a preset speech recognition model, combined with a preset loss formula and a selector, the problem of insufficient accuracy in stuttering speech recognition in existing technologies is solved, thereby improving the accuracy and resource efficiency of stuttering speech recognition.

CN119229871BActive Publication Date: 2025-12-26PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411237904.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-12-26
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

Existing pre-defined ASR models have low accuracy in recognizing stuttering speech, and the lack of effective stuttering datasets leads to insufficient recognition accuracy.

Method used

By acquiring initial text information, a target stuttering predictor is used to predict and generate target text information, which is then input into a preset speech recognition model for recognition. The text information is classified and processed by a preset loss formula and a preset selector, thereby improving the accuracy of stuttering speech recognition.

Benefits of technology

It improves the accuracy of stuttering speech recognition, reduces resource waste by generating and predicting target text information, and achieves controllability of the amount of target text information and accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229871B_ABST
    Figure CN119229871B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and provides a stuttering speech recognition method and device, computer equipment and a storage medium, the method comprising the following steps: acquiring initial text information for recognizing target stuttering speech; inputting the initial text information into a target stuttering predictor for prediction, generating target text information according to a prediction result of the target stuttering predictor; and inputting the target text information into a preset speech recognition model, recognizing the target stuttering speech corresponding to the target text information in the preset speech recognition model, so that the accuracy of stuttering speech recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a stuttering speech recognition method and device, a computer device and a storage medium. BACKGROUND

[0002] In the technical field of artificial intelligence, data imbalance is a common problem, especially in the specific application of speech recognition (ASR) and text-to-speech (TTS) technologies. Due to the scarcity of stuttering data sets, the existing preset ASR model has insufficient recognition accuracy for stuttering speech, resulting in low recognition accuracy of stuttering speech. Therefore, how to predict stuttering events and improve the accuracy of stuttering speech recognition has become a technical problem to be solved. SUMMARY

[0003] The main purpose of the present application is to provide a stuttering speech recognition method, device, computer device and storage medium to improve the accuracy of stuttering speech recognition.

[0004] In a first aspect, the present application provides a stuttering speech recognition method, which comprises the following steps:

[0005] obtaining initial text information for recognizing target stuttering speech; inputting the initial text information into a target stuttering predictor for prediction, generating target text information according to the prediction result of the target stuttering predictor; inputting the target text information into a preset speech recognition model, and recognizing the target stuttering speech corresponding to the target text information in the preset speech recognition model.

[0006] In a second aspect, the present application further provides a stuttering speech recognition device, which comprises:

[0007] An initial text information determination module is configured to obtain initial text information for recognizing target stuttering speech; a prediction module is configured to input the initial text information into a target stuttering predictor for prediction, and generate target text information according to the prediction result of the target stuttering predictor; and a target speech recognition module is configured to input the target text information into a preset speech recognition model, and recognize the target stuttering speech corresponding to the target text information in the preset speech recognition model.

[0008] In a third aspect, the present application further provides a computer device, which comprises a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the stuttering speech recognition method as described above is implemented.

[0009] In a fourth aspect, the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the stammer speech recognition method is implemented.

[0010] The present application provides a stammer speech recognition method, device, computer equipment and storage medium. The initial text information for identifying the target stammer speech is obtained. The initial text information is input into a target stammer predictor for prediction, and the target text information is generated according to the prediction result of the target stammer predictor. The target text information is input into a preset speech recognition model, and the target stammer speech corresponding to the target text information is identified in the preset speech recognition model. The initial text information is input into the target stammer predictor, the target prediction category is obtained according to the prediction result of the target stammer predictor, the target prediction category is added to the preset position of the initial text information, the target text information is obtained, the target stammer speech is input into the preset speech recognition model, and the target stammer speech is obtained. The accuracy of stammer speech recognition is improved. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0012] Figure 1 A flowchart of a stammer speech recognition method provided by an embodiment of the present application.

[0013] Figure 2 A schematic flowchart of determining target text information provided by an embodiment of the present application.

[0014] Figure 3 A schematic diagram of a preset speech recognition model provided by an embodiment of the present application.

[0015] Figure 4 Another schematic flowchart of determining target text information provided by an embodiment of the present application.

[0016] Figure 5 A schematic diagram of a stammer speech recognition device provided by an embodiment of the present application.

[0017] Figure 6 A structural schematic block diagram of a computer equipment provided by an embodiment of the present application. DETAILED DESCRIPTION

[0018] With reference to the drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.

[0019] The flowcharts shown in the drawings are only illustrative, and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so that the actual execution order can be changed according to the actual situation.

[0020] The embodiments of the present application provide a stuttering speech recognition method and device, computer equipment and a storage medium.

[0021] Some embodiments of the present application will be described in detail below with reference to the drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other.

[0022] Please refer to Figure 1 , Figure 1 A flowchart of a stuttering speech recognition method provided by an embodiment of the present application is shown. The stuttering speech recognition method can be used in a terminal or a server to improve the recognition accuracy of stuttering speech. The terminal can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant and a wearable device. The server can be a standalone server, a server cluster, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, etc. basic cloud computing services. As shown in Figure 1 The stuttering speech recognition method specifically includes steps S101 to S103.

[0023] S101, obtaining initial text information for recognizing target stuttering speech.

[0024] The number of initial text information is at least one.

[0025] Specifically, the type of the obtained initial text information can be four types of initial text: initial text information with sound prolongation, initial text information with pause category, initial text information with both sound prolongation and pause, and normal text information (non-pause category).

[0026] It should be noted that other suitable text categories can also be selected in actual applications, or some categories can be added or deleted according to actual application needs, and the present application does not make any limitation in this regard.

[0027] For example, the text information with sound prolongation can be: Today, the weather is very good, let's go to the park together. The text information with pause can be: Today... the weather is very good... let's go to the park... The text information with both sound prolongation and pause can be: Today, the weather is very good, let's go to the park.

[0028] It should be noted that the different types of initial text information obtained above are only used for illustrative purposes, and other arbitrary suitable initial text information can also be obtained according to actual needs in actual applications, and the present application does not make any limitation in this regard.

[0029] S102, input the initial text information into the target stuttering predictor for prediction, and generate target text information according to the prediction result of the target stuttering predictor.

[0030] Specifically, before inputting the initial text information into the target stuttering predictor, the text format of the initial text information is modified according to the text type that can be accepted by the target stuttering predictor, so as to ensure that the target stuttering predictor can be compatible with the initial text information. The text type that can be accepted by the target stuttering predictor can be any one of.txt file,.doc,.docx.

[0031] It should be noted that the text type that can be accepted by the target stuttering predictor is only used for illustrative purposes, and any other suitable text type can also be selected according to specific conditions in actual applications, and the present application does not make any limitation in this regard.

[0032] For example, the initial text information can be input into the target stuttering predictor based on the application programming interface (Application Programming Interface, API) of the target stuttering predictor. The API of the target stuttering predictor can be any suitable API type such as Web API, Library API, operating system API, database API, and the present application does not make any limitation in this regard.

[0033] Specifically, after inputting the initial text information into the target stuttering predictor for prediction, the target text information is generated according to the prediction result of the target stuttering predictor.

[0034] In some embodiments, the initial text information includes first text information and second text information, and before adding the target prediction category to the preset position of the initial text information to generate the target text information, the method further includes: performing classification processing on all of the initial text information based on a preset selector to obtain the first text information and the second text information classified by the preset selector; and adding the target prediction category corresponding to the first text information to the preset position of the first text information to generate the target text information corresponding to the first text information.

[0035] In some embodiments, the number of the first text information and the number of the second text information are both not greater than the number of the initial text information.

[0036] For example, the preset selector can be a decision tree, a support vector machine, a naive Bayes classifier, etc.

[0037] It should be noted that the above-mentioned preset selectors are only used for illustrative purposes, and other suitable preset selectors can be selected according to actual needs in actual applications, and the present application does not limit this.

[0038] After the initial text information is classified into the first text information and the second text information according to the preset selector, the target prediction category corresponding to each first text information is obtained, and the target prediction category corresponding to each first text information is added to the preset position of the corresponding first text information to obtain the target text information corresponding to the first text information. No related processing of adding the target prediction category is performed on the obtained second text information.

[0039] The above-mentioned processing method based on the preset selector only determines the target text information corresponding to the first text information, and does not perform any operation on the second text information, which realizes the controllability of the number of obtained target text information, can generate a specific number of target text information according to actual needs, avoids the waste of resources for generating target text information due to the excessive number of generated target text information, and improves the efficiency of generating target text information.

[0040] It should be noted that the preset position can be any suitable position in the first text information, such as the front end of the first text information, the middle of the first text information, the end of the first text information, etc. The specific position of adding the target prediction category in the first text information can be reasonably set according to actual conditions, and the present application does not limit this.

[0041] In some embodiments, the preset selector classifies all of the initial text information to obtain the first text information and the second text information after classification, including: determining a first parameter; if the first parameter is greater than or equal to the number of all of the initial text information, the preset selector determines all of the initial text information as the first text information; if the first parameter is less than the number of all of the initial text information, the preset selector determines the first initial text information to the initial text information at the first parameter position as the first text information, and determines the remaining initial text information as the second text information.

[0042] In the formula, the preset selector can also be a keyword selector, a regular expression selector, a sentence pattern selector, a theme selector, or any other suitable selector, and the present application does not limit the preset selector.

[0043] For example, if the preset selector is a keyword selector, the initial text information is input into the keyword selector, and after selection by the keyword selector, the initial text information can be determined as any one of the text information with repeated words, the text information with sound prolongation, or the text information with pauses. For example, the initial text information is: I want to eat an apple. After inputting the initial text information into the keyword selector, the initial text information is classified into the text information with sound prolongation after selection by the keyword selector. Similarly, each initial text information is input into the keyword selector in turn, and each initial text information is classified into the corresponding text type by the keyword selector. The specific classification process of the keyword selector for each initial text information is the same as the above type, and the present application does not repeat it here.

[0044] For example, the first parameter can be represented by the letter N, and the specific value of the first parameter N can be set to any suitable positive integer such as 10, 20, 100, etc., and the present application does not limit the first parameter.

[0045] It should be noted that the first parameter can also be represented by the letters T, X, etc., and any other suitable letter can be selected according to the actual application, and the present application does not limit the first parameter.

[0046] For example, if the number of all the initial text information is 20, N = 30, the first parameter is greater than the number of all the initial text information, the preset selector determines all the initial text information as the first text information. Then, the target prediction category corresponding to each first text information is added to the preset position of the first text information to generate the target text information corresponding to the first text information. Similarly, when the first parameter is equal to the number of all the initial text information, all the initial text information is also determined as the first text information, and then the target text information corresponding to all the first text information is obtained.

[0047] For example, if the number of all the initial text information is 20, N = 15, the first parameter is less than the number of all the initial text information.

[0048] The preset selector determines the first initial text information to the initial text information at the first parameter position as the first text information, and determines the remaining initial text information as the second text information. For example, all the initial text information is: initial text information 1, initial text information 2, initial text information 3, …, initial text information 20, the first initial text information 1 to the initial text information 15 at the first parameter position are determined as the first text information, that is, the first 15 initial text information, that is, initial text information 1, initial text information 2, initial text information 3, …, initial text information 15, are determined as the first text information. The remaining initial text information, that is, initial text information 16, initial text information 17, …, initial text information 20, are determined as the second text information.

[0049] For another example, if the categories of the first five target text information are all text information with repeated words, and the subsequent three categories of initial text information are target text information, at this time, only the first five target text information with repeated words are analyzed, the first parameter N can be set to 5. Similarly, in other cases, the first parameter can also be set to other appropriate values, which are not limited in the application.

[0050] In summary, according to the set first parameter N, the first N initial text information is selected from the starting position, the first N initial text information is determined as the first text information, the target text information corresponding to the first text information is determined, and the remaining initial text information is determined as the second text information. When the number of initial text information is less than N, all the initial text information is directly determined as the first text information.

[0051] It should be noted that the purpose of setting the first parameter is to achieve controllability of the quantity of target text information.

[0052] Please refer to Figure 3 , Figure 3 is a schematic diagram of a preset speech recognition model in an embodiment of the present application. In some embodiments, the preset speech recognition model includes a front-end model and an acoustic model. The method of inputting the target text information into the preset speech recognition model and recognizing the target stuttering speech corresponding to the target text information in the preset speech recognition model includes: inputting the target text information into the front-end model to obtain target speech information to be generated output by the front-end model; and inputting the target speech information to be generated into the acoustic model, and performing speech recognition in the acoustic model to obtain the target stuttering speech.

[0053] wherein the target speech information to be generated is speech information containing phonemes.

[0054] For example, "Today the weather is very hot, the sky is clear and cloudless" can be "Today the weather is very hot, the sky is clear and cloudless" after processing by the front-end model, and then converted into corresponding target speech information to be generated in the front-end model, and the target speech information to be generated is input into the acoustic model for speech recognition to obtain the corresponding target stuttering speech. The target stuttering speech can be finally obtained through the joint training of the above-mentioned front-end model and acoustic model.

[0055] wherein the above-mentioned front-end model can be any one of a probabilistic language model, a rule-based language model, a speech rate predictor, etc., and the above-mentioned acoustic model can be any one of fastspeech, tracotron, WaveNet.

[0056] It should be noted that the specific types of the above-mentioned front-end model and acoustic model are only used for illustrative purposes, and other suitable models can also be selected in actual applications, which are not limited by the present application.

[0057] In some embodiments, before inputting the initial text information into the target stuttering predictor for prediction, the method further includes: training a prediction module in the target stuttering predictor based on a preset loss formula to determine a prediction weight of the preset loss formula.

[0058] For example, the preset loss formula is:

[0059]

[0060] wherein, L represents a loss value, y1 represents the first predicted category, y2 represents the second predicted category, and σ represents a weight of the second predicted category, wherein the first predicted category is text with pauses, and the second predicted category is non-pause text, represents the first predicted probability, represents the second predicted probability.

[0061] For example, if the second predicted category is less, in order to balance the loss of the first predicted category and the second predicted category and prevent the stuttering predictor from being biased to the first predicted category, the value of σ can be increased to increase the penalty for prediction errors of the second predicted category.

[0062] In another example, the initial text information with pause and prolongation labels in the present scheme is less, which is a sparse label. The loss formula can solve the problem of unbalanced initial text information, so as to realize the penalty for prediction errors of the initial text information with pause and prolongation labels corresponding to the target text information.

[0063] It should be noted that the above preset loss formula only considers two cases of pause and non-pause. In actual application, any combination of two cases can be selected according to actual needs, and the present application does not limit this.

[0064] S103, inputting the target text information into a preset speech recognition model, and identifying the target stuttering speech corresponding to the target text information in the preset speech recognition model.

[0065] Specifically, the interface type or text type of the target text information can be determined for the preset speech recognition model, and the target text information can be converted into a corresponding text type. For example, the interface type of the preset speech recognition model can be a command line interface, an application programming interface, a Web service, etc. The target text information can be input into the preset speech recognition model based on the interface or API of the preset speech recognition model.

[0066] In some embodiments, after the target text information is input into the preset speech recognition model to obtain the target stuttering speech corresponding to the target text information, the method further comprises: inputting the target stuttering speech into the preset speech recognition model, training the preset speech recognition model to obtain a target speech recognition model corresponding to the preset speech recognition model.

[0067] Specifically, the obtained multiple target stuttering speeches are sequentially input into the preset speech recognition model, and the preset speech recognition model is trained based on the target stuttering speech to improve the recognition accuracy of the preset speech recognition model.

[0068] Exemplarily, the preset speech recognition model can be any one of Fastspeech2 or Tactron.

[0069] It should be noted that the Fastspeech2 and Tactron given in the above preset speech recognition model are only exemplary, and in actual application, any other suitable preset speech recognition model can also be selected according to specific conditions, and the present application does not limit this.

[0070] In the stuttering speech recognition method provided by the embodiments of the present application, initial text information for identifying target stuttering speech is obtained; the initial text information is input into a target stuttering predictor for prediction, and target text information is generated according to the prediction result of the target stuttering predictor; and the target text information is input into a preset speech recognition model, and the target stuttering speech corresponding to the target text information is identified in the preset speech recognition model. By inputting the obtained target text information into a preset loss formula, the prediction error of the target text information corresponding to the initial text information with pause and prolongation labels is punished, and the accuracy of the prediction of the initial text information with pause and prolongation labels is improved. Before obtaining the target text information, the initial text information is input into a preset selector to classify a plurality of initial text information, obtain the target text information corresponding to the first text information after classification of the initial text information, and realize the controllability of the number of generated target text information, and improve the efficiency of generating target text information.

[0071] Please refer to Figure 2 , Figure 2 An exemplary flowchart for determining target text information is provided for an embodiment of the present application. As Figure 2 shown, the stuttering speech recognition method specifically includes steps S201 to S205.

[0072] S201, determining a first prediction category, a second prediction category, a third prediction category and a fourth prediction category.

[0073] Specifically, the first prediction category is initial text information corresponding to a pause category; the second prediction category is initial text information corresponding to a non-pause category; the third prediction category is initial text information with sound prolongation; and the fourth prediction category is initial text information with both sound prolongation and pause.

[0074] S202, inputting the initial text information into the target stuttering predictor to obtain a first prediction probability corresponding to the first prediction category, a second prediction probability corresponding to the second prediction category, a third prediction probability corresponding to the third prediction category, and a fourth prediction probability corresponding to the fourth prediction category.

[0075] Specifically, each time an initial text information is input into the target stuttering predictor, the initial text information is predicted in the target stuttering predictor to obtain specific values of a first prediction probability corresponding to the first prediction category, a second prediction probability corresponding to the second prediction category, a third prediction probability corresponding to the third prediction category, and a fourth prediction probability corresponding to the fourth prediction category.

[0076] For example, if the initial text information is: I want to eat ~ ~ apple, after the initial text information is input into the target stuttering predictor, the first prediction probability corresponding to the first prediction category (i.e., the initial text information corresponding to the pause category) is 30%; the second prediction probability corresponding to the second prediction category (i.e., the initial text information corresponding to the non-pause category) is 50%; the third prediction probability corresponding to the third prediction category (i.e., the initial text information with sound prolongation) is 90%; and the fourth prediction probability corresponding to the fourth prediction category (i.e., the initial text information with both sound prolongation and pause) is 40%. The specific steps of inputting other initial text information into the target stuttering predictor for prediction are similar to the steps given in the above embodiment, and will not be repeated here.

[0077] It should be noted that the initial text information in the above embodiment and the specific values of the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability corresponding to the initial text information are only used for illustrative purposes. In actual applications, the specific values of the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability can be obtained according to specific application conditions and specific prediction conditions of the target stuttering predictor, and the present application does not limit this.

[0078] S203, taking the maximum value of the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability as a target prediction probability.

[0079] Specifically, after obtaining the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability corresponding to the initial text information, the maximum value of the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability is taken as the target prediction probability.

[0080] For example, if the first prediction probability corresponding to a certain initial text information is 20%, the second prediction probability is 60%, the third prediction probability is 30%, and the fourth prediction probability is 40%, the second prediction probability is taken as the target prediction probability. Similarly, the target prediction probability corresponding to other initial text information can be obtained, and will not be repeated here.

[0081] S204, taking the prediction category corresponding to the target prediction probability as the target prediction category corresponding to the initial text information.

[0082] For example, if the second prediction probability has the maximum value among the first prediction probability, the second prediction probability, the third prediction probability and the fourth prediction probability, the second prediction probability is the target prediction probability. If the initial text information corresponding to the prediction category of the second prediction probability is the non-pause category, the target prediction category corresponding to the initial text information is the non-pause category. The target prediction category corresponding to other initial text information can be obtained in a similar manner, which will not be described herein.

[0083] S205, adding the target prediction category to the preset position of the initial text information to generate the target text information.

[0084] The preset position can be the front end of the initial text information, the middle of the initial text information or the rear end of the initial text information, or any other suitable position, which will not be described herein.

[0085] For example, after determining the target prediction category corresponding to the initial text information, the target prediction category is added to the preset position of the initial text information to generate the target text information. For example, the initial text information is: I want to eat rice. After prediction by the target stuttering predictor, the target prediction category is the first prediction category (i.e., the initial text information of the pause category), and the target prediction category is added to the front end of the initial text information (assuming that the preset position is the front end), and the target text information is obtained as: [pause category] I want to eat rice. For example, the initial text information is: I want to eat rice. After prediction by the target stuttering predictor, the target prediction category is the third prediction category (i.e., the initial text information with sound prolongation), and the target prediction category is added to the front end of the initial text information (assuming that the preset position is the front end), and the target text information is obtained as: [sound prolongation category] I want to eat rice. Similarly, when the target prediction category is the third prediction category or the fourth prediction category, the third prediction category or the fourth prediction category is added to the suitable preset position of the initial text information to obtain the corresponding target text information, which will not be described herein.

[0086] It should be noted that the above process of determining the target text information can also refer to the process or steps in Figure 4 , Figure 4 Another schematic flowchart for determining the target text information provided by an embodiment of the present application, Figure 4 The specific process in Figure 2 is similar, which will not be described herein.

[0087] The method for determining target text information provided in the embodiments of the present application comprises: determining a first predicted category, a second predicted category, a third predicted category and a fourth predicted category; inputting the initial text information into the target stuttering predictor to obtain a first predicted probability corresponding to the first predicted category, a second predicted probability corresponding to the second predicted category, a third predicted probability corresponding to the third predicted category and a fourth predicted probability corresponding to the fourth predicted category; taking the maximum value among the first predicted probability, the second predicted probability, the third predicted probability and the fourth predicted probability as a target predicted probability; taking the predicted category corresponding to the target predicted probability as a target predicted category corresponding to the initial text information; and adding the target predicted category to a preset position of the initial text information to generate the target text information. By determining the first predicted category, the second predicted category, the third predicted category and the fourth predicted category, and calculating the first predicted probability, the second predicted probability, the third predicted probability and the fourth predicted probability in the target stuttering predictor, the specific values of the first predicted probability, the second predicted probability, the third predicted probability and the fourth predicted probability are determined in a specific calculation manner, instead of relying on subjective feelings, so that the accuracy of the data is improved without being affected by other external factors. The maximum value among the first predicted probability, the second predicted probability, the third predicted probability and the fourth predicted probability is taken as the target predicted probability, and the target predicted category corresponding to the target predicted probability is determined. The target predicted category is added to the preset position of the initial text information to obtain the target text information. The target text information is obviously marked through the prediction of the target stuttering predictor, and an accurate text basis is provided for subsequent accurate generation of stuttering speech.

[0088] Please refer to Figure 5 , Figure 5 A schematic diagram of a stuttering speech recognition device provided in an embodiment of the present application. The stuttering speech recognition device can be configured in a server or a terminal, and is used to execute the stuttering speech recognition method described above.

[0089] As Figure 5 shown, the stuttering speech recognition device 500 comprises an initial text information determination module 501, a prediction module 502 and a target speech recognition module 503.

[0090] The initial text information determination module 501 is configured to acquire initial text information used to recognize target stuttering speech.

[0091] The prediction module 502 is configured to input the initial text information into a target stuttering predictor for prediction, and generate target text information according to the prediction result of the target stuttering predictor.

[0092] The target speech recognition module 503 is configured to input the target text information into a preset speech recognition model, and identify the target stuttering speech corresponding to the target text information in the preset speech recognition model.

[0093] In some embodiments, the initial text information determination module 501 is specifically configured to:

[0094] Obtain initial text information used for identifying a target stuttering speech.

[0095] In some embodiments, the prediction module 502 is specifically configured to:

[0096] Determine a first prediction category, a second prediction category, a third prediction category and a fourth prediction category; input the initial text information into the target stuttering prediction model to obtain a first prediction probability corresponding to the first prediction category, a second prediction probability corresponding to the second prediction category, a third prediction probability corresponding to the third prediction category and a fourth prediction probability corresponding to the fourth prediction category; take a maximum value among the first prediction probability, the second prediction probability, the third prediction probability and the fourth prediction probability as a target prediction probability; take a prediction category corresponding to the target prediction probability as a target prediction category corresponding to the initial text information; and add the target prediction category to a preset position of the initial text information to generate the target text information.

[0097] In some embodiments, the prediction module 502 is specifically further configured to:

[0098] Classify all the initial text information based on a preset selector to obtain the first text information and the second text information classified by the preset selector;

[0099] Add a target prediction category corresponding to the first text information to a preset position of the first text information to generate the target text information corresponding to the first text information.

[0100] In some embodiments, the prediction module 502 is specifically further configured to:

[0101] Determine a first parameter; if the first parameter is greater than or equal to a number of all the initial text information, the preset selector determines all the initial text information as the first text information; if the first parameter is less than the number of all the initial text information, the preset selector determines the first initial text information to the initial text information at the first parameter position as the first text information, and determines the remaining initial text information as the second text information.

[0102] In some embodiments, the prediction module 502 is specifically further configured to:

[0103] The prediction module in the target stuttering predictor is trained based on a preset loss formula to determine prediction weights of the preset loss formula.

[0104] In some embodiments, the target speech recognition module 503 is specifically configured to:

[0105] The target text information is input into the front-end model to obtain target speech information to be generated output by the front-end model, and the target speech information to be generated is input into the acoustic model to perform speech recognition in the acoustic model to obtain the target stuttering speech.

[0106] In some embodiments, the target speech recognition module 503 is specifically configured to:

[0107] The target stuttering speech is input into a preset speech recognition model, the preset speech recognition model is trained to obtain a target speech recognition model corresponding to the preset speech recognition model.

[0108] It should be noted that, for the convenience and brevity of description, the specific working processes of the above-described apparatus and modules can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0109] The apparatus described above can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 6 .

[0110] Please refer to Figure 6 , Figure 6 is a structural schematic block diagram of a computer device according to an embodiment of the present application. The computer device can be a server.

[0111] Referring to Figure 6 , the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a non-volatile storage medium and an internal memory.

[0112] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to perform any kind of stuttering speech recognition method.

[0113] The processor is configured to provide computing and control capabilities to support the operation of the entire computer device.

[0114] The internal memory provides an environment for the running of the computer program in the non-volatile storage medium, and the computer program, when executed by the processor, can cause the processor to perform any kind of stuttering speech recognition method.

[0115] The network interface is configured to perform network communication, such as sending the assigned task, etc. Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0116] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0117] In one embodiment, the processor is configured to run a computer program stored in the memory to perform the following steps:

[0118] obtaining initial text information for identifying target stuttering speech; inputting the initial text information into a target stuttering predictor for prediction, generating target text information according to the prediction result of the target stuttering predictor; inputting the target text information into a preset speech recognition model, and identifying the target stuttering speech corresponding to the target text information in the preset speech recognition model.

[0119] In one embodiment, when the initial text information is input into the target stuttering predictor for prediction, and the target text information is generated according to the prediction result of the target stuttering predictor, it is configured to implement:

[0120] determining a first prediction category, a second prediction category, a third prediction category and a fourth prediction category; inputting the initial text information into the target stuttering predictor to obtain a first prediction probability corresponding to the first prediction category, a second prediction probability corresponding to the second prediction category, a third prediction probability corresponding to the third prediction category and a fourth prediction probability corresponding to the fourth prediction category; taking the maximum value among the first prediction probability, the second prediction probability, the third prediction probability and the fourth prediction probability as a target prediction probability; taking the prediction category corresponding to the target prediction probability as a target prediction category corresponding to the initial text information; adding the target prediction category to a preset position of the initial text information to generate the target text information.

[0121] In one embodiment, the entire initial text information includes first text information and second text information, and before the target prediction category is added to the preset position of the initial text information to generate the target text information, the following is implemented:

[0122] performing classification processing on the entire initial text information based on a preset selector to obtain the first text information and the second text information classified by the preset selector; adding a target prediction category corresponding to the first text information to a preset position of the first text information to generate the target text information corresponding to the first text information.

[0123] In one embodiment, the classification processing on the entire initial text information based on the preset selector to obtain the first text information and the second text information classified by the preset selector is implemented as follows:

[0124] determining a first parameter; if the first parameter is greater than or equal to the number of the entire initial text information, the preset selector determines all the initial text information as the first text information; if the first parameter is less than the number of the entire initial text information, the preset selector determines the first initial text information to the initial text information at the first parameter position as the first text information, and determines the remaining initial text information as the second text information.

[0125] In one embodiment, the preset speech recognition model includes a front-end model and an acoustic model, and the inputting of the target text information into the preset speech recognition model and the recognition of the target text information corresponding to the target stuttering speech in the preset speech recognition model are implemented as follows:

[0126] input the target text information into the front-end model to obtain target speech information to be generated output by the front-end model; and input the target speech information to be generated into the acoustic model, and perform speech recognition in the acoustic model to obtain the target stuttering speech.

[0127] In one embodiment, before inputting the initial text information into the target stuttering predictor for prediction, the method is configured to:

[0128] Based on a preset loss formula, the prediction module in the target stuttering predictor is trained to determine the prediction weight of the preset loss formula.

[0129] In one embodiment, after inputting the target text information into the preset speech recognition model and identifying the target stuttering speech corresponding to the target text information in the preset speech recognition model, the method is configured to:

[0130] input the target stuttering speech into a preset speech recognition model, train the preset speech recognition model to obtain a target speech recognition model corresponding to the preset speech recognition model.

[0131] In the embodiments of the present application, a computer readable storage medium is also provided, which stores a computer program including program instructions. The processor executes the program instructions to implement any one of the stuttering speech recognition methods provided in the embodiments of the present application.

[0132] The computer readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0133] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A stuttering speech recognition method, characterized by, The method comprises: obtaining initial text information for identifying target stuttering speech; inputting the initial text information into a target stuttering predictor for prediction, and generating target text information according to a prediction result of the target stuttering predictor; inputting the target text information into a preset speech recognition model, and identifying the target stuttering speech corresponding to the target text information in the preset speech recognition model; wherein the inputting the initial text information into the target stuttering predictor for prediction, and generating the target text information according to the prediction result of the target stuttering predictor comprises: determining a first prediction category, a second prediction category, a third prediction category, and a fourth prediction category; inputting the initial text information into the target stuttering predictor to obtain a first prediction probability corresponding to the first prediction category, a second prediction probability corresponding to the second prediction category, a third prediction probability corresponding to the third prediction category, and a fourth prediction probability corresponding to the fourth prediction category; taking a maximum value among the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability as a target prediction probability; taking a prediction category corresponding to the target prediction probability as a target prediction category corresponding to the initial text information; adding the target prediction category to a preset position of the initial text information to generate the target text information; all initial text information includes first text information and second text information, and before the adding the target prediction category to the preset position of the initial text information to generate the target text information, the method further comprises: performing classification processing on all the initial text information based on a preset selector to obtain the first text information and the second text information classified by the preset selector; adding a target prediction category corresponding to the first text information to a preset position of the first text information to generate the target text information corresponding to the first text information; the performing classification processing on all the initial text information based on the preset selector to obtain the first text information and the second text information classified by the preset selector comprises: determining a first parameter; if the first parameter is greater than or equal to a number of all the initial text information, the preset selector determines all the initial text information as the first text information; if the first parameter is less than the number of all the initial text information, the preset selector determines the first initial text information to the initial text information at the first parameter position as the first text information, and determines the remaining initial text information as the second text information.

2. The stuttering speech recognition method of claim 1, wherein, the preset speech recognition model comprises a front-end model and an acoustic model, and the inputting the target text information into the preset speech recognition model to identify the target stuttering speech corresponding to the target text information in the preset speech recognition model comprises: inputting the target text information into the front-end model to obtain target speech information to be generated output by the front-end model; input the target speech information to be generated into the acoustic model, perform speech recognition in the acoustic model, and obtain the target stuttering speech.

3. The stuttering speech recognition method of claim 1, wherein, Before inputting the initial text information into the target stuttering predictor for prediction, the method further comprises: training a prediction module in the target stuttering predictor based on a preset loss formula to determine a prediction weight of the preset loss formula.

4. The stuttering speech recognition method of claim 1, wherein, After inputting the target text information into the preset speech recognition model and recognizing the target stuttering speech corresponding to the target text information in the preset speech recognition model, the method further comprises: inputting the target stuttering speech into a preset speech recognition model, training the preset speech recognition model, and obtaining a target speech recognition model corresponding to the preset speech recognition model.

5. A stuttering speech recognition apparatus characterized by comprising: comprises: an initial text information determination module configured to obtain initial text information used to recognize a target stuttering speech; a prediction module configured to input the initial text information into a target stuttering predictor for prediction, generate target text information according to a prediction result of the target stuttering predictor; a target speech recognition module configured to input the target text information into a preset speech recognition model, recognize the target stuttering speech corresponding to the target text information in the preset speech recognition model; wherein the inputting the initial text information into the target stuttering predictor for prediction and generating the target text information according to the prediction result of the target stuttering predictor comprises: determining a first prediction category, a second prediction category, a third prediction category, and a fourth prediction category; inputting the initial text information into the target stuttering predictor to obtain a first prediction probability corresponding to the first prediction category, a second prediction probability corresponding to the second prediction category, a third prediction probability corresponding to the third prediction category, and a fourth prediction probability corresponding to the fourth prediction category; taking a maximum value among the first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability as a target prediction probability; taking a prediction category corresponding to the target prediction probability as a target prediction category corresponding to the initial text information; adding the target prediction category to a preset position of the initial text information to generate the target text information; all initial text information includes first text information and second text information, and before adding the target prediction category to a preset position of the initial text information to generate the target text information, the apparatus is further configured to: classify all the initial text information based on a preset selector to obtain the first text information and the second text information classified by the preset selector; add a target prediction category corresponding to the first text information to a preset position of the first text information to generate the target text information corresponding to the first text information; the classifying all the initial text information based on the preset selector to obtain the first text information and the second text information classified by the preset selector comprises: determining a first parameter; If the first parameter is greater than or equal to the total number of the initial text information, the preset selector determines all of the initial text information as the first text information; If the first parameter is less than the total number of the initial text information, the preset selector determines the first initial text information to the initial text information at the first parameter position as the first text information, and determines the remaining initial text information as the second text information.

6. A computer device, comprising: The computer device comprises a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program and implement the stuttering speech recognition method according to any one of claims 1 to 4 when executing the computer program.

7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to enable the processor to implement the stuttering speech recognition method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Diagnosing and treatment of speech pathologies using analysis by synthesis technology

    US20210158834A1