Speech recognition method, device, computer equipment and storage medium
By introducing a third-party resource library into the speech recognition model for supplementary results, selecting the optimal text based on the domain confidence, solving the problem of low recognition accuracy of untrained speech content and improving the overall accuracy of speech recognition.
Patent Information
- Application Number
- CN202111483278.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-12-07
AI Technical Summary
The existing speech recognition model has low recognition accuracy on untrained speech content, resulting in insufficient speech recognition accuracy.
By obtaining the voice to be recognized, input a preset speech recognition model to obtain the speech recognition text, its content confidence and domain confidence, and determine whether the text confidence is lower than the threshold. If it is low, call the third-party resource library for the result supplement processing, and select the text with the highest domain confidence as the target speech recognition text.
The accuracy of speech recognition has been improved, especially for untrained speech content. Through supplementary processing of external resource libraries, the optimal text is selected as the recognition result, which improves the accuracy of recognition.
Smart Images

Figure CN114171023B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition method, apparatus, computer equipment, and storage medium. Background Art
[0002] In the IoT industry scenario, human-computer interaction, information search through engines, and terminal control all involve speech recognition technology. Speech is recognized through a speech recognition model, and after the corresponding speech text is identified, the information expression of the speech text is specifically understood through semantic understanding technology, so that the corresponding system can perform subsequent actions based on the understanding results.
[0003] Typically, a speech recognition model needs to be pre-trained before speech recognition is performed based on the trained speech recognition model. However, in many practical applications, the speech content to be recognized has not been trained into the model in advance, resulting in inaccurate speech recognition by the speech recognition model. The accuracy of speech recognition needs to be improved. Summary of the Invention
[0004] The embodiments of the present application provide a speech recognition method, apparatus, computer device, and storage medium, which can improve the accuracy of speech recognition.
[0005] In a first aspect, an embodiment of the present application provides a speech recognition method, which includes:
[0006] Get the speech to be recognized;
[0007] Inputting the speech to be recognized into a preset speech recognition model to obtain speech recognition text, content confidence and domain confidence of the speech recognition text;
[0008] Determining whether a text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence;
[0009] If the text confidence threshold of the speech recognition text is less than the confidence threshold, calling a third-party resource library to perform result supplementation processing on the speech recognition text to obtain a supplemented speech recognition text;
[0010] The speech text with the highest domain confidence in the supplemented speech recognition text is determined as the target speech recognition text.
[0011] In a second aspect, an embodiment of the present application further provides a speech recognition device, comprising:
[0012] An acquisition unit, configured to acquire the speech to be recognized;
[0013] An input unit, configured to input the speech to be recognized into a preset speech recognition model to obtain a speech recognition text, a content confidence level, and a domain confidence level of the speech recognition text;
[0014] a judging unit, configured to judge whether a text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence;
[0015] A supplementing unit, configured to, when a text confidence threshold of the speech recognition text is less than the confidence threshold, call a third-party resource library to perform result supplement processing on the speech recognition text to obtain a supplemented speech recognition text;
[0016] The first determining unit is configured to determine the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text.
[0017] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0018] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and the program instructions can implement the above method when executed by a processor.
[0019] The embodiments of the present application provide a speech recognition method, apparatus, computer device, and storage medium. The method comprises: obtaining a speech to be recognized; inputting the speech to be recognized into a preset speech recognition model to obtain a speech recognition text, content confidence, and domain confidence of the speech recognition text; determining whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence; if the text confidence threshold of the speech recognition text is less than the confidence threshold, calling a third-party resource library to perform result supplementation processing on the speech recognition text to obtain a supplemented speech recognition text; and determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text. In an embodiment of the present application, when the confidence of the text recognized by the speech recognition model is relatively low, it is necessary to call a third-party resource library to supplement the results of the speech recognition text, and then select the text with the highest domain confidence from the supplemented speech recognition text as the target speech recognition text. It can be seen that when the accuracy of the speech recognition text obtained according to the preset speech recognition model is not high, the speech recognition text can be supplemented by other databases, and then the optimal text can be selected from the supplemented text as the target speech recognition text. Therefore, this solution can improve the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A schematic diagram of an application scenario of the speech recognition method provided in an embodiment of the present application;
[0022] Figure 2 A flowchart of a speech recognition method provided in an embodiment of the present application;
[0023] Figure 3 A flowchart of a speech recognition method provided in another embodiment of the present application;
[0024] Figure 4 A schematic block diagram of a speech recognition device provided in an embodiment of the present application;
[0025] Figure 5 A schematic block diagram of a speech recognition device provided in another embodiment of the present application;
[0026] Figure 6 A schematic block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0028] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0029] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0030] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0031] Embodiments of the present application provide a speech recognition method, apparatus, computer device, and storage medium.
[0032] The executor of the speech recognition method may be the speech recognition device provided in the embodiment of the present application, or a computer device in which the speech recognition device is integrated. The speech recognition device may be implemented in hardware or software, and the computer device may be a terminal or a server. The terminal may be a smart phone, a tablet computer, a PDA, a laptop computer, a smart speaker, etc.
[0033] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of the speech recognition method provided in the embodiment of the present application. The speech recognition method is applied to Figure 1 In the computer device 10, the computer device 10 obtains the user's speech to be recognized; inputs the speech to be recognized into a preset speech recognition model to obtain a speech recognition text, content confidence and domain confidence of the speech recognition text; judges whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence; if the text confidence threshold of the speech recognition text is less than the confidence threshold, calls a third-party resource library to perform result supplement processing on the speech recognition text to obtain a supplemented speech recognition text; and determines the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text.
[0034] Figure 2 This is a flow chart of the speech recognition method provided in the embodiment of the present application. Figure 2 As shown, the method includes the following steps S110-160.
[0035] S110: Acquire the speech to be recognized.
[0036] In this embodiment, the user can wake up the computer device, and then input the voice to be recognized into the computer device by speaking, so that the computer device acquires the voice to be recognized.
[0037] S120: Input the speech to be recognized into a preset speech recognition model to obtain speech recognition text, content confidence and domain confidence of the speech recognition text.
[0038] After the computer device obtains the voice to be recognized, it inputs the voice to be recognized into a pre-trained voice recognition model, and then outputs the voice recognition text with the highest content confidence, the content confidence corresponding to the voice recognition text, and the domain confidence through the voice recognition model. Among them, the content confidence can reflect the recognition accuracy of the voice recognition text, and the domain confidence can reflect which service fields the voice recognition text belongs to. Among them, the service field includes the voice interaction field and the terminal control field, among others. Among them, the terminal control field specifically includes the air conditioning control field, the curtain control field, the TV control field, and the like.
[0039] It should be noted that the speech recognition model in this embodiment can not only realize the function of converting speech into text, but also determine the domain confidence of the domain to which the recognized text belongs. In some embodiments, the speech recognition model can be a convolutional neural network model, such as an N-gram neural network model.
[0040] S130. Determine whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence. If so, execute steps S140-S150; if not, execute step S160.
[0041] Wherein, step S130 includes: determining the text confidence according to the content confidence, the preset content confidence weight, the domain confidence and the preset domain confidence weight; and then judging whether the text confidence is less than the confidence threshold.
[0042] Specifically, in some embodiments, text confidence = content confidence * preset content confidence weight + domain confidence * preset domain confidence weight.
[0043] It should be noted that the domain confidence in the speech recognition model in this embodiment includes multiple domain sub-confidences. This embodiment needs to calculate the text confidence based on the domain sub-confidence with the highest value among the multiple domain sub-confidences.
[0044] S140: Call a third-party resource library to perform result supplementation processing on the speech recognition text to obtain supplemented speech recognition text.
[0045] When the text confidence threshold of the speech recognition text is less than the confidence threshold, it can be considered that there is no corresponding recognition result in the speech recognition model and the recognition result accuracy is relatively low. In this case, it is necessary to call a third-party resource library to supplement the speech recognition text results, where the third-party resource library includes an Internet search engine, a social network search interface, and / or multiple domain resource libraries.
[0046] In some embodiments, when the third-party resource library includes multiple domain resource libraries, step S140 includes: obtaining the text pinyin corresponding to the speech recognition text; determining the target domain resource library from the multiple domain resource libraries based on the domain confidence; and finally determining the supplemented speech recognition text in the target domain resource library based on the text pinyin.
[0047] Wherein, the domain confidence includes a plurality of domain sub-confidences, and the determining of the target domain resource library from the plurality of domain resource libraries according to the domain confidence includes: extracting a preset number of confidences with the largest values from the plurality of domain sub-confidences to obtain a plurality of target domain sub-confidences; determining the domain resource library corresponding to each of the target domain sub-confidences from the plurality of domain resource libraries to obtain a plurality of target domain resource libraries;
[0048] For example, if the preset number is 3, it is necessary to extract the top 3 confidences with the largest values from multiple domain sub-confidences as the target domain sub-confidences, where the domain sub-confidence carries the domain label of the corresponding domain. This embodiment will determine the domain corresponding to the domain sub-confidence based on the domain label carried by the domain sub-confidence, and then select the same 3 domains as the domain corresponding to the domain sub-confidence from multiple domain resource libraries as the target domain resource libraries.
[0049] At this time, determining the supplemented speech recognition text in the target domain resource library according to the text pinyin includes: performing speech text adjustment on the text pinyin in each target domain resource library to obtain multiple supplemented speech recognition texts.
[0050] Specifically, a search for the same pronunciation (text pinyin) is performed in each domain resource library, and the searched recognition text is used as the supplemented speech recognition text to obtain multiple supplemented speech recognition texts, wherein the supplemented speech recognition text includes the text searched through the domain resource library and the speech recognition text recognized by the speech recognition model (that is, the original recognized text).
[0051] For example, the text of the speech to be recognized recognized by the speech recognition model is "Turn on the Bluetooth of the Tianyi smart speaker", but the confidence threshold for calculating the text is relatively low at this time, so the third-party resource library is called to supplement the results of the speech recognition text. The corresponding text obtained through supplementary retrieval is "Turn on the Bluetooth of the Tianyi smart speaker". At this time, the domain confidence is detected again, and it is found that the domain confidence obtained by the supplementary retrieval is relatively high. At this time, the text obtained by the supplementary retrieval is used as the final recognized text.
[0052] S150: Determine the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text.
[0053] In some embodiments, specifically, S150 includes: determining the domain confidence of each of the supplemented speech recognition texts according to a preset domain confidence calculation model; determining the confidence with the largest median of the domain confidences of multiple supplemented speech recognition texts as the target domain confidence; determining the supplemented speech recognition text corresponding to the target domain confidence as the target speech recognition text, wherein the speech recognition model includes a domain confidence calculation model and a content confidence calculation model, and the domain confidence calculation model is used to calculate the domain confidence of the input speech recognition text, specifically, calculating the domain confidence based on the similarity of each word in each domain in the speech recognition text.
[0054] That is, the domain confidence calculation model can calculate the domain confidence of the input text (each supplemented speech recognition text) according to the domain probability to which each word in the domain text belongs, obtain the domain confidence corresponding to each supplemented speech recognition text, and then determine the domain confidence with the highest value as the target domain confidence, and finally determine the supplemented speech recognition text corresponding to the target domain confidence as the final recognition text.
[0055] S160: Determine the speech recognition text as the target speech recognition text.
[0056] In this embodiment, if the text confidence threshold of the speech recognition text is greater than or equal to the confidence threshold, it means that the text recognized by the speech recognition model at this time is relatively accurate, and it is considered that the speech recognition model has accurate recognition results. At this time, the speech recognition text can be directly used as the target speech recognition text.
[0057] Figure 3 This is a flow chart of a speech recognition method provided by another embodiment of the present application. Figure 3 As shown, the speech recognition method of this embodiment includes steps S210-S290. Steps S210-S260 are similar to steps S110-S160 in the above embodiment and are not described in detail here. The following describes in detail the additional steps S270-S290 in this embodiment.
[0058] S270: Determine a semantic understanding model for the domain corresponding to the target speech recognition text.
[0059] The target speech recognition text in this embodiment carries a domain label. In order to improve the accuracy of semantic understanding, the computer device in this embodiment sets corresponding semantic understanding models for different service fields. After the target speech recognition text is determined, the semantic understanding model of the corresponding field of the target speech recognition text will be called.
[0060] S280. Perform semantic understanding on the target speech recognition text according to the semantic understanding model to obtain a semantic understanding result.
[0061] After the corresponding semantic understanding model is called, semantic understanding will be performed in the semantic understanding model. For example, if the target speech recognition text is "Turn on the Bluetooth of Tianyi smart speaker", the semantic understanding result obtained by the semantic understanding model is "Call Tianyi smart speaker and turn on the Bluetooth function of the Tianyi smart speaker".
[0062] S290. Execute processing corresponding to the semantic understanding result.
[0063] After obtaining the semantic understanding results, corresponding processing will be performed based on the semantic understanding results, calling the business system automatic control, business system automatic search or other intelligent functions of the business system, for example, turning on the Bluetooth function of the Tianyi smart speaker.
[0064] It can be seen that this application can introduce external information resources into the speech recognition system to complete the continued recognition of inaccurate content due to the untimely update of the language recognition model; based on the judgment of the field to which the speech belongs, different supplementary information resources are called to perform new continued recognition; based on the judgment of the field to which the speech belongs, the recognition results that are closest to the semantics of the user's recognition needs are recommended.
[0065] During the speech recognition process, the recognition of content that is not included in the language recognition model in the speech recognition engine is completed. During the recognition process, the industry scenario vocabulary dictionary (settable), social network search interface, and knowledge base content in this field can be called to search for content with the same pronunciation (pinyin search). Select the search content text in one or several fields as a supplement to the recognition results. Further based on the semantic understanding calculation and mining of the recognition results in the new content resources, in addition, the present application can further provide semantically related recognition results as a supplementary recognition result selection for users. The present application solves the problem that the language recognition model in the speech recognition engine cannot be updated in a timely manner, and forms an open architecture for the speech recognition engine to the Internet / mobile Internet, so that the recognition results are more oriented to search and application needs.
[0066] In some embodiments, before executing this application, it is necessary to build a speech recognition model and a semantic understanding model, and the steps are as follows:
[0067] 1. Obtain the scene and command words for controlling the terminal, such as the terminal name (standard name, abbreviation, nickname, name of a specific industry field, etc.);
[0068] 2. Build an industry scenario semantic understanding model within the semantic understanding model. Based on "industry name" and "terminal name," search industry resource libraries, including professional websites on the internet industry. Build a semantic model for industry terminal understanding.
[0069] 3. Based on the terminal's "control command words," repeat the above steps to build a "control command" understanding model in the semantic understanding model (essentially, obtaining a larger set of command words for a specific terminal or a category of terminals in a specific scenario within the industry);
[0070] 4. Build a speech recognition model based on the user's commonly used command language.
[0071] This application solves the problem of language model not being able to be updated in time in speech recognition engines, and forms an open architecture for speech recognition engines facing the Internet / mobile Internet.
[0072] This application can be applied in speech recognition systems, call center automatic question-answering systems, and other types of search engine systems.
[0073] For example, applying speech recognition methods and systems in call centers can achieve speech recognition and semantic understanding for the services provided by the call center. This can also identify hot topics in the call center. It can also identify new user descriptions of an item in call center recordings, or identify new events and content within the call center. By leveraging the timeliness and targeted nature of internet information resources, it can also identify requests that have not been trained with a language model.
[0074] In this application, the information of the vertical field is first imported into the system, and the corresponding features are extracted to train the speech recognition model, semantic understanding model, search engine model, etc. When the front end obtains the voice input signal, the audio signal is recognized by the speech recognition model. However, if the input signal is found to be similar to the vertical field information that has been trained, targeted recognition is performed. After the recognition is completed, the recognition result model is scored. When the score exceeds a certain threshold, the corresponding text is output with a field label. After that, the semantic understanding model calls the semantic understanding model rules of the corresponding field to perform corresponding semantic understanding and information search. This method can form the natural language recognition capability of a certain vertical field, and form the corresponding semantic understanding and information search capabilities of this field. The vertical field semantic features can be called by different speech recognition engines to form the core feature library of semantic understanding of this type of field knowledge. In this way, expansion and improvement can gradually form the speech recognition capability coverage of the selected business field, and can provide corresponding semantic understanding capabilities and information search capabilities.
[0075] Due to the limitation of the service field and the process of organizing and sorting out the service field knowledge in advance, the output speech comprehension ability is more targeted; in addition, this application can also provide secondary disambiguation capabilities through the natural language understanding ability of text, so that possible errors in the speech recognition output results can be corrected, and more accurate results can be achieved.
[0076] In summary, the present application obtains a speech to be recognized; inputs the speech to be recognized into a preset speech recognition model to obtain a speech recognition text, the content confidence level of the speech recognition text, and the domain confidence level; determines whether the text confidence level of the speech recognition text is less than a preset confidence threshold based on the content confidence level and the domain confidence level; if the text confidence level of the speech recognition text is less than the confidence threshold, invokes a third-party resource library to perform supplementary processing on the speech recognition text to obtain a supplemented speech recognition text; and determines the speech text with the highest domain confidence level in the supplemented speech recognition text as the target speech recognition text. In an embodiment of the present application, when the confidence of the text recognized by the speech recognition model is relatively low, it is necessary to call a third-party resource library to supplement the results of the speech recognition text, and then select the text with the highest domain confidence from the supplemented speech recognition text as the target speech recognition text. It can be seen that when the accuracy of the speech recognition text obtained according to the preset speech recognition model is not high, the speech recognition text can be supplemented by other databases, and then the optimal text can be selected from the supplemented text as the target speech recognition text. This solution can improve the accuracy of speech recognition.
[0077] Figure 4 This is a schematic block diagram of a speech recognition device provided by an embodiment of the present application. Figure 4 As shown, corresponding to the above speech recognition method, the present application also provides a speech recognition device. The speech recognition device includes a unit for executing the above speech recognition method, and the device can be configured in a desktop computer, tablet computer, laptop computer, etc. Specifically, please refer to Figure 4 The speech recognition device includes an acquisition unit 401, an input unit 402, a judgment unit 403, a supplementation unit 404 and a first determination unit 405.
[0078] An acquisition unit 401 is used to acquire a speech to be recognized;
[0079] An input unit 402 is configured to input the speech to be recognized into a preset speech recognition model to obtain a speech recognition text, a content confidence level, and a domain confidence level of the speech recognition text;
[0080] A judging unit 403 is configured to judge whether the text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence;
[0081] The supplementing unit 404 is configured to, when the text confidence threshold of the speech recognition text is less than the confidence threshold, call a third-party resource library to perform result supplement processing on the speech recognition text to obtain a supplemented speech recognition text;
[0082] The first determining unit 405 is configured to determine the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text.
[0083] In some embodiments, the third-party resource library includes multiple domain resource libraries, and the supplementing unit 404 is specifically configured to:
[0084] Obtain the text pinyin corresponding to the speech recognition text;
[0085] determining a target domain resource library from the plurality of domain resource libraries according to the domain confidence;
[0086] The supplemented speech recognition text is determined in the target domain resource library according to the pinyin of the text.
[0087] In some embodiments, the domain confidence includes multiple domain sub-confidences. When performing the step of determining the target domain resource repository from the multiple domain resource repositories according to the domain confidence, the supplementing unit 404 is further configured to:
[0088] Extracting a preset number of confidences with the largest values from the plurality of domain sub-confidences to obtain a plurality of target domain sub-confidences;
[0089] Determining domain resource libraries corresponding to the sub-confidences of the target domains from the plurality of domain resource libraries, to obtain a plurality of target domain resource libraries;
[0090] The step of determining the supplemented speech recognition text in the target domain resource library according to the pinyin of the text includes:
[0091] The pinyin of the text is respectively adjusted in each target domain resource library to obtain a plurality of supplemented speech recognition texts.
[0092] In some embodiments, the first determining unit 405 is specifically configured to:
[0093] Determining the domain confidence of each of the supplemented speech recognition texts according to a preset domain confidence calculation model;
[0094] Determining the confidence with the largest median value among the domain confidences of the plurality of supplemented speech recognition texts as the target domain confidence;
[0095] The supplemented speech recognition text corresponding to the target domain confidence is determined as the target speech recognition text.
[0096] In some embodiments, the determining unit 403 is specifically configured to:
[0097] Determining the text confidence according to the content confidence, the preset content confidence weight, the domain confidence, and the preset domain confidence weight;
[0098] Determine whether the text confidence is less than the confidence threshold.
[0099] Figure 5 This is a schematic block diagram of a speech recognition device provided by another embodiment of the present application. Figure 5 As shown, the speech recognition device of this embodiment is based on the above embodiment and further includes a second determining unit 406 , a comprehension unit 407 , a processing unit 408 and a third determining unit 409 .
[0100] A second determining unit 406 is configured to determine a semantic understanding model for the target speech recognition text corresponding to a domain;
[0101] An understanding unit 407 is configured to perform semantic understanding on the target speech recognition text according to the semantic understanding model to obtain a semantic understanding result;
[0102] The processing unit 408 is used to perform processing corresponding to the semantic understanding result.
[0103] The third determining unit 409 is configured to determine the speech recognition text as the target speech recognition text when the text confidence threshold of the speech recognition text is greater than or equal to the confidence threshold.
[0104] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned speech recognition device and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.
[0105] The above-mentioned speech recognition device can be implemented in the form of a computer program. The computer program can be used in Figure 6 Runs on the computer device shown.
[0106] See also Figure 6 , Figure 6This is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 600 can be a terminal or a server. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.
[0107] See Figure 6 The computer device 600 includes a processor 602 , a memory, and a network interface 605 connected via a system bus 601 , wherein the memory may include a non-volatile storage medium 603 and an internal memory 604 .
[0108] The non-volatile storage medium 603 may store an operating system 6031 and a computer program 6032. The computer program 6032 includes program instructions, which, when executed, may enable the processor 602 to perform a speech recognition method.
[0109] The processor 602 is used to provide computing and control capabilities to support the operation of the entire computer device 600.
[0110] The internal memory 604 provides an environment for the operation of the computer program 6032 in the non-volatile storage medium 603. When the computer program 6032 is executed by the processor 602, the processor 602 can perform a speech recognition method.
[0111] The network interface 605 is used to communicate with other devices over the network. Figure 6 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 600 to which the solution of the present application is applied. The specific computer device 600 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0112] The processor 602 is configured to execute a computer program 6032 stored in the memory to implement the following steps:
[0113] Get the speech to be recognized;
[0114] Inputting the speech to be recognized into a preset speech recognition model to obtain speech recognition text, content confidence and domain confidence of the speech recognition text;
[0115] Determining whether a text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence;
[0116] If the text confidence threshold of the speech recognition text is less than the confidence threshold, calling a third-party resource library to perform result supplementation processing on the speech recognition text to obtain a supplemented speech recognition text;
[0117] The speech text with the highest domain confidence in the supplemented speech recognition text is determined as the target speech recognition text.
[0118] In some embodiments, the third-party resource library includes multiple domain resource libraries. When the processor 602 implements the step of calling the third-party resource library to supplement the speech recognition text and obtain the supplemented speech recognition text, the processor 602 specifically implements the following steps:
[0119] Obtain the text pinyin corresponding to the speech recognition text;
[0120] determining a target domain resource library from the plurality of domain resource libraries according to the domain confidence;
[0121] The supplemented speech recognition text is determined in the target domain resource library according to the pinyin of the text.
[0122] In some embodiments, the domain confidence includes multiple domain sub-confidences. When the processor 602 implements the step of determining the target domain resource library from the multiple domain resource libraries according to the domain confidence, the processor 602 specifically implements the following steps:
[0123] Extracting a preset number of confidences with the largest values from the plurality of domain sub-confidences to obtain a plurality of target domain sub-confidences;
[0124] Determining domain resource libraries corresponding to the sub-confidences of the target domains from the plurality of domain resource libraries, to obtain a plurality of target domain resource libraries;
[0125] The step of determining the supplemented speech recognition text in the target domain resource library according to the pinyin of the text includes:
[0126] The pinyin of the text is respectively adjusted in each target domain resource library to obtain a plurality of supplemented speech recognition texts.
[0127] In some embodiments, when the processor 602 implements the step of determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text, the processor 602 specifically implements the following steps:
[0128] Determining the domain confidence of each of the supplemented speech recognition texts according to a preset domain confidence calculation model;
[0129] Determining the confidence with the largest median value among the domain confidences of the plurality of supplemented speech recognition texts as the target domain confidence;
[0130] The supplemented speech recognition text corresponding to the target domain confidence is determined as the target speech recognition text.
[0131] In some embodiments, when the processor 602 implements the step of determining whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence, the processor 602 specifically implements the following steps:
[0132] Determining the text confidence according to the content confidence, the preset content confidence weight, the domain confidence, and the preset domain confidence weight;
[0133] Determine whether the text confidence is less than the confidence threshold.
[0134] In some embodiments, after implementing the step of determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text, the processor 602 further implements the following steps:
[0135] Determining a semantic understanding model for the target speech recognition text;
[0136] Performing semantic understanding on the target speech recognition text according to the semantic understanding model to obtain a semantic understanding result;
[0137] Execute processing corresponding to the semantic understanding result.
[0138] In some embodiments, after implementing the step of determining whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence, the processor 602 further implements the following steps:
[0139] If the text confidence threshold of the speech recognition text is greater than or equal to the confidence threshold, the speech recognition text is determined as the target speech recognition text.
[0140] It should be understood that in the embodiment of the present application, the processor 602 may be a central processing unit (CPU), and the processor 602 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0141] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0142] Therefore, the present application also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps:
[0143] Get the speech to be recognized;
[0144] Inputting the speech to be recognized into a preset speech recognition model to obtain speech recognition text, content confidence and domain confidence of the speech recognition text;
[0145] Determining whether a text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence;
[0146] If the text confidence threshold of the speech recognition text is less than the confidence threshold, calling a third-party resource library to perform result supplementation processing on the speech recognition text to obtain a supplemented speech recognition text;
[0147] The speech text with the highest domain confidence in the supplemented speech recognition text is determined as the target speech recognition text.
[0148] In some embodiments, the third-party resource library includes multiple domain resource libraries. When the processor executes the program instructions to implement the step of calling the third-party resource library to supplement the speech recognition text and obtain the supplemented speech recognition text, the processor specifically implements the following steps:
[0149] Obtain the text pinyin corresponding to the speech recognition text;
[0150] determining a target domain resource library from the plurality of domain resource libraries according to the domain confidence;
[0151] The supplemented speech recognition text is determined in the target domain resource library according to the pinyin of the text.
[0152] In some embodiments, the domain confidence includes multiple domain sub-confidences. When the processor executes the program instructions to implement the step of determining the target domain resource library from the multiple domain resource libraries based on the domain confidence, the processor specifically implements the following steps:
[0153] Extracting a preset number of confidences with the largest values from the plurality of domain sub-confidences to obtain a plurality of target domain sub-confidences;
[0154] Determining domain resource libraries corresponding to the sub-confidences of the target domains from the plurality of domain resource libraries, to obtain a plurality of target domain resource libraries;
[0155] The step of determining the supplemented speech recognition text in the target domain resource library according to the pinyin of the text includes:
[0156] The pinyin of the text is respectively adjusted in each target domain resource library to obtain a plurality of supplemented speech recognition texts.
[0157] In some embodiments, when the processor executes the program instructions to implement the step of determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text, the processor specifically implements the following steps:
[0158] Determining the domain confidence of each of the supplemented speech recognition texts according to a preset domain confidence calculation model;
[0159] Determining the confidence with the largest median value among the domain confidences of the plurality of supplemented speech recognition texts as the target domain confidence;
[0160] The supplemented speech recognition text corresponding to the target domain confidence is determined as the target speech recognition text.
[0161] In some embodiments, when the processor executes the program instructions to implement the step of determining whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence, the processor specifically implements the following steps:
[0162] Determining the text confidence according to the content confidence, the preset content confidence weight, the domain confidence, and the preset domain confidence weight;
[0163] Determine whether the text confidence is less than the confidence threshold.
[0164] In some embodiments, after executing the program instructions to implement the step of determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text, the processor further implements the following steps:
[0165] Determining a semantic understanding model for the target speech recognition text;
[0166] Performing semantic understanding on the target speech recognition text according to the semantic understanding model to obtain a semantic understanding result;
[0167] Execute processing corresponding to the semantic understanding result.
[0168] In some embodiments, after executing the program instructions to implement the step of determining whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence, the processor further implements the following steps:
[0169] If the text confidence threshold of the speech recognition text is greater than or equal to the confidence threshold, the speech recognition text is determined as the target speech recognition text.
[0170] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0171] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0172] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and other division methods may be used in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not implemented.
[0173] The steps in the method of the embodiment of the present application can be adjusted in order, combined, and deleted according to actual needs. The units in the device of the embodiment of the present application can be combined, divided, and deleted according to actual needs. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.
[0175] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A speech recognition method, characterized in that: include: Get the speech to be recognized; Inputting the speech to be recognized into a preset speech recognition model to obtain speech recognition text, content confidence and domain confidence of the speech recognition text; Determining whether a text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence; If the text confidence threshold of the speech recognition text is less than the confidence threshold, calling a third-party resource library to perform result supplementation processing on the speech recognition text to obtain a supplemented speech recognition text; Determine the speech text with the highest domain confidence among the supplemented speech recognition texts as the target speech recognition text; The third-party resource library includes an Internet search engine, a social network search interface and / or multiple domain resource libraries; When the third-party resource library includes multiple domain resource libraries, the calling of the third-party resource library to perform result supplementation processing on the speech recognition text to obtain the supplemented speech recognition text includes: Obtain the text pinyin corresponding to the speech recognition text; determining a target domain resource library from the plurality of domain resource libraries according to the domain confidence; The supplemented speech recognition text is determined in the target domain resource library according to the pinyin of the text.
2. The method according to claim 1, characterized in that The domain confidence includes a plurality of domain sub-confidences, and determining a target domain resource library from the plurality of domain resource libraries according to the domain confidence includes: Extracting a preset number of confidences with the largest values from the plurality of domain sub-confidences to obtain a plurality of target domain sub-confidences; Determining domain resource libraries corresponding to the sub-confidences of the target domains from the plurality of domain resource libraries, to obtain a plurality of target domain resource libraries; The step of determining the supplemented speech recognition text in the target domain resource library according to the pinyin of the text includes: The pinyin of the text is respectively adjusted in each target domain resource library to obtain a plurality of supplemented speech recognition texts.
3. The method according to claim 2, characterized in that The step of determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text includes: Determining the domain confidence of each of the supplemented speech recognition texts according to a preset domain confidence calculation model; Determining the confidence with the largest median value among the domain confidences of the plurality of supplemented speech recognition texts as the target domain confidence; The supplemented speech recognition text corresponding to the target domain confidence is determined as the target speech recognition text.
4. The method according to claim 1, wherein The determining, based on the content confidence and the domain confidence, whether the text confidence of the speech recognition text is less than a preset confidence threshold, includes: Determining the text confidence according to the content confidence, the preset content confidence weight, the domain confidence, and the preset domain confidence weight; Determine whether the text confidence is less than the confidence threshold.
5. The method according to claim 1, wherein After determining the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text, the method further includes: Determining a semantic understanding model for the target speech recognition text; Performing semantic understanding on the target speech recognition text according to the semantic understanding model to obtain a semantic understanding result; Execute processing corresponding to the semantic understanding result.
6. The method according to any one of claims 1 to 5, characterized in that After determining whether the text confidence of the speech recognition text is less than a preset confidence threshold based on the content confidence and the domain confidence, the method further includes: If the text confidence threshold of the speech recognition text is greater than or equal to the confidence threshold, the speech recognition text is determined as the target speech recognition text.
7. A speech recognition device, characterized in that: include: An acquisition unit, configured to acquire the speech to be recognized; An input unit, configured to input the speech to be recognized into a preset speech recognition model to obtain a speech recognition text, a content confidence level, and a domain confidence level of the speech recognition text; a judging unit, configured to judge whether a text confidence of the speech recognition text is less than a preset confidence threshold according to the content confidence and the domain confidence; A supplementing unit, configured to, when a text confidence threshold of the speech recognition text is less than the confidence threshold, call a third-party resource library to perform result supplement processing on the speech recognition text to obtain a supplemented speech recognition text; A first determining unit is configured to determine the speech text with the highest domain confidence in the supplemented speech recognition text as the target speech recognition text; The third-party resource library includes an Internet search engine, a social network search interface and / or multiple domain resource libraries; when the third-party resource library includes multiple domain resource libraries, the supplementing unit is specifically used to: Obtain the text pinyin corresponding to the speech recognition text; determining a target domain resource library from the plurality of domain resource libraries according to the domain confidence; The supplemented speech recognition text is determined in the target domain resource library according to the pinyin of the text.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 6 can be implemented.
Citation Information
Patent Citations
Predictive speech recognition method and device based on big data
CN109273004A
Method and device for recognizing natural language, vehicle-mounted multi-media host and computer readable storage medium
CN109785840A