Speech recognition method, device, equipment and storage medium
By building a speech recognition decoding network based on vertical keyword and sentence decoding networks, combined with cloud servers and multi-model decoding, the problem of inaccurate vertical keyword recognition is solved, and efficient speech recognition is achieved in specific scenarios, especially the accurate recognition of names of people and places.
Patent Information
- Application Number
- CN202111274880.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-10-29
AI Technical Summary
Existing speech recognition technology has poor recognition effect in scenarios involving vertical keywords, especially names of people and places, and cannot meet user needs, affecting user experience.
Build a speech recognition and decoding network based on the vertical keyword set and sentence decoding network, combine the cloud server to build and transmit the vertical keyword set in real time, use the speech recognition and decoding network and the general speech recognition model to perform multi-model decoding, and ensure accurate recognition of vertical keywords through acoustic score incentives.
The accuracy of speech recognition has been improved, especially in specific scenarios involving vertical keywords. It can accurately identify keywords such as names of people and places, thereby improving the user experience.
Smart Images

Figure CN113920999B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a speech recognition method, apparatus, device and storage medium. Background Art
[0002] With the rapid development of technologies such as mobile Internet and artificial intelligence, human-computer interaction scenarios have appeared in large numbers in people's daily life and production processes. As an important interface for human-computer interaction, speech recognition is becoming more and more widely used.
[0003] Currently, the most effective approach for speech recognition is to use neural network technology to learn from massive amounts of data, creating a speech recognition model. This model performs very well in common scenarios. In theory, with sufficient data and sufficient vocabulary coverage, excellent recognition results can be achieved.
[0004] However, in voice recognition scenarios involving vertical keywords, such as calling mobile phone contacts, sending messages to mobile phone contacts, checking city weather conditions, navigation and positioning, etc., the existing voice recognition effect is very poor and is usually unable to accurately recognize user voices, especially for vertical keywords such as names of people and places in user voices, which often cannot be successfully recognized. Summary of the Invention
[0005] Based on the above technical status, the embodiments of the present application propose a speech recognition method, device, equipment and storage medium, which can accurately recognize the speech to be recognized, especially can accurately recognize the speech in specific scenarios involving vertical keywords, especially can accurately recognize vertical keywords in the speech.
[0006] A speech recognition method, comprising:
[0007] Obtaining the acoustic state sequence of the speech to be recognized;
[0008] Based on the vertical category keyword set and sentence decoding network of the scene to be recognized, a speech recognition decoding network is constructed, wherein the sentence decoding network is constructed by at least performing sentence induction processing on the text corpus of the scene to be recognized;
[0009] The acoustic state sequence is decoded using the speech recognition decoding network to obtain a speech recognition result.
[0010] A speech recognition method, comprising:
[0011] Obtaining the acoustic state sequence of the speech to be recognized;
[0012] The acoustic state sequence is decoded using a speech recognition decoding network to obtain a first speech recognition result, and the acoustic state sequence is decoded using a general speech recognition model to obtain a second speech recognition result; the speech recognition decoding network is constructed based on a vertical category keyword set and a sentence decoding network in the scene to which the speech to be recognized belongs;
[0013] Performing acoustic score incentive on the first speech recognition result;
[0014] A final speech recognition result is determined at least from the first speech recognition result and the second speech recognition result after stimulation.
[0015] A speech recognition device, comprising:
[0016] An acoustic recognition unit, configured to obtain an acoustic state sequence of a speech to be recognized;
[0017] A network construction unit, configured to construct a speech recognition decoding network based on a set of vertically categorized keywords and a sentence decoding network for the scene to which the speech to be recognized belongs, wherein the sentence decoding network is constructed at least by performing sentence induction processing on a text corpus for the scene to which the speech to be recognized belongs;
[0018] A decoding processing unit is used to decode the acoustic state sequence using the speech recognition decoding network to obtain a speech recognition result.
[0019] A speech recognition device, comprising:
[0020] An acoustic recognition unit, configured to obtain an acoustic state sequence of a speech to be recognized;
[0021] A multidimensional decoding unit, configured to decode the acoustic state sequence using a speech recognition decoding network to obtain a first speech recognition result, and to decode the acoustic state sequence using a universal speech recognition model to obtain a second speech recognition result; the speech recognition decoding network is constructed based on a set of vertical keywords and a sentence decoding network for the scene to which the speech to be recognized belongs;
[0022] an acoustic excitation unit, configured to perform acoustic score excitation on the first speech recognition result;
[0023] The decision processing unit is used to determine a final speech recognition result from at least the first speech recognition result after stimulation and the second speech recognition result.
[0024] A speech recognition device, comprising:
[0025] memory and processor;
[0026] The memory is connected to the processor and is used to store programs;
[0027] The processor is used to implement the above-mentioned speech recognition method by running the program stored in the memory.
[0028] A storage medium stores a computer program, which implements the above-mentioned speech recognition method when executed by a processor.
[0029] The speech recognition method proposed in the present application can construct a speech recognition decoding network based on the set of vertical category keywords in the scene to which the speech to be recognized belongs and a pre-constructed sentence decoding network for the scene. Then, the speech recognition decoding network contains various speech sentence information in the scene to which the speech to be recognized belongs, and also contains various vertical category keywords in the scene to which the speech to be recognized belongs. The speech recognition decoding network can be used to decode speech composed of any sentence and any vertical category keyword in the scene to which the speech to be recognized belongs. Therefore, by constructing the above-mentioned speech recognition decoding network, the speech to be recognized can be accurately recognized, especially the speech in a specific scene involving vertical category keywords can be accurately recognized, and in particular, the vertical category keywords in the speech can be accurately recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0031] Figure 1 This is a flow chart of a speech recognition method provided in an embodiment of the present application;
[0032] Figure 2 Schematic diagram of a word-level sentence decoding network provided in an embodiment of the present application;
[0033] Figure 3 This is a flow chart of another speech recognition method provided in an embodiment of the present application;
[0034] Figure 4 This is a flow chart of another speech recognition method provided in an embodiment of the present application;
[0035] Figure 5 This is a flow chart of another speech recognition method provided in an embodiment of the present application;
[0036] Figure 6 This is a flow chart of another speech recognition method provided in an embodiment of the present application;
[0037] Figure 7 This is a schematic diagram of a text sentence network provided by an embodiment of the present application;
[0038] Figure 8 is a schematic diagram of a pronunciation-level sentence decoding network provided in an embodiment of the present application;
[0039] Figure 9 Schematic diagram of a word-level name network provided by an embodiment of the present application;
[0040] Figure 10 The embodiment of this application provides Figure 9 Schematic diagram of the corresponding pronunciation-level name network;
[0041] Figure 11 is a flowchart of a process for correcting a first speech recognition result using a second speech recognition result, provided in an embodiment of the present application;
[0042] Figure 12 is a flowchart of a process for determining a final speech recognition result from a first speech recognition result and a second speech recognition result, provided by an embodiment of the present application;
[0043] Figure 13 is a schematic diagram of a state network of speech recognition results provided by an embodiment of the present application;
[0044] Figure 14 Yes Figure 13 Schematic diagram of the state network after path expansion of the speech recognition result shown;
[0045] Figure 15 This is a structural diagram of a speech recognition device provided in an embodiment of the present application;
[0046] Figure 16 is a structural diagram of another speech recognition device provided in an embodiment of the present application;
[0047] Figure 17 This is a structural diagram of a speech recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The technical solutions of the embodiments of the present application are applicable to speech recognition application scenarios. By adopting the technical solutions of the embodiments of the present application, speech content can be recognized more accurately, especially in specific business scenarios involving vertical keywords. Speech content can be recognized more accurately, especially vertical keywords in speech can be accurately recognized, thereby improving the overall speech recognition effect.
[0049] The above-mentioned vertical keywords generally refer to different keywords belonging to the same type. For example, names of people, places, application names, etc. constitute different vertical keywords. For example, the different names in the user's address book constitute the name vertical keywords, the different place names in the user's area constitute the place name vertical keywords, and the names of the various applications installed on the user's terminal constitute the application name vertical keywords.
[0050] The above-mentioned business scenarios involving vertical keywords refer to business scenarios that contain vertical keywords in the corresponding interactive voice, such as voice dialing, voice navigation and other business scenarios. Since the user must say the name of the person calling or the name of the place to navigate, for example, the user may say "Call XX" or "Navigate to YY", where "XX" may be any person's name in the user's mobile phone address book, and "YY" may be a place name in the user's area. It can be seen that the voice in these business scenarios contains vertical keywords (such as names of people and places), so these business scenarios are business scenarios involving vertical keywords.
[0051] With the widespread use of artificial intelligence and smart terminals, human-computer interaction scenarios are becoming more and more common, and voice recognition is an important interface for human-computer interaction. For example, on smart terminals, many manufacturers have built-in voice assistants in the terminal operating system, allowing users to control the terminal through voice. For example, users can use voice to call or send text messages to contacts in the communication, or use voice to query the city weather, or use voice to open or close terminal applications, etc. These interaction scenarios are specific business scenarios compared to ordinary voice recognition business scenarios. Most of the voices in these scenarios involve vertical keywords (such as names of people and places in the address book, and names of terminal applications).
[0052] Compared with ordinary text keywords, vertical keywords are characterized by frequent changes, unpredictability, and user customization. In addition, vertical keywords account for an extremely low proportion in the massive speech recognition training corpus. As a result, conventional speech recognition solutions that train speech recognition models using massive corpora are often unable to cope with speech recognition tasks involving vertical keywords.
[0053] For example, compared to conventional textual expectations, names appear very rarely. Therefore, even in massive training corpora, names are rare, which makes it impossible for the model to fully learn the characteristics of names through massive corpora. Moreover, names are user-defined text content, which is inexhaustible and unpredictable. It is unrealistic to generate all names manually. Furthermore, the names of contacts stored in the user's address book may not be standard names. They may be nicknames, code names, or nicknames. Users may even modify, add, or delete contacts in the address book at any time. This makes the names of different users in the address book highly diverse, making it impossible for the speech recognition model to learn all the characteristics of names in a unified way.
[0054] Therefore, the conventional technical solution of training a speech recognition model through massive corpus and using the speech recognition model to implement speech recognition function is not fully competent for speech recognition tasks in business scenarios involving vertical keywords, especially for vertical keywords in speech, which often cannot be successfully recognized, seriously affecting the user experience.
[0055] In view of the above-mentioned technical status, an embodiment of the present application proposes a speech recognition method, which can improve the speech recognition effect, especially the speech recognition effect in business scenarios involving vertical keywords.
[0056] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0057] This application embodiment proposes a speech recognition method, see Figure 1 As shown, the method includes:
[0058] S101: Acquire an acoustic state sequence of a speech to be recognized.
[0059] Specifically, based on the application scenario introduction of the technical solution of the embodiment of the present application introduced above, the above-mentioned speech to be recognized is specifically speech data in a business scenario involving vertical keywords, and the speech to be recognized contains speech content of vertical keywords.
[0060] The above-mentioned speech to be recognized is subjected to endpoint detection, windowing and framing, feature extraction and other processing to obtain its audio features, which can be Mel-frequency cepstral coefficient MFCC features or any other type of audio features.
[0061] After acquiring the audio features of the speech to be recognized, these features are input into the acoustic model for acoustic recognition, resulting in a posterior score for the acoustic state of each frame of audio, i.e., an acoustic state sequence. The acoustic model is primarily a neural network structure that uses forward computation to identify the acoustic state and posterior score corresponding to each audio frame. The acoustic state corresponding to the audio frame is specifically the pronunciation unit corresponding to the audio frame, such as the phoneme or phoneme sequence corresponding to the audio frame.
[0062] The conventional speech recognition technology solution is an acoustic model + language model architecture. That is, the acoustic model is first used to acoustically recognize the speech to be recognized, realizing the mapping of speech features to phoneme sequences; then, the language model is used to recognize the phoneme sequence, realizing the mapping of phonemes to text.
[0063] In conventional speech recognition solutions, the acoustic state sequence of the speech to be recognized, obtained through acoustic recognition, is input into a language model for decoding, thereby determining the text content corresponding to the speech to be recognized. This language model is trained based on a massive amount of training data and is capable of mapping phonemes to text.
[0064] Unlike conventional speech recognition solutions, the embodiment of the present application does not use the above-mentioned speech model trained based on massive corpus to decode the acoustic state sequence, but instead uses a decoding network constructed in real time for decoding. For details, please see the following content.
[0065] S102: Construct a speech recognition decoding network based on the vertical keyword set and sentence decoding network in the scene to which the speech to be recognized belongs.
[0066] Specifically, unlike conventional language models, the embodiment of the present application constructs a speech recognition and decoding network in real time when recognizing speech in vertical keyword business scenarios, which is used to decode the acoustic state sequence of the speech to be recognized and obtain speech recognition results.
[0067] The above-mentioned speech recognition decoding network is constructed by a set of vertical keywords in the scene to which the speech to be recognized belongs, and a pre-constructed sentence decoding network in the scene to which the speech to be recognized belongs.
[0068] The sentence decoding network for the scene to be recognized speech is constructed by at least performing sentence induction processing on the text corpus for the scene to be recognized speech;
[0069] The scenario to which the speech to be recognized belongs specifically refers to the service scenario to which the speech to be recognized belongs. For example, if the speech to be recognized is "I want to make XX a call," the speech to be recognized belongs to the service of making a phone call, and therefore the scenario to which the speech to be recognized belongs is the phone call scenario. For another example, if the speech to be recognized is "Navigate to XX," the speech to be recognized belongs to the service of navigation, and therefore the scenario to which the speech to be recognized belongs is the navigation scenario.
[0070] After research, the inventors of this application found that in business scenarios involving vertical keywords, a considerable portion of user voice sentence patterns are fixed sentence patterns. For example, in the scenarios of making calls or sending text messages, users' commonly used sentence patterns are usually "I want to give XX a call" or "send a message to XX for me"; in the voice navigation scenario, users' commonly used sentence patterns are usually "go to XX (place name)" or "navigate to XX (place name)".
[0071] Therefore, in a certain business scenario involving vertical keywords, the sentence patterns of user voices are regular, or exhaustive. By summarizing these sentence patterns, a sentence pattern network corresponding to the scenario can be obtained, which is named a sentence pattern decoding network in the embodiment of the present application. It can be understood that the sentence pattern decoding network constructed based on the above method can contain sentence pattern information corresponding to the scenario. When the text corpus of all sentence patterns in a certain scenario is summarized to obtain a sentence pattern decoding network, the sentence pattern decoding network can contain any sentence pattern in the scenario.
[0072] As a preferred implementation, the embodiment of the present application is constructed by performing sentence induction and grammatical slot definition processing on the text corpus in the scenario to which the speech to be recognized belongs.
[0073] The above-mentioned definition of the grammatical slot in the sentence pattern is specifically to determine the grammatical type of the text slot in the sentence pattern. In an embodiment of the present application, the text slots in the text sentence are divided into ordinary grammatical slots and replacement grammatical slots, wherein the text slot where the non-vertical keyword in the text sentence is located is defined as an ordinary grammatical slot, and the text slot where the vertical keyword in the text sentence is located is defined as a replacement grammatical slot.
[0074] As a simple example, in the scenario of making a phone call or sending a text message, by summarizing the sentence structure and defining the grammatical slots for the text corpus "I want to give XX a call" or "send a message to XX for me", we can get the following: Figure 2The sentence decoding network shown. The sentence decoding network consists of nodes and directed arcs connecting the nodes, wherein the directed arcs correspond to ordinary grammar slots and replacement grammar slots, and the directed arcs carry label information for recording the text content in the slot. Specifically, the ordinary grammar slot entries are segmented and connected in series through nodes and directed arcs. The directed arcs between the two nodes are marked with word information, and the left and right sides of the colon represent input and output information respectively. Here, the input and output information are set to be the same, and multiple words after the segmentation of a single entry are connected in series. Different entries in the same grammar slot are connected in parallel. The replacement grammar slot uses the placeholder "#placeholder#" for placeholding, and no expansion is performed. The nodes are numbered in sequence, wherein words with the same starting node identifier share a starting node, and words with the same ending node identifier share an ending node. Figure 2 This example illustrates a simple network diagram for decoding word-level sentences in a phone book. The regular grammar slot before the replacement grammar slot contains three entries: "I want to give," "send a message to," and "give a call." The regular grammar slot after the replacement grammar slot contains three entries: "for me," "a call," and "a call with her number." The connection between nodes 10 and 18 indicates that you can go directly from node 10 to the end node. The "" on the arc represents silence.
[0075] The specific construction process of the above sentence decoding network can be found in the detailed description of the embodiments below.
[0076] The above-mentioned vertical keyword set for the business scenario to which the voice to be recognized belongs refers to a set consisting of all vertical keyword sets for the business scenario to which the voice to be recognized belongs. For example, assuming that the voice to be recognized is a voice in a voice call or voice text message scenario, the vertical keyword set for the business scenario to which the voice to be recognized belongs can specifically be a set of names consisting of the names in the user's address book; assuming that the voice to be recognized is a voice in a voice navigation scenario, the vertical keyword set for the business scenario to which the voice to be recognized belongs can specifically be a set of place names consisting of the various place names in the user's area.
[0077] The speech recognition decoding network is obtained by adding the vertical category keywords from the vertical category keyword set in the business scenario of the speech to be recognized to the replacement grammar slot of the sentence decoding network. It can be seen that the decoding network not only contains all the speech sentence patterns in the business scenario of the speech to be recognized, but also contains all the vertical category keywords in the scenario. Therefore, the speech recognition decoding network can recognize the speech sentence patterns in the business scenario of the speech to be recognized, and can recognize the vertical category keywords in the speech, that is, it can recognize the speech in the business scenario.
[0078] The specific construction process of the above-mentioned speech recognition decoding network will be described in detail in subsequent embodiments.
[0079] It should be noted that, as a preferred implementation, when constructing a speech recognition and decoding network, the embodiment of the present application is specifically constructed on the server, that is, the vertical keyword set under the business scenario to which the speech to be recognized belongs is transmitted to the cloud server, so that the cloud server constructs a speech recognition and decoding network based on the vertical keyword set under the business scenario to which the speech to be recognized belongs and the pre-constructed sentence decoding network.
[0080] For example, when a user says the voice command "I want to give XXa call" to the mobile phone terminal, the mobile phone terminal transmits the local address book (i.e., a set of keywords for names) to the cloud server. The cloud server builds a speech recognition decoding network based on the names in the address book and the sentence decoding network in the call scenario. Then, the speech recognition decoding network contains various sentence patterns for making calls, as well as the names of the people in the address book who are making the call. Using this decoding network, it is possible to recognize the voice of the user calling any member of the current address book.
[0081] In conventional technical solutions, the speech recognition decoding network is built locally on the user terminal, not in real time. Instead, it is pre-built and repeatedly called. Due to the relatively low computing resources of terminal devices, the network construction speed is slow, and the network decoding speed is limited. Moreover, the decoding network built non-real time cannot be updated in time when the vertical keyword set is updated, affecting the speech recognition effect.
[0082] The embodiment of the present application, on the other hand, constructs a speech recognition and decoding network on a cloud server. Furthermore, during the speech recognition process, the vertical category keyword set is input in real time by executing step S102, and the speech recognition and decoding network is constructed. This ensures that the constructed speech recognition and decoding network contains the latest vertical category keyword set, that is, the vertical category keyword set required for this recognition, thereby accurately identifying the vertical category keywords. At the same time, based on the powerful computing power of the cloud server, the speech recognition and decoding network will have stronger decoding performance.
[0083] Moreover, thanks to the centralized processing power of the cloud server, there is no need to build a separate speech recognition and decoding network for each terminal; only the cloud server needs to build one. For any terminal connected to the cloud server, as long as the terminal transmits the speech information to be recognized and the vertical keyword set of the business scenario to which the speech to be recognized belongs (such as the terminal's locally stored address book) to the cloud server, the cloud server can build an appropriate speech recognition and decoding network for the speech to be recognized and decode the speech to be recognized.
[0084] S103: Decode the acoustic state sequence using the speech recognition decoding network to obtain a speech recognition result.
[0085] As can be seen from the above description, the speech recognition and decoding network constructed above includes sentence patterns in the business scenario of the speech to be recognized and a set of vertical keywords in the business scenario of the speech to be recognized. Therefore, the speech content of the speech to be recognized can be recognized using this speech recognition and decoding network.
[0086] For example, if the user terminal collects a voice signal of the user calling someone in the terminal address book, in order to recognize the voice, according to the introduction of step S102 above, the local address book of the terminal (the address book is used as a vertical keyword set) and the pre-built sentence decoding network corresponding to the calling business scenario are used to build a speech recognition decoding network containing the names of the people in the local address book. For details, please refer to Figure 2 The speech recognition and decoding network architecture shown in Figure 1 contains not only all the phone call speech patterns, but also all the names in the local address book. In theory, when a user uses a voice-controlled terminal to call anyone in the terminal address book, the user's spoken sentences are all included in the speech recognition and decoding network.
[0087] Exemplarily, in the above-mentioned speech recognition and decoding network, there are multiple identical or different sentence paths composed of different vertical category keywords. When a certain acoustic state sequence matches the pronunciation of one or several sentence paths in the speech recognition and decoding network, it can be determined that the text content of the acoustic state sequence is the text content of the sentence path. Therefore, the speech recognition result finally decoded may be the text of one or several paths in the speech recognition and decoding network, that is, the final speech recognition result may be one or more.
[0088] For example, assuming the user voice captured by the terminal is "I want to give John a call," after acoustic recognition of the voice to obtain its acoustic state sequence, a voice recognition and decoding network is constructed using the terminal's address book. This voice recognition and decoding network includes the sentence pattern "I want to give XX a call" and the name "John." The network also includes other sentence patterns and other names. Based on this voice recognition and decoding network, the acoustic state sequence of the voice is pronunciation-matched with each path in the network. It can be determined that the acoustic state sequence matches the pronunciation of the path "I want to give John a call," and the voice recognition result "I want to give John a call" is obtained, thus achieving recognition of the user voice.
[0089] From the above introduction, it can be seen that the speech recognition method proposed in the embodiment of the present application can construct a speech recognition decoding network based on the set of vertical category keywords in the business scenario to which the speech to be recognized belongs and the pre-constructed sentence decoding network in the business scenario. Then, in the speech recognition decoding network, various speech sentences in the business scenario to which the speech to be recognized belongs are included, and various vertical category keywords in the business scenario to which the speech to be recognized belongs are also included. The speech recognition decoding network can be used to decode speech composed of any sentence and any vertical category keyword in the business scenario to which the speech to be recognized belongs. Therefore, by constructing the above-mentioned speech recognition decoding network, the speech to be recognized can be accurately recognized, especially the speech in a specific scenario involving vertical category keywords can be accurately recognized, and in particular, the vertical category keywords in the speech can be accurately recognized.
[0090] As a preferred embodiment, while using the above-mentioned speech recognition decoding network to decode the acoustic state sequence of the speech to be recognized, a general speech recognition model is also used to decode the acoustic state sequence of the speech to be recognized.
[0091] For the sake of convenience in distinction, the result obtained by decoding the acoustic state sequence of the speech to be recognized using the above-mentioned speech recognition decoding network is named the first speech recognition result, and the result obtained by decoding the acoustic state sequence of the speech to be recognized using the above-mentioned general speech recognition model is named the second speech recognition result.
[0092] See also Figure 3As shown, after executing step S301 and obtaining the acoustic state sequence of the speech to be recognized, step S302 and step S303 are respectively executed to construct a speech recognition decoding network, and the acoustic state sequence is decoded by using the speech recognition decoding network to obtain a first speech recognition result; and step S304 is executed to decode the acoustic state sequence by using a general speech recognition model to obtain a second speech recognition result.
[0093] The first speech recognition result and the second speech recognition result can be one or more. In order to ensure the speech recognition effect, at most five speech recognition results output by each model are retained to participate in the determination of the final speech recognition result.
[0094] The general speech recognition model mentioned above is a conventional speech recognition model obtained through massive corpus training. It recognizes the text content corresponding to the speech by learning the characteristics of the speech, rather than having a standardized sentence structure like the above-mentioned speech recognition decoding network. Therefore, the sentence structure that can be recognized by this general speech recognition model is more flexible. Using this general speech recognition model to decode the acoustic state sequence of the speech to be recognized can more flexibly recognize the content of the speech to be recognized, without being restricted by the sentence structure of the speech to be recognized.
[0095] When the speech to be recognized is not a certain sentence pattern in the above-mentioned speech recognition and decoding network, it cannot be correctly decoded by the speech recognition and decoding network, or the first speech recognition result obtained is inaccurate. However, due to the application of the universal speech recognition model, the speech to be recognized can still be recognized and decoded to obtain the second speech recognition result.
[0096] After obtaining the first speech recognition result and the second speech recognition result, step S305 is executed to determine a final speech recognition result at least from the first speech recognition result and the second speech recognition result.
[0097] As an exemplary embodiment, after obtaining the first speech recognition result and the second speech recognition result, based on the acoustic scores of the first speech recognition result and the second speech recognition result, one or more of the highest scores are selected from the first speech recognition result and the second speech recognition result through the acoustic score PK as the final speech recognition result.
[0098] Among them, the acoustic scores of the above-mentioned first speech recognition result and the second speech recognition result refer to the score of the entire decoding result determined according to the decoding scores of each acoustic state sequence element when decoding the acoustic state sequence of the speech to be recognized. For example, the sum of the decoding scores of each acoustic state sequence element can be used as the score of the entire decoding result. The decoding score of the acoustic state sequence element refers to the probability score of the acoustic state sequence element (such as a phoneme or a phoneme unit) being decoded as a certain text. Therefore, the score of the entire decoding result is the probability score of the entire acoustic state sequence being decoded as a certain text. The acoustic score of the speech recognition result reflects the score of the speech being recognized as the speech recognition result, and the score can be used to characterize the accuracy of the speech recognition result.
[0099] Therefore, based on the acoustic scores of one or more first speech recognition results and the acoustic scores of one or more second speech recognition results, the accuracy of each recognition result can be reflected. Through acoustic score PK, that is, through acoustic score comparison, one or more speech recognition results with the highest scores are selected from these recognition results and can be used as the final speech recognition result.
[0100] Figure 3 Steps S301 to S303 in the method embodiment shown correspond to Figure 1 For details of steps S101 to S103 in the method embodiment shown in FIG. Figure 1 The method embodiment is introduced.
[0101] Furthermore, Figure 4 A flow chart of another speech recognition method proposed in an embodiment of the present application is shown.
[0102] and Figure 3 The difference between the above-mentioned speech recognition method and the above-mentioned speech recognition method is that Figure 4 As shown, the speech recognition method proposed in the embodiment of the present application not only uses the constructed speech recognition decoding network and the general speech recognition model to decode the acoustic state sequence of the speech to be recognized to obtain the first speech recognition result and the second speech recognition result, but also executes step S405, and decodes the acoustic state sequence through the pre-trained scenario customized model to obtain the third speech recognition result.
[0103] After respectively obtaining the first speech recognition result, the second speech recognition result and the third speech recognition result, step S406 is executed to determine a final speech recognition result from the first speech recognition result, the second speech recognition result and the third speech recognition result.
[0104] The scenario-customized model mentioned above refers to a speech recognition model obtained by training speech recognition on the speech in the scenario to be recognized. This scenario-customized model has the same model architecture as the general speech recognition model mentioned above. Unlike the general speech recognition model, this scenario-customized model is not trained using a large amount of general corpus, but is trained using corpus from the scenario to be recognized. Therefore, compared with the general speech recognition model, this scenario-customized model has higher sensitivity and higher recognition rate for the speech in the business scenario to be recognized. This scenario-customized model can more accurately recognize speech in specific business scenarios than the general speech recognition model, without being limited to predetermined sentence patterns like the speech recognition decoding network mentioned above.
[0105] Therefore, on the basis of the above-mentioned speech recognition decoding network and the above-mentioned general speech recognition model, a scenario customization model is added, so that the three models can respectively decode the acoustic state sequence of the speech to be recognized, and can perform speech recognition on the speech to be recognized more comprehensively and deeply in a variety of ways.
[0106] For the speech recognition results output by the three models respectively, the introduction of the above step S305 can be referred to as an example. By comparing the acoustic scores of the first speech recognition result, the second speech recognition result and the third speech recognition result, one or more speech recognition results with the highest or higher acoustic scores are selected as the final speech recognition result.
[0107] Figure 4 Steps S401 to S404 in the method embodiment shown correspond to Figure 3 For details of steps S301 to S304 in the method embodiment shown in FIG. Figure 3 The contents of the corresponding method embodiments will not be repeated here.
[0108] The main idea of the above-mentioned multi-model decoding-based speech recognition method is to perform decoding through multiple models and then select the final recognition result from multiple recognition results through acoustic score PK. In actual applications, it is found that when the acoustic scores of the first speech recognition result output by the above-mentioned speech recognition decoding network and the second speech recognition result output by the general speech recognition model, or the third speech recognition result output by the scenario-customized model are relatively close, the first speech recognition result is often PKed by the second or third speech recognition result.
[0109] In fact, when the scores of the speech recognition results output by the three models are similar, the sentence structure of each speech recognition result is basically the same, and there will be differences only in the position of the vertical category keywords. The first speech recognition result contains more accurate vertical category keyword information. If the first speech recognition result is PKed, it may result in inaccurate recognition of the vertical category keywords. Therefore, when the scores of the recognition results output by each model are similar, the recognition result containing the accurate vertical category keywords should be given the upper hand.
[0110] However, according to the description of the above embodiments, the first speech recognition result cannot win.
[0111] In response to the above situation, when performing speech recognition result PK, you can first incentivize the score of the slot where the vertical keyword in the first speech recognition result is located to increase it by a certain proportion, that is, to perform acoustic score incentives on the first speech recognition result, so that when different models output the same sentence structure, the first speech recognition result can win.
[0112] The above introduction only summarizes the idea and necessity of acoustic score excitation. For specific acoustic score excitation processing, please refer to the following embodiments.
[0113] Based on the idea of acoustic score excitation, the present embodiment proposes another speech recognition method, see Figure 5 As shown, the method includes:
[0114] S501: Acquire an acoustic state sequence of a speech to be recognized.
[0115] S502: Decode the acoustic state sequence using a speech recognition decoding network to obtain a first speech recognition result. And, S503: Decode the acoustic state sequence using a universal speech recognition model to obtain a second speech recognition result; the speech recognition decoding network is constructed based on a set of vertical keywords and a sentence decoding network for the scene to be recognized.
[0116] S504: Perform acoustic score incentive on the first speech recognition result.
[0117] S505: Determine a final speech recognition result at least from the first speech recognition result and the second speech recognition result after stimulation.
[0118] Specifically, the specific processing contents of the above steps S501, S502, S503, and S505 can be found in Figures 1-4 The corresponding contents in the corresponding speech recognition method embodiment will not be repeated here.
[0119] The speech recognition and decoding network is constructed based on a set of vertical keywords for the scene to be recognized, and a sentence decoding network obtained by pre-processing sentence structure and grammatical slot definition of the text corpus for the scene to be recognized. The specific content of the speech recognition and decoding network can be found in the above-mentioned embodiments, and the construction process of the network can be found in the detailed description of the embodiments below.
[0120] Different from the speech recognition method introduced in the above embodiment, the speech recognition method proposed in the embodiment of the present application first performs acoustic score incentive on the first speech recognition result before performing acoustic score PK between the first speech recognition result and the second speech recognition result to determine the final speech recognition result.
[0121] Specifically, the first speech recognition result is given an acoustic score incentive, specifically, the acoustic score of the slot where the vertical category keyword in the first speech recognition result is located is incentivized, that is, the acoustic score of the slot where the vertical category keyword in the first speech recognition result is located is scaled according to the incentive coefficient. The specific value of the incentive coefficient is determined by the business scenario and the actual situation of the speech recognition result. For the specific content of the acoustic score incentive, please refer to the introduction of the embodiment of the acoustic score incentive later.
[0122] Through the above-mentioned acoustic score excitation processing, when the scores of the first speech recognition result and the second speech recognition result are similar, that is, when the sentence structure is the same, the acoustic score of the first speech recognition result can be made higher than that of the second speech recognition result, so that the first speech recognition result wins in the acoustic score PK. In this way, when the scores of the first speech recognition result and the second speech recognition result are similar, it is ensured that the vertical category keywords in the final speech recognition result are relatively more accurate recognition results.
[0123] It can be seen that the speech recognition method proposed in the embodiment of the present application can construct a speech recognition decoding network based on the set of vertical category keywords in the business scenario to which the speech to be recognized belongs and the pre-built sentence decoding network in the business scenario. The speech recognition decoding network can be used to decode speech consisting of any sentence pattern and any vertical category keyword in the business scenario to which the speech to be recognized belongs. Therefore, based on the above-mentioned speech recognition decoding network, it is possible to accurately recognize the speech to be recognized, especially to accurately recognize the speech in a specific scenario involving vertical category keywords, and in particular, to accurately recognize the vertical category keywords in the speech.
[0124] At the same time, the speech recognition method proposed in the embodiments of this application utilizes both the aforementioned speech recognition decoding network and a general speech recognition model for decoding and recognition. The general speech recognition model has greater sentence structure flexibility than the aforementioned speech recognition decoding network. By using multiple models to decode the acoustic state sequence of the speech to be recognized, it can perform more comprehensive and in-depth speech recognition of the speech to be recognized in a variety of ways.
[0125] In addition, in the case of multi-model decoding and recognition, the embodiment of the present application performs acoustic score excitation on the speech recognition results output by the speech recognition decoding network. Since the above-mentioned speech recognition decoding network has a higher recognition accuracy for vertical category keywords than the general speech recognition model, based on the above-mentioned acoustic score excitation processing, when the speech recognition results output by the speech recognition decoding network have a score similar to that of the speech recognition results output by the general speech recognition model, the speech recognition results output by the speech recognition decoding network can win, thereby ensuring that the vertical category keywords in the final speech recognition result are correctly recognized.
[0126] As a preferred embodiment, the present application also proposes another speech recognition method, which is relative to Figure 5 The speech recognition method shown in the figure adds a scenario-customized model to decode the acoustic state sequence of the speech to be recognized.
[0127] See also Figure 6 As shown, in addition to executing step S602, decoding the acoustic state sequence using a speech recognition decoding network to obtain a first speech recognition result; and executing step S603, decoding the acoustic state sequence using a general speech recognition model to obtain a second speech recognition result, step S604 is also executed, decoding the acoustic state sequence through a pre-trained scenario customized model to obtain a third speech recognition result.
[0128] The above-mentioned scenario-customized model is obtained by performing speech recognition training on the speech in the scenario to which the speech to be recognized belongs.
[0129] Specifically, the functions of the above-mentioned scene customization model and the beneficial effects brought about by the addition of the scene customization model can be referred to the above-mentioned corresponding Figure 4 The contents of the embodiment of the speech recognition method shown are not repeated here.
[0130] Finally, by executing step S606, a final speech recognition result is determined from the first speech recognition result after stimulation, the second speech recognition result, and the third speech recognition result.
[0131] For example, as described in the above embodiment, by performing acoustic score comparison on the first, second, and third speech recognition results, multiple or one speech recognition results with higher or highest acoustic scores can be selected as the final speech recognition result. For the specific processing process, please refer to the corresponding content in the above embodiment.
[0132] Figure 6 For the specific contents of steps S601 to S603 and S605, please refer to the specific processing contents of the corresponding steps in the above embodiment, which will not be repeated here.
[0133] The aforementioned speech recognition methods, when ultimately comparing and deciding between multiple speech recognition results, rely solely on the acoustic scores of the speech recognition results. This completely ignores the impact of the language model on recognition performance, particularly in the aforementioned speech recognition decoding networks or scenario-specific models. This simplistic and straightforward comparison strategy can significantly impact recognition performance and, in severe cases, lead to false triggering, impacting the user experience.
[0134] In this regard, an embodiment of the present application proposes to perform language model excitation on the speech recognition result based on the acoustic score PK, so that the language model information is integrated into the speech recognition result, and finally, the final speech recognition result is selected through the language score PK.
[0135] As an optional method for determining the speech recognition result, referring to the above embodiment, the first speech recognition result, the second speech recognition result, and the third speech recognition result are obtained respectively through the speech recognition decoding network, the general speech recognition model, and the scenario-customized model, and after the first speech recognition result is acoustically stimulated, the final speech recognition result is determined from the stimulated first speech recognition result, the second speech recognition result, and the third speech recognition result. The following processing can be performed:
[0136] First, according to the acoustic score of the first speech recognition result after acoustic score excitation and the acoustic score of the second speech recognition result, a candidate speech recognition result is determined from the first speech recognition result and the second speech recognition result.
[0137] Specifically, the processing in this step is the same as the acoustic score PK introduced above. The first speech recognition result and the second speech recognition result after acoustic score excitation are subjected to acoustic score PK, and one or more speech recognition results with the highest acoustic score are selected as candidate speech recognition results.
[0138] Then, language model excitation is performed on the candidate speech recognition result and the third speech recognition result respectively.
[0139] Specifically, the above-mentioned language model excitation refers to matching the speech recognition results with the vertical category keywords in the scene to which the speech to be recognized belongs. If the match is successful, the speech recognition results are path expanded, and then the expanded speech recognition results are re-scored based on the clustered language model to complete the excitation of the speech recognition results on the language model. The specific language model excitation processing process will be specifically introduced in the embodiments below.
[0140] Finally, based on the language scores of the candidate speech recognition results after language model excitation and the language scores of the third speech recognition results after language model excitation, a final speech recognition result is determined from the candidate speech recognition results and the third speech recognition results.
[0141] Specifically, referring to the aforementioned acoustic score PK strategy, a language score PK is performed on the candidate speech recognition results after language model stimulation and the third speech recognition result. One or more speech recognition results with the highest language scores are selected as the final speech recognition result. The specific language score PK processing process can be referred to the acoustic score PK processing process described in the above embodiment and will not be detailed here.
[0142] The execution order of the above-mentioned steps such as acoustic score excitation, language model excitation, and selection of candidate speech recognition results can be flexibly adjusted without affecting the overall functional implementation.
[0143] For example, as an optional method for deciding and selecting speech recognition results, refer to the above-mentioned embodiment, after obtaining the first speech recognition result, the second speech recognition result and the third speech recognition result respectively through the speech recognition decoding network, the general speech recognition model and the scenario customization model, the first speech recognition result, the second speech recognition result and the third speech recognition result are respectively subjected to language model excitation; then, based on the language scores of the first speech recognition result, the second speech recognition result and the third speech recognition result after language model excitation, the final speech recognition result is determined from the first speech recognition result, the second speech recognition result and the third speech recognition result through the language score PK.
[0144] The following will describe in detail the processing steps in each of the speech recognition methods described in the above embodiments using different embodiments. It should be understood that since the various speech recognition methods described above have overlapping or identical processing steps, the specific implementations of the processing steps described in the following embodiments are applicable to the corresponding or related processing steps of the speech recognition methods described in the above embodiments.
[0145] First, this application introduces the construction process of the sentence decoding network used to construct the speech recognition decoding network in each of the above-mentioned speech recognition method embodiments. The construction process of the sentence decoding network described below is only an exemplary and preferred implementation scheme. When actually applying the technical solutions of the embodiments of this application, it is also possible to refer to the functions of the sentence decoding network embodied in this embodiment and adopt other methods to construct it.
[0146] The sentence decoding network for the business scenario of the speech to be recognized described in the above embodiments can be constructed by executing the following steps A1-A3:
[0147] A1. Construct a text sentence network by performing sentence pattern induction and grammatical slot definition processing on the corpus data in the scene to which the speech to be recognized belongs.
[0148] The corpus data in the business scenario to which the voice to be recognized belongs is the annotated data of voice collected from actual business scenarios. For example, in the scenario where the user makes a phone call or sends a text message by voice, the user's voice command for making a phone call or sending a text message is collected and annotated as the corpus data in the scenario of making a phone call or sending a text message. Alternatively, it is also possible to manually expand it based on experience to obtain corpus data that conforms to grammatical logic and conforms to the business scenario. For example, "I want to give John a call" and "send a message to Peter for me" are two corpus data use cases. Because sentence induction and grammatical slot definition are performed directly based on the corpus in the future, this stage strives to collect corpus with high coverage, but there is no requirement for coverage of vertical keywords.
[0149] As mentioned above, in a certain business scenario involving vertical keywords, the sentence patterns of user voice are usually regular, or exhaustible. By summarizing these sentence patterns and classifying and defining the grammatical slots in the sentence patterns, a sentence pattern network corresponding to the business scenario can be obtained. The embodiment of this application names it a text sentence pattern network.
[0150] The above-mentioned definition of the grammatical slots in the sentence pattern is specifically to determine the grammatical type of the text slots in the sentence pattern. In an embodiment of the present application, the text slots in the text sentence are divided into ordinary grammatical slots and replacement grammatical slots, wherein the text slots corresponding to non-vertical keywords are defined as ordinary grammatical slots, and the text slots corresponding to vertical keywords are defined as replacement grammatical slots. In the ordinary grammatical slots, the non-vertical keyword content in the text sentence is stored, and in the replacement grammatical slots, the placeholders corresponding to the vertical keywords are stored. According to the different positions of the vertical keyword text slots in the text, the number of ordinary grammatical slots can be one or more, and each vertical keyword text slot corresponds to a replacement grammatical slot.
[0151] The text sentence network constructed in the above way contains network nodes and directed arcs connecting nodes. The text sentence network is defined based on ABNF (Augmented Backus Naur Form). Figure 7 As shown, the directed arcs of the text sentence network carry label information, and the label information corresponds to the placeholder of the replacement grammar slot corresponding to the directed arc or the text of the ordinary grammar slot corresponding to the directed arc.
[0152] Figure 7 It is a text sentence network defined based on the corpus data collected in the scene of making a phone call or sending a text message. <xxxx>The directed arc of the " tag is called a common grammar slot. A common grammar slot contains at least one entry, and all entries must be collected during the text sentence network definition phase. Figure 7 There are two common syntax slots in <phone>and <sth>,in <phone>The text entries before the names of people in the address book in the corresponding scenario corpus use case, such as "I want to give" and "senda message" are <phone>Two entries in the grammar slot; <sth>Indicates the text entry after the address book name in the use case, such as "a call" and "for me". <sth>The two entries in the grammar slot. The directed arc labeled "xxx" is called an alternative grammar slot. This slot does not require an actual entry during the sentence definition phase, but only requires a "#placeholder#" placeholder. The actual entry is dynamically entered when building the speech recognition decoding network. Figure 7 The "name" in the grammar is a replacement grammar slot. The subsequently dynamically created vertical keyword network will be inserted into the replacement grammar slot to form a complete speech recognition decoding network. The last type of directed arc with a "-" label is called a virtual arc. The virtual arc is a directed arc without grammar slot and term information, indicating that the path is optional. The virtual arc must have a corresponding grammar slot to indicate that a decoding path that can skip the grammar slot must be created during the subsequent network construction process. For example, Figure 7 The grammatical slots corresponding to the virtual arcs between nodes 2 and 4 and between nodes 2 and 3 are <sth>Grammar slot, indicating that the recognition result may not go through this grammar slot. Figure 7 The text sentence network defines two sentence patterns, namely <phone>+name and <phone>+name+ <sth>The speech recognition results obtained by the network must conform to one of these two sentence structures. For example, the use case "give a call to John" only contains common grammatical slots. <phone>= "give a call to" and the replacement grammar slot name = "John". Since the common grammar slots have more entries in actual use, for the sake of convenience, only a small number of entries are used to describe related concepts in the subsequent embodiments.
[0153] Furthermore, all grammatical slots in the text sentence network may be numbered with an ID, defined as the slot_id field of the grammatical slot, and a globally unique identifier may be set.
[0154] Ultimately, the sentence patterns, grammatical slots, and entries of common grammatical slots defined by the sentence pattern network together constitute the text sentence pattern network.
[0155] A2. Segmenting the entries in the common grammar slots of the text sentence network and expanding the word nodes according to the segmentation results to obtain a word-level sentence decoding network.
[0156] Specifically, based on the grammatical slot information in the text sentence network constructed above, the common grammatical slot entries are parsed, the entries in the common grammatical slots are segmented and the word nodes are expanded according to the segmentation results to construct a word-level sentence decoding network.
[0157] The word-level sentence decoding network comprises a plurality of nodes and directed arcs between the nodes. When constructing the word-level sentence decoding network, each entry in the common grammar slot is first segmented to obtain the words corresponding to each entry.
[0158] Then, we use the words corresponding to the same entry to perform word node expansion. That is, we concatenate the word segmentation results of the same entry through nodes and directed arcs to obtain the word string corresponding to the entry. The directed arc between two nodes is annotated with the word information obtained by word segmentation. The left and right sides of the colon represent the input and output information, respectively. Here, the input and output information are assumed to be the same.
[0159] Finally, referring to the above-mentioned node expansion method, multiple words after the segmentation of a single entry are connected in series, and the corresponding word strings of different entries in the same common grammar slot are connected in parallel. The grammar slot is replaced with the placeholder "#placeholder#" for placeholding. No expansion is performed, and the network nodes are numbered in sequence, where words with the same starting node identifier share a starting node, and words with the same ending node identifier share an ending node, and a word-level sentence decoding network can be obtained.
[0160] like Figure 2 A relatively simple example of a communication book word-level sentence decoding network diagram, grammar slot <phone>Contains three entries: "I want to give", "send a message to" and "give a call", grammar slot <sth>The three terms "for me", "a call", and "a call with her number" are included. The connection between node 10 and node 18 indicates that you can go directly from node 10 to the end node. The "" on the arc represents silence.
[0161] A3. Replace each word in the ordinary grammar slot of the word-level sentence decoding network with the corresponding pronunciation, and expand the pronunciation node according to the pronunciation corresponding to the word to obtain a pronunciation-level sentence decoding network. The pronunciation-level sentence decoding network serves as the sentence decoding network for the scenario to which the speech to be recognized belongs.
[0162] Specifically, first, each word in the common grammar slot of the word-level sentence decoding network is replaced with the corresponding pronunciation.
[0163] For example, the corresponding relationship between existing words and pronunciations can be queried through a pronunciation dictionary to determine the pronunciation corresponding to each word marked on a directed arc in the common grammar slot of the word-level sentence decoding network. Based on this, the word marked on the directed arc is replaced with the pronunciation corresponding to the word.
[0164] Then, each pronunciation in the word-level sentence decoding network is divided into pronunciation units, and pronunciation nodes are expanded using the pronunciation units corresponding to the pronunciations to obtain a pronunciation-level sentence decoding network.
[0165] That is, for each pronunciation on a directed arc in the word-level sentence decoding network, its pronunciation unit is determined and divided into pronunciation units. This application exemplarily divides the pronunciation into phoneme sequences. For example, the pronunciation of the word "I" is the single phoneme "ay", and the pronunciation of the word "give" is the phoneme string "g ih v".
[0166] On this basis, the pronunciation nodes are expanded and connected in series according to the arrangement order and number of the pronunciation units. For example, similar to the above-mentioned word node expansion, the various phonemes of the same pronunciation are connected in series through nodes and directed arcs in order to obtain a phoneme string corresponding to the pronunciation. The pronunciation is then replaced with the phoneme string corresponding to the pronunciation, and the word-level sentence decoding network is expanded into a pronunciation-level sentence decoding network. During the pronunciation node expansion process, the replacement grammar slot is still not expanded. The pronunciation-level sentence decoding network serves as a sentence decoding network for the business scenario to which the speech to be recognized belongs.
[0167] Figure 8 A schematic diagram showing a simple utterance-level sentence decoding network.
[0168] In addition, in order to facilitate the subsequent network structured storage, subnetwork dynamic insertion update and decoding traversal in the computer, the nodes in the sentence decoding network are numbered in order, and the pronunciation units with the same starting node identifier share a starting node, and the pronunciation units with the same ending node identifier share an ending node. A single node in the network includes a total of 3 attribute fields: id number, number of incoming arcs and number of outgoing arcs, which constitute a node storage triplet. The node incoming arc is a directed arc pointing to the node, and the outgoing arc is a directed arc emitted from the node. A single directed arc in the network includes a total of 4 attribute fields: left node number, right node number, pronunciation information on the arc, and the grammar slot identifier slot_id to which it belongs, and records the total number of nodes and the total number of arcs in the network at the same time.
[0169] Furthermore, the left and right node position information of the grammatical slot in the sentence decoding network can also be recorded, such as Figure 8 As shown, the normal syntax slot <phone>The left node position in the network is 0, the right node position is 10, the left node position of the replacement grammar slot name is 10, the right node is 11, and the normal grammar slot <sth>The left node is 11 and the right node is 16.
[0170] Furthermore, the arcs that meet the conditions in the obtained pronunciation-level sentence decoding network can be merged and optimized, and redundant nodes can be deleted to reduce the complexity of the network. The specific method used is the same as the general decoding network optimization method and will not be described in detail.
[0171] The decoding network obtained after completing the above steps is the sentence decoding network, which can be loaded into the cloud speech recognition service as a global resource. However, since the replacement grammar slot has not yet recorded the actual contact information, it does not yet have actual decoding capabilities.
[0172] Based on the sentence decoding network constructed above, combined with the vertical keyword set in the business scenario to which the speech to be recognized belongs, a speech recognition decoding network can be constructed that is truly used to decode the acoustic state sequence of the speech to be recognized.
[0173] The embodiments of this application will further introduce the construction process of the speech recognition decoding network by example.
[0174] As an implementation method, the embodiment of the present application constructs a speech recognition decoding network by executing the following steps B1-B3:
[0175] B1. Obtain a pre-built sentence decoding network for the scenario to which the speech to be recognized belongs.
[0176] Specifically, a sentence decoding network for the business scenario of the speech to be recognized can be constructed in advance according to the above embodiment, and then the sentence decoding network can be directly called. Alternatively, the sentence decoding network can be constructed in real time when executing step B1, as described in the above embodiment.
[0177] B2. Build a vertical keyword network based on the vertical keyword in the vertical keyword set under the scenario to which the voice to be recognized belongs.
[0178] Specifically, the vertical keyword set for the business scenario to which the voice to be recognized belongs refers to a set consisting of all vertical keywords for the business scenario to which the voice to be recognized belongs. For example, assuming that the voice to be recognized is a voice in a voice call or voice text message scenario, the vertical keyword set for the business scenario to which the voice to be recognized belongs can specifically be a set of names consisting of the names in the user's address book; assuming that the voice to be recognized is a voice in a voice navigation scenario, the vertical keyword set for the business scenario to which the voice to be recognized belongs can specifically be a set of place names consisting of the names of various places in the user's area.
[0179] The embodiment of the present application takes the address book as a vertical keyword set and takes the construction of a name network in the communication as an example to introduce the specific implementation process of constructing a vertical keyword network.
[0180] Constructing a vertical keyword network can refer to the above process of constructing a sentence decoding network. The difference is that the constructed vertical keyword network no longer contains grammatical slot information, and all grammatical slots are assumed to be replacement grammatical slots.
[0181] First, a word-level vertical keyword network is constructed based on each vertical keyword in the vertical keyword set under the business scenario to which the voice to be recognized belongs.
[0182] For example, each name in the address book is segmented separately to obtain the individual words contained in the name, and the individual words are connected in series through nodes and directed arcs to obtain the word string corresponding to the name; the word strings corresponding to different names are connected in parallel to construct a word-level name network.
[0183] like Figure 9 The word-level name network constructed for the three contact terms "Jack Alen", "Tom" and "Peter".
[0184] Then, each word in the word-level vertical keyword network is replaced with the corresponding pronunciation, and the pronunciation nodes are expanded according to the pronunciation corresponding to the word to obtain the pronunciation-level vertical keyword network.
[0185] For example, for each word in the word-level name network, its pronunciation and the phonemes contained in the pronunciation are determined respectively, and each phoneme is connected through nodes and directed arcs to form a phoneme string. Then, the pronunciation is replaced with the phoneme string corresponding to the pronunciation to obtain the pronunciation-level name network.
[0186] The construction of the above-mentioned word-level vertical keyword network and the processing of pronunciation replacement can refer to the corresponding processing content when constructing the sentence decoding network introduced in the above embodiment.
[0187] Through the above introduction, we get the pronunciation-level vertical keyword network, which is the final vertical keyword network constructed. In this network, the node with 0 incoming arcs is the starting node of the network, and the node with 0 outgoing arcs is the ending node of the network.
[0188] like Figure 10 Shown is the Figure 9 The pronunciation-level name network shown in FIG1 is obtained by performing pronunciation replacement on the word-level name network. In this network, node 0 is the starting node of the network and node 8 is the ending node of the network.
[0189] B3. Insert the vertical keyword network into the sentence decoding network to obtain a speech recognition decoding network.
[0190] From the above introduction, it can be seen that by constructing a word-level network and performing pronunciation expansion on the word-level network, the sentence decoding network and vertical keyword network finally constructed are both composed of nodes and directed arcs connecting the nodes, and pronunciation information or placeholders are stored on the directed arcs between the nodes. Specifically, the pronunciation information of the text in the slot is stored on the directed arcs of the ordinary grammar slots of the corresponding sentence decoding network, the placeholders are stored on the directed arcs of the replacement grammar slots of the corresponding sentence decoding network, and the pronunciation information of the text in the slot is stored on the directed arcs of each grammar slot of the corresponding vertical keyword network.
[0191] On this basis, the left and right nodes of the replacement grammar slot of the vertical keyword network and the sentence decoding network are connected respectively through directed arcs. That is, the directed arc where the replacement grammar slot in the sentence decoding network is located is replaced by the vertical keyword network to construct a speech recognition decoding network.
[0192] When the vertical keyword network is connected to the left and right nodes of the replacement grammar slot of the sentence decoding network respectively, in order to ensure that valid pronunciation information is stored on the directed arcs between each pair of adjacent nodes after the connection, the embodiment of the present application connects the right node of each outgoing arc of the starting node of the vertical keyword network with the left node of the replacement grammar slot through a directed arc, and each connected directed arc stores the pronunciation information on each outgoing arc of the starting node of the vertical keyword network, and connects the left node of each incoming arc of the ending node of the vertical keyword network with the right node of the replacement grammar slot through a directed arc, and each connected directed arc stores the pronunciation information on each incoming arc of the ending node of the vertical keyword network, thereby constructing a speech recognition decoding network.
[0193] As a more preferred embodiment, in order to improve the efficiency of inserting the vertical keyword network into the sentence decoding network, the embodiment of the present application stores the unique identifier corresponding to the keyword on the first arc and the last arc of each keyword in the vertical keyword network, and the unique identifier can be exemplarily set to the hash code of the keyword. Figure 10 In the pronunciation-level name network, the directed arc between nodes (0, 1) and the directed arc between nodes (7, 8) respectively store hash codes corresponding to the name "JackAlen".
[0194] Correspondingly, the embodiment of the present application also sets up a set of keyword information that has been connected to the network. In the set of keyword information that has been connected to the network, the unique identifier of the keyword that has been inserted into the sentence decoding network and the left and right node numbers of the directed arc where the unique identifier is located in the sentence decoding network are stored accordingly.
[0195] Exemplarily, the above-mentioned set of keyword information that has been connected to the network can adopt a HashMap storage structure with a key:value structure, where the key is the hash code corresponding to the above-mentioned vertical keyword, and the value is the set of node number pairs of the directed arc where the hash code is located. Initially, the HashMap is empty. During the entire identification service process, the HashMap uniquely saves the hash code and node number pairs of all dynamically passed vertical keyword entries, but does not record the user ID and the mapping relationship between the user ID and the vertical keyword set.
[0196] The above-mentioned setting of the keyword information set that has been inserted into the network can facilitate the clarification of the vertical keyword information that has been inserted into the sentence decoding network, that is, it can clarify the vertical keyword information that already exists in the speech recognition decoding network. In this way, each time a vertical keyword network insertion is performed, the vertical keyword information set that has been inserted can be queried to determine whether the vertical keyword has been inserted. When it is determined that the vertical keyword to be inserted already exists in the speech recognition decoding network, the insertion of the current vertical keyword can be canceled, and the insertion operation of other vertical keywords can be continued.
[0197] Based on the above idea, when inserting the vertical keyword network into the sentence decoding network, each outgoing arc of the starting node of the vertical keyword network is traversed. For each outgoing arc traversed, according to the unique identifier on the outgoing arc and the set of keyword information that has been entered into the network, it is determined whether the keyword corresponding to the unique identifier has been inserted into the sentence decoding network.
[0198] If the keyword corresponding to the unique identifier is not inserted into the sentence decoding network, the right node of the traversed outgoing arc is connected to the left node of the replacement grammar slot through a directed arc, the pronunciation information of the traversed outgoing arc is stored on the directed arc, and the number of incoming arcs or outgoing arcs of the nodes at both ends of the directed arc is updated.
[0199] Exemplarily, when inserting the name network into the sentence decoding network, each outgoing arc of the starting node in the name network is traversed, and for each outgoing arc traversed, the hash code of the name on the outgoing arc is obtained. The hash code is compared with all hash codes in the keyword information set that has been added to the network. If the hash code matches any hash code in the keyword information set that has been added to the network, it means that the name corresponding to the hash code has been inserted into the sentence decoding network. At this time, the outgoing arc is skipped and the hash code judgment of the next outgoing arc is performed.
[0200] If the hash code of the traversed outgoing arc does not match all the hash codes in the set of keyword information that have been entered into the network, it means that the name corresponding to the hash code has not been inserted into the sentence decoding network. At this time, the right node of the traversed outgoing arc is connected to the left node of the replacement grammar slot of the sentence decoding network through a directed arc, and the pronunciation information of the traversed outgoing arc is stored on the connected directed arc, and the number of incoming arcs or outgoing arcs of the nodes at both ends of the directed arc is updated.
[0201] The above process realizes the connection between the right node of each outgoing arc of the start node of the vertical keyword network and the left node of the replacement grammar slot of the sentence decoding network.
[0202] Accordingly, each incoming arc of the end node of the vertical keyword network is traversed in sequence. For each traversed incoming arc, based on the unique identifier on the incoming arc and the set of networked keyword information, it is determined whether the keyword corresponding to the unique identifier has been inserted into the sentence decoding network;
[0203] If the keyword corresponding to the unique identifier is not inserted into the sentence decoding network, the left node of the traversed input arc is connected to the right node of the replacement grammar slot through a directed arc, and the pronunciation information of the traversed input arc is stored on the directed arc.
[0204] Exemplarily, when inserting the name network into the sentence decoding network, each incoming arc of the end node in the name network is traversed, and for each traversed incoming arc, the hash code of the name on the incoming arc is obtained. The hash code is compared with all hash codes in the keyword information set that has been entered into the network. If the hash code matches any hash code in the keyword information set that has been entered into the network, it means that the name corresponding to the hash code has been inserted into the sentence decoding network. At this time, the incoming arc is skipped and the hash code judgment of the next incoming arc is performed.
[0205] If the hash code of the traversed input arc does not match all the hash codes in the set of keyword information that have been entered into the network, it means that the name corresponding to the hash code has not been inserted into the sentence decoding network. At this time, the left node of the traversed input arc is connected to the right node of the replacement grammar slot of the sentence decoding network through a directed arc, and the pronunciation information of the traversed input arc is stored on the connected directed arc, and the number of input arcs or output arcs of the nodes at both ends of the directed arc is updated.
[0206] The above process realizes the connection between the left node of each incoming arc of the end node of the vertical keyword network and the right node of the replacement grammar slot of the sentence decoding network.
[0207] Through the above processing, each keyword in the vertical keyword network is inserted into the sentence decoding network. In specific implementation, the execution order of the above-mentioned insertion of the start node and end node of the vertical keyword network can be flexibly arranged. For example, the insertion operation can be performed on the start node of the vertical keyword network first, or on the end node of the vertical keyword network first, or both can be performed simultaneously.
[0208] Furthermore, when the keyword in the vertical keyword network is inserted into the sentence decoding network, that is, after the right node of the first arc of the network path where the keyword is located is connected to the left node of the replacement grammar slot of the sentence decoding network, and the left node of the last arc of the network path where the keyword is located is connected to the right node of the replacement grammar slot of the sentence decoding network, the embodiment of the present application will store the unique identifier of the keyword and the left and right node numbers of the directed arc where the unique identifier is located in the sentence decoding network in the corresponding manner in the keyword information set that has been connected to the network.
[0209] For example, for Figure 10 The pronunciation-level name network, when the node 1 of the network path where the name "Jack Alen" is located is connected to Figure 8 The left node 10 of the replacement grammar slot of the sentence decoding network shown is connected to the node 7 of the network path where "Jack Alen" is located. Figure 8 After the right node 11 of the replacement grammar slot of the sentence decoding network shown is connected, the hash code of the name "JackAlen" is used as the key and the left and right node numbers of the directed arc where the hash code is located are used as the value, and stored in the networked keyword information set.
[0210] Since the embodiment of the present application builds a speech recognition and decoding network on a cloud server, each user can upload a vertical keyword set to the cloud server, and the cloud server can build a large-scale or ultra-large-scale speech recognition and decoding network based on the vertical keyword set uploaded by each user, so as to meet the calling needs of various users.
[0211] The setting of the keyword information set that has been added to the network can improve the efficiency of inserting vertical keywords into the speech recognition decoding network, and at the same time can facilitate the selection of a specific decoding path based on the current user's speech recognition needs.
[0212] For example, in a business scenario where a user makes a phone call by voice, the voice recognition and decoding network constructed according to the above scheme not only includes the address book entry information dynamically passed in during the current session, but also includes the address book information of other historical sessions, and is uniformly stored in the form of a hash code in the set of keyword information that has been connected to the network.
[0213] When recognizing the current conversation, it should be done within the context of the currently passed address book to accurately identify the contact the user wants to call. At this point, the speech recognition decoding network path needs to be updated so that the decoding path is limited to the currently passed address book.
[0214] The specific implementation method is:
[0215] Traverse each unique identifier in the keyword information set that has been connected to the network; if the traversed unique identifier is not the unique identifier of any keyword in the vertical keyword set under the business scenario to which the voice to be recognized belongs, then disconnect the directed arc between the left and right node numbers corresponding to the unique identifier.
[0216] Still taking the above-mentioned business scenario of the user making a phone call as an example, in order to decode the user's voice, a speech recognition decoding network is constructed using the user's address book and sentence decoding network. After the current address book is inserted into the speech recognition decoding network, each hash code in the keyword information set that has been connected to the network is traversed. If the traversed hash code belongs to the hash code of the name in the address book that is currently passed in, no processing is performed; if it is not the hash code of the name in the address book that is currently passed in, the left and right node numbers corresponding to the hash code are determined by querying the keyword information set that has been connected to the network, and the directed arc between the left and right node numbers is disconnected. Then, for the speech recognition decoding network that actually participates in decoding, only the decoding path of the currently passed in address book is connected, so it can only decode the speech recognition result of calling any person in the current address book, which is also in line with user expectations.
[0217] It can be seen that after the above processing, when decoding the current speech, the decoding path can be limited to the range of the vertical category keyword set passed in this time, which is beneficial to narrowing the path search range, improving decoding efficiency and reducing decoding errors.
[0218] From the introduction of the embodiments of the above-mentioned speech recognition methods, it can be seen that the speech recognition decoding network constructed based on the vertical category keyword set and the sentence decoding network is the main network for realizing speech recognition involving vertical category keywords.
[0219] The inventors of this application discovered during their research that the speech recognition decoding network has two deficiencies: one is the problem of false triggering of vertical keywords, and the other is the problem of insufficient network coverage.
[0220] The false triggering problem is the most common problem in speech recognition based on a fixed-sentence speech recognition decoding network, and it is also the most difficult problem to solve. False triggering means that the actual content of an audio segment is not the sentence pattern in the fixed-sentence speech recognition decoding network, but the final result is the result of the fixed-sentence speech recognition decoding network. Vertical keyword false triggering means that there is no vertical keyword in the actual result or the vertical keyword is not in the vertical keyword set passed in this time, but the final result is that the result of the speech recognition decoding network wins, giving an incorrect vertical keyword. For example, in a phone call scenario, there is no name in the actual result or the name is not in the address book passed in this time, but the final result is that the result output by the speech recognition decoding network wins, giving an incorrect name.
[0221] False triggering of vertical keywords is generally divided into the following four types: (1) False triggering when the vertical keyword and the real result have the same pronunciation; (2) False triggering when the vertical keyword and the real result have similar pronunciation; (3) False triggering due to the large difference in pronunciation between the vertical keyword and the real result caused by the incentive strategy; (4) There is no vertical keyword in the real result, but the speech recognition decoding network recognizes the false triggering of the vertical keyword.
[0222] Among them, regarding the fourth false triggering situation mentioned above, its root cause can be attributed to insufficient sentence coverage in the speech recognition decoding network.
[0223] In conjunction with this vertical keyword false triggering situation, the present application embodiment proposes a corresponding solution to the problem of insufficient sentence coverage of the speech recognition decoding network. For other false triggering situations, other solutions will be used in subsequent embodiments to solve and optimize them.
[0224] As discussed in the previous article, the speech recognition decoding network is built based on the sentence decoding network, that is, it is built based on sentence patterns. The advantage of this network construction method is that it can accurately match speech sentence patterns. The disadvantage is that if the sentence pattern is not in the network, there is no way to recognize it. The corpus used to build the network often cannot cover all sentence patterns in all scenarios, so the speech recognition decoding network will introduce a problem, that is, the sentence pattern coverage is not high enough.
[0225] The general speech recognition model is trained based on massive amounts of data and features scalable sentence structures with a rich variety of sentence structures. In specific business scenarios, errors in the general speech recognition model's results are primarily due to errors in the recognition of vertical keywords. This is primarily because the general speech recognition model incorporates the language scores from the training corpus, and the corpus often doesn't accurately reflect the scores of vertical keywords. Although the general speech recognition model may identify incorrect vertical keywords, its sentence structure is correct, primarily because the sentence structure can be fitted using the training data.
[0226] In summary, the sentence structure information in the recognition results of the general speech recognition model is more reliable, while the vertical keyword information in the recognition results of the speech recognition decoding network is more reliable. Based on this idea, this case proposes a sentence structure based on the general speech recognition model to solve the problem of low sentence structure coverage in the speech recognition decoding network.
[0227] The specific solution is to correct the first speech recognition result based on the second speech recognition result, on the premise that the speech recognition decoding network and the general speech recognition model constructed according to the technical ideas of this application are used to decode the acoustic state sequence of the speech to be recognized to obtain the first speech recognition result and the second speech recognition result.
[0228] Combined with the previous discussion, we can see that, relatively speaking, the vertical keyword content in the first voice recognition result is more accurate, while the sentence structure information in the second voice recognition result is more accurate. Therefore, when using the second voice recognition result to correct the first voice recognition result, the correction should be made to the non-vertical keyword content in the first voice recognition result to make its sentence structure more accurate.
[0229] Therefore, the embodiment of the present application divides the content in the first voice recognition result into vertical keyword content and non-vertical keyword content. Correspondingly, according to the content correspondence between the second voice recognition result and the first voice recognition result, the content in the second voice recognition result is divided into reference text content and non-reference text content, wherein the reference text content in the second voice recognition result refers to the text content in the second voice recognition result that matches the non-vertical keyword content in the first voice recognition result.
[0230] The text content that matches the non-vertical keyword content in the first voice recognition result can specifically be the character string that is most similar to the non-vertical keyword content in the first voice recognition result, or the character string content whose similarity is greater than a set threshold.
[0231] Based on the above-mentioned text content division, the above-mentioned correction of the first voice recognition result according to the second voice recognition result is specifically to use the reference text content in the second voice recognition result to correct the non-vertical keyword content in the first voice recognition result to obtain the corrected first voice recognition result.
[0232] The above correction process can be specifically implemented by executing the following steps C1-C3:
[0233] C1. Determine vertical keyword content and non-vertical keyword content from the first voice recognition result, and determine text content corresponding to the non-vertical keyword content in the first voice recognition result from the second voice recognition result as reference text content.
[0234] As a preferred implementation, the embodiment of the present application matches the first speech recognition result and the second speech recognition result based on the edit distance algorithm, and determines the reference text content from the second speech recognition result.
[0235] Specifically, first, an edit distance matrix between the first speech recognition result and the second speech recognition result is determined according to an edit distance algorithm. The edit distance matrix includes the edit distance between each character in the first speech recognition result and each character in the second speech recognition result.
[0236] Then, based on the edit distance matrix and the non-vertical keyword content in the first speech recognition result, the text content corresponding to the non-vertical keyword content in the first speech recognition result is determined from the second speech recognition result as the reference text content.
[0237] According to the position of the vertical category keywords in the first voice recognition result, the non-vertical category keyword content in the first voice recognition result may be divided into a character string before the vertical category keyword and / or a character string after the vertical category keyword.
[0238] Taking the string preceding the vertical category keyword as an example, based on this partial string and the edit distance matrix described above, the string with the smallest edit distance to this partial string is selected from the second speech recognition result. This is the reference text content corresponding to this partial string. Similarly, for the string following the vertical category keyword, the corresponding reference text content can also be determined using the above method.
[0239] C2. Determine the corrected non-vertical keyword content based on the reference text content in the second voice recognition result and the non-vertical keyword content in the first voice recognition result.
[0240] As an optional implementation, the embodiment of the present application determines the target text content in the second voice recognition result or the non-vertical keyword content in the first voice recognition result as the corrected non-vertical keyword content based on the character difference between the reference text content in the second voice recognition result and the non-vertical keyword content in the first voice recognition result;
[0241] Among them, the target text content in the second voice recognition result refers to the text content in the second voice recognition result that corresponds to the position of the non-vertical keyword content in the first voice recognition result.
[0242] Specifically, the position of the target text content in the second voice recognition result is the same as the position of the non-vertical keyword content in the first voice recognition result.
[0243] For example, assuming that the non-vertical category keyword content in the first speech recognition result is the text content located before the vertical category keyword in the first speech recognition result, then the target text content in the second speech recognition result is specifically the text content before the position of the vertical category keyword in the first speech recognition result is mapped to the position in the second speech recognition result. The mapping of the position of the vertical category keyword in the first speech recognition result to the second speech recognition result can be achieved based on the above-mentioned edit distance matrix.
[0244] Below, we'll use a phone call scenario as an example to explain the process for determining corrected non-name content. The following explanation uses the example of a person's name in the first speech recognition result being mid-sentence, focusing on the correction process for the character string preceding the name. As described below, corresponding corrections can also be made to the character string following the name.
[0245] After determining the reference text content from the second voice recognition result based on the edit distance matrix between the first voice recognition result and the second voice recognition result, and the non-vertical keyword content in the first voice recognition result, the reference text content in the second voice recognition result is compared with the non-vertical keyword content in the first voice recognition result to determine whether the reference text content in the second voice recognition result is the same as the non-vertical keyword content in the first voice recognition result.
[0246] If they are the same, the target text content in the second speech recognition result is determined as the corrected non-vertical keyword content;
[0247] If they are different, determining whether the second speech recognition result has more characters than the non-vertical keyword content in the first speech recognition result, and whether the difference in the number of characters between the two does not exceed a set threshold;
[0248] If the second speech recognition result has more characters than the non-vertical keyword content in the first speech recognition result, and the difference in the number of characters between the two does not exceed a set threshold, the target text content in the second speech recognition result is determined to be the corrected non-vertical keyword content;
[0249] If the number of characters in the second voice recognition result is less than that in the non-vertical keyword content in the first voice recognition result, and / or the difference in the number of characters between the two exceeds a set threshold, the non-vertical keyword content in the first voice recognition result is determined as the corrected non-vertical keyword content.
[0250] Specifically, see Figure 11 The correction processing process shown calculates the local maximum subsequence of the character string before the name and the second speech recognition result based on the edit distance matrix between the first speech recognition result and the second speech recognition result, and the character string before the name in the first speech recognition result. The local maximum subsequence is the reference text content determined from the second speech recognition result.
[0251] Then, it is determined whether the largest subsequence is the character string before the person's name, that is, whether the content of the reference text is the same as the character string before the person's name.
[0252] If they are the same, the target text content in the second speech recognition result is determined to be the corrected string before the name. In this case, the target text content in the second speech recognition result, specifically the position of the name in the first speech recognition result, is mapped to the text content before the position in the second speech recognition result based on the edit distance matrix between the first and second speech recognition results.
[0253] If they are different, determining whether the second speech recognition result has more characters than the character string before the person's name;
[0254] If the number of characters in the second voice recognition result is not more than the number of characters in the character string before the name, the character string before the name in the first voice recognition result is kept unchanged, that is, the character string before the name in the first voice recognition result is determined as the revised character string before the name.
[0255] If the number of characters in the second voice recognition result is greater than the number of characters in the character string before the name, it is further determined whether the difference in the number of characters between the two exceeds a set threshold, specifically, whether the number of characters in the second voice recognition result exceeds the number of characters in the character string before the name by more than 20% of the number of characters in the character string before the name.
[0256] If the threshold is exceeded, the character string before the name in the first speech recognition result is kept unchanged, that is, the character string before the name in the first speech recognition result is determined as the corrected character string before the name.
[0257] If it does not exceed the set threshold, the target text content in the second speech recognition result is determined as the corrected character string before the name.
[0258] Similarly, for the character string after the name in the first speech recognition result, the above description can also be referred to to determine the corrected character string after the name.
[0259] C3. Utilize the corrected non-vertical keyword content and the vertical keyword content to combine and obtain a corrected first speech recognition result.
[0260] Specifically, the corrected non-vertical keyword content and the vertical keyword content are combined according to the positional relationship between the original non-vertical keyword content and the vertical keyword content, and the obtained combination result is used as the corrected first speech recognition result.
[0261] For example Figure 11 As shown, the corrected character string before the name, the character string after the name, and the corrected character string after the name are sequentially concatenated to obtain a corrected first speech recognition result.
[0262] By correcting the first speech recognition result according to the above-mentioned processing procedure, the output result of the speech recognition decoding network can not only contain more accurate vertical keyword information, but also be able to integrate the accurate sentence information recognized by the general speech recognition model, thereby making the recognition result of the speech recognition decoding network more accurate and the sentence coverage higher. At the same time, it can solve the problem of false recognition triggering caused by the low sentence coverage of the speech recognition decoding network.
[0263] As described in the above embodiment, when multiple models decode the acoustic state sequence of the speech to be recognized respectively, in order to make the result of the speech recognition decoding network win in the final PK, the result of the speech recognition decoding network will be stimulated, including acoustic stimulation and language stimulation. During the stimulation process, if the stimulation is inappropriate, it may cause insufficient stimulation, or it may cause excessive stimulation and cause false triggering, resulting in the second and third types of vertical keyword false triggering as described above, namely: (2) false triggering when the pronunciation of the vertical keyword and the real result is similar; (3) false triggering introduced by the stimulation strategy due to the large difference in pronunciation between the vertical keyword and the real result.
[0264] In order to improve the above problems, so that through reasonable incentives, we can ensure that the results of the speech recognition and decoding network are not eliminated by the results of other models, and try to avoid the false triggering of the speech recognition and decoding network due to excessive incentives, the embodiment of the present application studies the above incentive scheme and proposes an optimal incentive scheme.
[0265] The embodiment of the present application first proposes that when a speech recognition decoding network and a general speech recognition model are used to decode the acoustic state sequence of the speech to be recognized, when determining the final speech recognition result from the first speech recognition result output by the speech recognition decoding network and the second speech recognition result output by the general speech recognition model, reference can be made to Figure 12 The process shown in FIG. 1 is to determine the final speech recognition result by processing the following steps D1-D3:
[0266] D1. Determine a degree of match between the first speech recognition result and the second speech recognition result by comparing the first speech recognition result and the second speech recognition result, and determine a confidence level of the first speech recognition result based on the degree of match between the first speech recognition result and the second speech recognition result.
[0267] Specifically, the character edit distance between the first speech recognition result and the second speech recognition result is calculated according to an edit distance algorithm, thereby determining the matching degree between the two.
[0268] After determining the matching degree between the first speech recognition result and the second speech recognition result, the confidence level of the first speech recognition result is further determined based on the matching degree between the first speech recognition result and the second speech recognition result.
[0269] For example, when determining the confidence level of the first speech recognition result, it is first determined whether the matching degree between the first speech recognition result and the second speech recognition result is greater than a set matching degree threshold. The matching degree threshold is a value with the highest recognition contribution rate calculated from statistics of multiple test sets.
[0270] If the match score is greater than the set matching threshold, the confidence level of the first speech recognition result is calculated based on the acoustic scores of each frame of the first speech recognition result. That is, the acoustic scores of each frame of the first speech recognition result are accumulated or weighted, and the accumulated result is used as the confidence level of the first speech recognition result.
[0271] If it is not greater than the set matching threshold, a decoding network is constructed using the vertical category keyword content in the first speech recognition result and the second speech recognition result, and the acoustic state sequence is re-decoded using the decoding network, and the first speech recognition result is updated using the decoding result; based on the acoustic scores of each frame of the updated first speech recognition result, the confidence of the first speech recognition result is calculated and determined.
[0272] Specifically, if the match between the first and second speech recognition results is no greater than a set match threshold, a miniature decoding network is constructed using the vertical category keywords in the first speech recognition result and the sentence structure of the second speech recognition result (that is, the content in the second speech recognition result excluding the content corresponding to the vertical category keywords). This miniature decoding network has only one decoding path.
[0273] The miniature decoding network is used to re-decode the acoustic state sequence of the speech to be recognized to obtain a decoding result, which is used as a new first speech recognition result. Then, the acoustic scores of each frame of the updated first speech recognition result during decoding are accumulated or weighted to form the confidence level of the final first speech recognition result.
[0274] When the confidence level of the first speech recognition result is greater than a preset confidence level threshold, step D2 is executed to select a final speech recognition result from the first speech recognition result and the second speech recognition result based on the acoustic score of the first speech recognition result and the acoustic score of the second speech recognition result.
[0275] Specifically, the above-mentioned confidence threshold refers to a score threshold determined through experiments, which enables the first speech recognition result to win when compared with other speech recognition results in a score competition.
[0276] If the confidence level of the first speech recognition result is greater than the preset confidence level threshold, it indicates that the first speech recognition result, due to its own confidence level, can withstand being easily eliminated in an acoustic score comparison with other speech recognition results. Therefore, in this case, an acoustic score comparison can be performed directly based on the acoustic scores of the first speech recognition result and the second speech recognition result, and one or more speech recognition results with the highest acoustic scores can be selected as the final speech recognition result.
[0277] When the confidence level of the first speech recognition result is not greater than a preset confidence threshold, step D3 is executed to perform acoustic score excitation on the first speech recognition result, and a final speech recognition result is selected from the first speech recognition result and the second speech recognition result based on the acoustic score of the first speech recognition result after excitation and the acoustic score of the second speech recognition result.
[0278] Specifically, if the confidence of the first speech recognition result is not greater than the preset confidence threshold, it means that the confidence of the first speech recognition result is low, that is, its acoustic score is low. When the acoustic score is PKed with other speech recognition results, it will be PKed by other speech recognition results, so that the more accurate vertical keyword information contained in the first speech recognition result will be lost in the final selected speech recognition result, which may cause recognition errors, especially vertical keyword recognition errors.
[0279] At this time, in order to ensure that the first voice recognition result is not easily PKed out in the subsequent acoustic score PK, the embodiment of the present application performs acoustic score incentives on the first voice recognition result, specifically, incentivizes the acoustic score of the slot where the vertical category keyword in the first voice recognition result is located, so that the acoustic score of the slot where the vertical category keyword in the first voice recognition result is located is increased by a certain proportion, thereby improving the acoustic score of the first voice recognition result, to ensure that the first voice recognition result is not easily PKed out in the subsequent acoustic score PK with other voice recognition results.
[0280] After the first speech recognition result is acoustically scored, an acoustic score comparison can be performed on the two based on the acoustic scores of the first speech recognition result and the second speech recognition result after the stimulation, and one or more speech recognition results with the highest acoustic scores are selected as the final speech recognition results.
[0281] The following describes a specific implementation scheme for performing acoustic score incentive on the first speech recognition result involved in the above embodiments.
[0282] The acoustic score of the first speech recognition result is incentivized, that is, the acoustic score of the slot where the vertical category keyword in the first speech recognition result is located is incentivized. Specifically, the acoustic score of the slot where the vertical category keyword in the first speech recognition result is located is multiplied by an excitation coefficient, and then the acoustic score of the first speech recognition result is recalculated based on the acoustic score of the vertical category keyword after excitation.
[0283] Specifically, by performing the following processing E1-E3, acoustic score excitation for the first speech recognition result can be achieved:
[0284] E1. Determine an acoustic excitation coefficient based on at least the vertical keyword content and the non-vertical keyword content in the first speech recognition result.
[0285] The acoustic excitation coefficient determines the intensity of the acoustic score excitation for the first speech recognition result. If the excitation coefficient is too large, it will cause over-excitation, resulting in the recognition false triggering problem described above; if the excitation coefficient is too small, the excitation purpose will not be achieved, and the first speech recognition result may be PKed by other speech recognition results in the score PK.
[0286] Therefore, the determination of the acoustic excitation coefficient is the key to solving the problem of false triggering of vertical keywords mentioned in the above embodiments, as well as the problem of losing important vertical keyword information during acoustic PK.
[0287] The embodiment of the present application stipulates that when determining the acoustic excitation coefficient, it should be determined based on at least the vertical keyword content and non-vertical keyword content in the first speech recognition result. In addition, it should also be determined in combination with actual business scenarios, empirical parameters under actual business scenarios, etc.
[0288] As an optional implementation, the acoustic excitation coefficient can be calculated and determined based on the acoustic score excitation prior coefficient in the business scenario to which the voice to be recognized belongs, the number of characters and the number of phonemes in the vertical category keywords in the first voice recognition result, and the total number of characters and the total number of phonemes in the first voice recognition result.
[0289] Specifically, in this embodiment, the acoustic excitation coefficient RC is calculated according to the following formula:
[0290] RC=1.0+α*{β*(SlotWC / SentWC)+(1-β)*(SlotPC / SentPC)}
[0291] Among them, α is the scenario prior coefficient. In the embodiment of the present application, it is the excitation prior coefficient under the business scenario to which the voice to be recognized belongs. The positive and negative values of the coefficient represent positive excitation and negative excitation, respectively. The prior coefficient α is an open parameter that can be dynamically set in each recognition session according to the needs of the recognition system. For example, based on natural language processing (NLP) technology, the upper-level system can predict the user's behavioral intention through the context of the user interaction and adjust the coefficient in real time and dynamically to meet the requirements of various scenarios.
[0292] In addition, since the vertical keywords in the business scenarios involving vertical keywords have inconsistent word counts and word pronunciation phoneme lengths, in order to improve the generalization ability of the acoustic excitation coefficient and its adaptability in different slots and avoid unreasonable excitation, the design of the acoustic excitation coefficient also fully considers the number of words (word count in slot, SlotWC) and the number of phonemes (phoneme count in slot, SlotPC) in the slot where the vertical keyword is located (i.e., the number of characters and the number of phonemes in the vertical keyword), as well as the proportional relationship between the number of words in the sentence (word count in sentence, SentWC) and the number of phonemes (phoneme count in sentence, SentPC) (i.e., the total number of characters and the total number of phonemes in the first speech recognition result), and sets influence weights β for the number of characters and the number of phonemes respectively, so that the sum of the influence weights of characters and phonemes is 1. This approach limits the range of the acoustic excitation coefficient and achieves adaptability of excitation under sentences and keywords of different lengths. In certain scenarios, if the calculated excitation coefficient RC is greater than 1.0, the recognition result will be more biased towards the first speech recognition result output by the speech recognition decoding network, and vice versa. Furthermore, acoustic score excitation only applies to the acoustic scores of the slots containing vertical keywords, eliminating interference from irrelevant context and thus avoiding false triggering of recognition caused by excessive excitation.
[0293] As another optional implementation, the score confidence of the vertical category keyword content in the first speech recognition result can be calculated based on the number of phonemes and acoustic scores of the vertical category keyword content in the first speech recognition result, and the number of phonemes and acoustic scores of the non-vertical category keyword content in the first speech data result; then, the acoustic excitation coefficient can be determined based on the score confidence of the vertical category keyword content in the first speech recognition result.
[0294] Specifically, by analyzing the acoustic state sequences of speech recognition results, we found that the state sequence scores corresponding to incorrectly recognized words are relatively low, which is why correctly recognized results often have the highest scores. In short, the average score of the acoustic sequence of incorrectly recognized results is relatively low.
[0295] In business scenarios involving vertical keywords, even if the vertical keyword is incorrectly recognized, the overall sentence structure is still recognized normally. In other words, in the entire speech recognition result, the local recognition effect of the vertical keyword is worse than the recognition effect of the non-vertical keyword portion of the entire sentence. Based on this idea, this case proposes to solve the problem of false triggering caused by similar pronunciations of vertical keywords and official results by using the confidence score of the vertical keyword content.
[0296] The vertical keyword score confidence scheme is to separate the vertical keyword and non-vertical keyword in the first speech recognition result output by the speech recognition decoding network, and calculate the total acoustic score of the vertical keyword and the ratio of the number of effective acoustic modeling factors occupied by the vertical keyword, as well as the total acoustic score of the non-vertical keyword part and the ratio of the number of effective acoustic modeling factors occupied by the non-vertical keyword part, and then divide the two to obtain the score confidence value of the vertical keyword.
[0297] The score confidence S of the above vertical keywords c It can be calculated by the following formula:
[0298]
[0299] The acoustic score of the vertical keyword in the first speech recognition result is recorded as S p The number of valid acoustic phonemes occupied by the slot where the vertical keyword is located is recorded as N p , the total acoustic score of non-vertical keywords is recorded as S a The number of valid acoustic phonemes occupied by the slot where the non-vertical keyword is located is recorded as N a .
[0300] After obtaining the score confidence of the vertical keyword content in the first speech recognition result, the acoustic excitation coefficient for stimulating the vertical keyword can be determined based on the score confidence.
[0301] As an exemplary implementation method, the score confidence of the vertical category keyword content in the first speech recognition result is compared with a preset confidence threshold. The confidence threshold is determined based on the probability of false recognition triggering caused by the acoustic excitation coefficient. When the score confidence of the vertical category keyword content in the first speech recognition result is greater than the confidence threshold, it can be considered that the vertical category keyword is prone to false triggering, and when determining the acoustic excitation coefficient for the vertical category keyword, the acoustic excitation coefficient should be lowered; when the score confidence of the vertical category keyword content in the first speech recognition result is not greater than the confidence threshold, it can be considered that the vertical category keyword is prone to being PKed, and when determining the acoustic excitation coefficient for the vertical category keyword, the acoustic excitation coefficient should be increased.
[0302] Furthermore, when determining the acoustic excitation coefficient based on the score confidence of the vertical keyword content in the first speech recognition result, the score confidence of the vertical keyword content in the first speech recognition result and the relationship between the predetermined acoustic excitation coefficient and the recognition effect and recognition false trigger can also be combined to determine the acoustic excitation coefficient.
[0303] Specifically, the embodiment of the present application analyzes the recognition results of multiple test sets, statistically analyzes the impact of the size of the acoustic excitation coefficient on the recognition effect and recognition false triggering, and determines the relationship between the acoustic excitation coefficient and the recognition effect and recognition false triggering.
[0304] Based on the relationship between the acoustic excitation coefficient, recognition effect, and false positives, a value that strikes a balance between improving the recognition rate and reducing false positives is selected when determining the specific value of the acoustic excitation coefficient. The principle of selection is to ensure that the number of false positives is significantly smaller than the number of false positives that improve the recognition effect. The acoustic excitation coefficient selected in this embodiment of the application is an excitation coefficient that results in a number of false positives that is one percent of the number that promotes the improvement in recognition effect.
[0305] E2. Using the acoustic excitation coefficient, update the acoustic score of the vertical keyword content in the first speech recognition result.
[0306] Specifically, the acoustic binary of the slot where the vertical keyword content in the first speech recognition result is located is multiplied by the acoustic excitation coefficient determined in the above step to obtain the updated acoustic score of the vertical keyword content.
[0307] E3. Recalculate and determine the acoustic score of the first speech recognition result based on the updated acoustic scores of the vertical category keyword content in the first speech recognition result and the acoustic scores of the non-vertical category keyword content in the first speech recognition result.
[0308] Specifically, the acoustic score of the vertical category keyword content is used to replace the acoustic score of the vertical category keyword content in the first speech recognition result, and then the acoustic scores of each character of the first speech recognition result are re-summed or weighted to obtain the updated acoustic score of the first speech recognition result.
[0309] Below, the specific implementation scheme of the language model excitation involved in the above embodiments is introduced. In the following embodiment, taking the language model excitation of the third speech recognition result as an example, the specific processing content of the language model excitation is introduced. It should be understood that the specific processing process of language model excitation is not limited by the excitation object, and the language model excitation scheme can also be applied to the excitation of other speech recognition results. For example, the implementation scheme of language model excitation described in the following embodiment is also applicable to the language model excitation of the candidate speech recognition results selected from the first speech recognition result and the second speech recognition result.
[0310] Language model excitation is to recalculate the score of the speech recognition result through the language model, so that the score of the speech recognition result carries a language component.
[0311] The language model incentive mechanism is mainly implemented through two aspects: one is the clustering class language model, and the other is to expand the path of the speech recognition result based on the strategy of matching vertical category keywords with the pronunciation sequence of the speech recognition result, and determine the language score of the speech recognition result based on the expanded path and the above-mentioned clustering language model.
[0312] First, let's introduce the clustering model. In specific speech recognition scenarios involving vertical keywords, such as making phone calls, sending text messages, checking the weather, and navigation, we can restrict the vertical keywords in each scenario to a limited range through enumeration or user-provided methods. The context in which the vertical keywords appear typically follows a specific sentence structure.
[0313] In addition to using general training corpus, the clustering class language model also performs special processing for this type of specific sentence patterns or statements. The clustering language model will define a class for all specific scenarios, and each type of scenario will use a special word (class) to mark and distinguish it as the category label corresponding to the scenario. After all category labels are defined, vertical category keywords such as names of people, cities, audio and video names in the training corpus will be replaced with corresponding category labels to form target corpus. These target corpora will be added to the original training corpus and used again for speech recognition training of the above-mentioned clustering language model. This processing method makes the special word class represent the probability of a class of words, so the probability of the N-gram language model where the special word class is located in the clustering model will be significantly higher than the probability of the specific vertical category keyword itself.
[0314] Based on the above-mentioned category labels, the embodiment of the present application performs path expansion on the third voice recognition result according to the vertical keyword set under the business scenario to which the voice to be recognized belongs and the category label corresponding to the business scenario.
[0315] Exemplarily, first, the vertical category keywords in the third speech recognition result are compared with the vertical category keywords in the vertical category keyword set in the business scenario to which the speech to be recognized belongs.
[0316] As mentioned above, in specific voice recognition business scenarios that contain vertical category keywords, such as making phone calls, checking the weather, and navigation, the vertical category keywords will be confined to a limited range. Through enumeration, user provision, and other methods, all vertical category keywords in the corresponding business scenario can be used as a static resource. Utilizing the pronunciation dictionary resources, the vertical category keywords for the business scenario to which the voice to be recognized belongs and the pronunciation string information of the third voice recognition result are generated respectively. By comparing the pronunciation information of the vertical category keywords in the third voice recognition result with the pronunciation information of the vertical category keywords in the vertical category keyword set for the business scenario to which the voice to be recognized belongs, it is determined whether the vertical category keywords in the third voice recognition result match any vertical category keywords in the vertical category keyword set for the business scenario to which the voice to be recognized belongs.
[0317] If the vertical category keyword in the third speech recognition result matches any vertical category keyword in the vertical category keyword set under the business scenario to which the speech to be recognized belongs, a new path is extended between the left and right nodes of the slot where the vertical category keyword in the third speech recognition result is located, and the category label corresponding to the business scenario to which the speech to be recognized belongs is stored on the new path.
[0318] For example, Figure 13 The following is the result of speech recognition in the phone call business scenario. <s> Call Zhang San< / s> ” state network (lattice) diagram.
[0319] For the name "Zhang San" in the speech recognition result, when it is confirmed through pronunciation matching that it matches "Zhang San" in the address book uploaded by the user, Figure 13 In the state network shown, a new path is extended between the left and right nodes of the slot where "Zhang San" is located. The new path shares the starting node and the ending node with "Zhang San" in the original state network, and the category label "class" corresponding to the current business scenario is marked on the new path. Specifically, the "class" can be "name". The state network after the path extension is as follows Figure 14 shown.
[0320] According to the above processing, after completing the path expansion of the third speech recognition result, the language model scores of the third speech recognition result and the extended path of the third speech recognition result are determined respectively based on the recognition results of the training corpus by the clustering language model corresponding to the category label corresponding to the business scenario to which the speech to be recognized belongs.
[0321] Specifically, the recognition result of the clustering language model for the training corpus includes the N-gram language model probability of each word in the recognition result, and this probability is the language score of the word.
[0322] After completing the path expansion of the third speech recognition result, select the clustering language model corresponding to the third speech recognition result to re-score the third speech recognition result and the extended path of the third speech recognition result, and determine the language model scores of the third speech recognition result and the extended path of the third speech recognition result respectively.
[0323] In this embodiment of the present application, since the general speech recognition model and the scenario-customized model are trained based on different corpora, they correspond to different clustering class language models. However, since the speech recognition decoding network and the scenario-customized model are both domain-specific models, they share the same clustering class language model. Therefore, when re-scoring the speech recognition results, different clustering language models are adapted based on the model that derived the speech recognition results for re-scoring.
[0324] In particular, when the above-mentioned language model excitation scheme is used to perform language model excitation on the candidate speech recognition results selected from the first speech recognition result and the second speech recognition result, since the candidate speech recognition may be the result output by the speech recognition decoding network (i.e., any one or more of the first speech recognition results), it may also be the result output by the general speech recognition model (i.e., one or more of the second speech recognition results), therefore, according to the source of the candidate speech recognition result, a clustering language model of the same type as its source should be selected for re-scoring.
[0325] The model structures of the above-mentioned different types of clustering language models are the same, but their training corpora are different. For example, the clustering language model of the same type as the general speech recognition model is obtained based on massive corpus training, and it is assumed that it is named Model A; and the clustering language model of the same type as the scenario customization model is obtained based on scenario corpora training, and it is assumed that it is named Model B; since Model A and Model B are trained based on different types of training corpora, they are different types of clustering language models. The clustering language model of the same type as the speech recognition decoding network is also trained based on scenario corpora, and it is assumed that it is named Model C; since Model B and Model C are trained based on the same type of training corpora, they are the same type of clustering language models.
[0326] Table 1 shows the Figure 14 The calculation method for re-checking the scores of the two paths in .
[0327] Table 1
[0328]
[0329] According to the above introduction, combined with Table 1, the language model scores scoreA and scoreB of the third speech recognition result and the extension path of the third speech recognition result can be determined respectively.
[0330] Finally, according to the language model score of the third speech recognition result, the language score of the third speech recognition result after language model excitation is determined using the language model score of the extended path of the third speech recognition result.
[0331] Specifically, the language model score scoreA of the third speech recognition result and the language model score scoreB of the extended path of the third speech recognition result are fused in a certain proportion, and the sum of their fusion coefficients is 1 to obtain the language score of the third speech recognition result after language model excitation.
[0332] For example, the language score Score of the third speech recognition result after language model excitation can be calculated using the following formula:
[0333] Score=γ*scoreA+(1–γ)*scoreB
[0334] Here, γ is an empirical coefficient, and its value is determined through testing, specifically with the goal of obtaining a correct language score and then selecting a correct speech recognition result from a large number of speech recognition results based on the language score PK.
[0335] The above embodiments introduce the specific implementation schemes of acoustic score excitation and language model excitation. In the above implementation schemes, especially in the acoustic score excitation implementation scheme, the problem of false triggering of vertical keywords is fully considered. By reasonably setting the excitation coefficient, it is possible to solve the false triggering caused by the excitation when the pronunciation of the vertical keyword and the real result is similar, as well as the false triggering problem introduced when the pronunciation difference between the vertical keyword and the real result is large.
[0336] As mentioned in the above embodiment, the problem of false triggering caused by the same pronunciation of the vertical category keyword and the real result cannot be solved by the above-mentioned solution of controlling the excitation coefficient. The reason is that the speech recognition decoding network constructed in the embodiment of the present application is a sentence network that relies on the acoustic model, which does not contain language information, so it cannot fundamentally solve the situation where the vertical category keyword and the real result have the same pronunciation. In order to reduce the impact of such false triggering, the embodiment of the present application uses multiple candidates to display the results to the user.
[0337] Specifically, since the general speech recognition model and the speech recognition decoding network share an acoustic model, when their output results have the same pronunciation, their acoustic scores must be the same. Therefore, when the acoustic scores of the first speech recognition result output by the speech recognition decoding network and the second speech recognition result output by the general speech recognition model are the same, the first speech recognition result and the second speech recognition result are used together as the final speech recognition result, that is, the first speech recognition result and the second speech recognition result are output at the same time, and the user selects the correct speech recognition result from them. Among them, the output order of the first speech recognition result and the second speech recognition result when they are output at the same time can be flexibly adjusted, and it is preferred to output them in the order of the first speech recognition result first and the second speech recognition result later.
[0338] It should be noted that the concept of outputting speech recognition results in a multi-candidate format is also applicable to the scoring of speech recognition results from multiple models. For example, in the above embodiment, when the output results of the speech recognition decoding network, the general speech recognition model, and the scenario-customized model are scored and compared to determine the final speech recognition result, if multiple different speech recognition results have the same score, these can be output simultaneously, allowing the user to select the correct speech recognition result.
[0339] So far, the above embodiments of the present application have respectively introduced the processing of each proposed speech recognition method, especially the typical processing steps in various speech recognition methods in detail. It should be noted that in order to keep the specification concise, the specific implementation methods of the same or corresponding processing steps in various speech recognition methods can refer to each other, and the embodiments of the present application will no longer list and describe them one by one. The processing steps in each speech recognition method can refer to each other and be combined to form a technical solution that does not exceed the scope of protection of this application.
[0340] In addition, corresponding to the above-mentioned speech recognition method, the embodiment of the present application also proposes a speech recognition device, see Figure 15 As shown, the speech recognition device includes:
[0341] Acoustic recognition unit 001, used to obtain the acoustic state sequence of the speech to be recognized;
[0342] A network construction unit 002 is configured to construct a speech recognition decoding network based on a set of vertically categorized keywords and a sentence decoding network for the scene to which the speech to be recognized belongs, wherein the sentence decoding network is constructed by at least performing sentence induction processing on a text corpus for the scene to which the speech to be recognized belongs;
[0343] The decoding processing unit 003 is used to decode the acoustic state sequence using the speech recognition decoding network to obtain a speech recognition result.
[0344] As an optional implementation, the above-mentioned construction of a speech recognition decoding network based on the vertical keyword set and sentence decoding network in the business scenario to which the speech to be recognized belongs includes:
[0345] The vertical keyword set under the scene to which the speech to be recognized belongs is transmitted to the cloud server, so that the cloud server builds a speech recognition decoding network based on the vertical keyword set under the scene to which the speech to be recognized belongs and the sentence decoding network.
[0346] As an optional implementation, the speech recognition result is used as the first speech recognition result;
[0347] The decoding processing unit 003 is further configured to:
[0348] Decoding the acoustic state sequence using a general speech recognition model to obtain a second speech recognition result;
[0349] A final speech recognition result is determined at least from the first speech recognition result and the second speech recognition result.
[0350] As an optional implementation manner, the decoding processing unit 003 is further configured to:
[0351] Decoding the acoustic state sequence using a pre-trained scenario-customized model to obtain a third speech recognition result; wherein the scenario-customized model is obtained by performing speech recognition training on speech in the scenario to which the speech to be recognized belongs;
[0352] Determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result includes:
[0353] A final speech recognition result is determined from the first speech recognition result, the second speech recognition result, and the third speech recognition result.
[0354] As an optional implementation manner, determining a final speech recognition result from the first speech recognition result, the second speech recognition result, and the third speech recognition result includes:
[0355] Performing language model excitation on the first speech recognition result, the second speech recognition result, and the third speech recognition result respectively;
[0356] According to the language scores of the first speech recognition result, the second speech recognition result and the third speech recognition result after stimulation, a final speech recognition result is determined from the first speech recognition result, the second speech recognition result and the third speech recognition result.
[0357] As an optional implementation manner, determining a final speech recognition result from the first speech recognition result, the second speech recognition result, and the third speech recognition result includes:
[0358] Performing acoustic score excitation on the first speech recognition result, and performing language model excitation on the third speech recognition result;
[0359] Determining a candidate speech recognition result from the first speech recognition result and the second speech recognition result according to the acoustic score of the first speech recognition result after acoustic score excitation and the acoustic score of the second speech recognition result;
[0360] Performing language model stimulation on the candidate speech recognition results;
[0361] A final speech recognition result is determined from the candidate speech recognition results and the third speech recognition result based on the language scores of the candidate speech recognition results after language model excitation and the language scores of the third speech recognition results after language model excitation.
[0362] Another embodiment of the present application also provides another speech recognition device, see Figure 16 As shown, the device includes:
[0363] Acoustic recognition unit 011, used to obtain the acoustic state sequence of the speech to be recognized;
[0364] A multidimensional decoding unit 012 is configured to decode the acoustic state sequence using a speech recognition decoding network to obtain a first speech recognition result, and to decode the acoustic state sequence using a universal speech recognition model to obtain a second speech recognition result; the speech recognition decoding network is constructed based on a set of vertical keywords and a sentence decoding network for the scene to which the speech to be recognized belongs;
[0365] An acoustic excitation unit 013, configured to perform acoustic score excitation on the first speech recognition result;
[0366] The decision processing unit 014 is configured to determine a final speech recognition result from at least the first speech recognition result and the second speech recognition result after stimulation.
[0367] As an optional implementation manner, the multi-dimensional decoding unit 012 is further configured to:
[0368] Decoding the acoustic state sequence using a pre-trained scenario-customized model to obtain a third speech recognition result; wherein the scenario-customized model is obtained by performing speech recognition training on speech in the scenario to which the speech to be recognized belongs;
[0369] The step of determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result after stimulation includes:
[0370] A final speech recognition result is determined from the first speech recognition result after stimulation, the second speech recognition result, and the third speech recognition result.
[0371] As an optional implementation manner, determining the final speech recognition result from the stimulated first speech recognition result, the second speech recognition result, and the third speech recognition result includes:
[0372] Determining a candidate speech recognition result from the first speech recognition result and the second speech recognition result according to the acoustic score of the first speech recognition result after acoustic score excitation and the acoustic score of the second speech recognition result;
[0373] Performing language model excitation on the candidate speech recognition result and the third speech recognition result respectively;
[0374] A final speech recognition result is determined from the candidate speech recognition results and the third speech recognition result based on the language scores of the candidate speech recognition results after language model excitation and the language scores of the third speech recognition results after language model excitation.
[0375] As an optional implementation, the sentence decoding network for the scene to be recognized speech is constructed by the following processing:
[0376] By performing sentence induction and grammar slot definition processing on the corpus data in the scene to which the speech to be recognized belongs, a text sentence network is constructed; wherein the text sentence network includes ordinary grammar slots corresponding to non-vertical category keywords and replacement grammar slots corresponding to vertical category keywords, and the replacement grammar slots store placeholders corresponding to vertical category keywords;
[0377] Segmenting the entries in the common grammar slots of the text sentence network and expanding the word nodes according to the segmentation results to obtain a word-level sentence decoding network;
[0378] Each word in the ordinary grammar slot of the word-level sentence decoding network is replaced with the corresponding pronunciation, and the pronunciation node is expanded according to the pronunciation corresponding to the word to obtain a pronunciation-level sentence decoding network. The pronunciation-level sentence decoding network serves as the sentence decoding network for the scenario to which the speech to be recognized belongs.
[0379] As an optional implementation, the entries in the common grammar slots of the text sentence network are segmented and word nodes are expanded according to the segmentation results to obtain a word-level sentence decoding network, including:
[0380] Segment each entry in the common grammar slot in the text sentence network to obtain the words corresponding to each entry;
[0381] Expand word nodes using the words corresponding to the same entry to obtain a word string corresponding to the entry;
[0382] The word strings corresponding to each lexical entry of the same common grammar slot are connected in parallel to obtain a word-level sentence decoding network.
[0383] As an optional implementation, each word in the common grammar slot of the word-level sentence decoding network is replaced with the corresponding pronunciation, and the pronunciation node is expanded according to the pronunciation corresponding to the word to obtain a pronunciation-level sentence decoding network, including:
[0384] Replacing each word in the common grammar slot of the word-level sentence decoding network with the corresponding pronunciation;
[0385] Each pronunciation in the word-level sentence decoding network is divided into pronunciation units, and pronunciation nodes are expanded using the pronunciation units corresponding to the pronunciations to obtain a pronunciation-level sentence decoding network.
[0386] As an optional implementation, a speech recognition decoding network is constructed based on a set of vertical keywords and a sentence decoding network in the scene to which the speech to be recognized belongs, including:
[0387] Obtaining a pre-built sentence decoding network for the scenario to which the speech to be recognized belongs;
[0388] Building a vertical keyword network based on the vertical keyword set in the scene to which the speech to be recognized belongs;
[0389] Insert the vertical keyword network into the sentence decoding network to obtain a speech recognition decoding network.
[0390] As an optional implementation manner, the vertical keyword network is constructed based on the vertical keyword in the vertical keyword set in the scene to which the speech to be recognized belongs, including:
[0391] Based on each vertical keyword in the vertical keyword set under the scene to which the speech to be recognized belongs, a word-level vertical keyword network is constructed;
[0392] Each word in the word-level vertical keyword network is replaced with the corresponding pronunciation, and the pronunciation nodes are expanded according to the pronunciation corresponding to the word to obtain the pronunciation-level vertical keyword network.
[0393] As an optional implementation, the vertical keyword network and the sentence decoding network are both composed of nodes and directed arcs connecting the nodes, and pronunciation information or placeholders are stored on the directed arcs between the nodes;
[0394] Inserting the vertical keyword network into the sentence decoding network to obtain a speech recognition decoding network includes:
[0395] The vertical keyword network is connected to the left and right nodes of the replacement grammar slot of the sentence decoding network through directed arcs to construct a speech recognition decoding network.
[0396] As an optional implementation, the vertical keyword network is connected to the left and right nodes of the replacement grammar slot of the sentence decoding network through directed arcs to construct a speech recognition decoding network, including:
[0397] A speech recognition decoding network is constructed by connecting the right node of each outgoing arc of the starting node of the vertical keyword network with the left node of the replacement grammar slot through a directed arc, and connecting the left node of each incoming arc of the ending node of the vertical keyword network with the right node of the replacement grammar slot through a directed arc.
[0398] As an optional implementation manner, the first arc and the last arc of each keyword in the vertical keyword network respectively store a unique identifier corresponding to the keyword;
[0399] The method comprises: connecting the right node of each outgoing arc of the starting node of the vertical keyword network with the left node of the replacement grammar slot through a directed arc, and connecting the left node of each incoming arc of the ending node of the vertical keyword network with the right node of the replacement grammar slot through a directed arc to construct a speech recognition decoding network, including:
[0400] Traversing each outgoing arc of the starting node of the vertical keyword network, for each traversed outgoing arc, determining whether the keyword corresponding to the unique identifier has been inserted into the sentence decoding network based on the unique identifier on the outgoing arc and the set of networked keyword information; wherein the set of networked keyword information correspondingly stores the unique identifier of the keyword that has been inserted into the sentence decoding network, and the left and right node numbers of the directed arc where the unique identifier is located in the sentence decoding network;
[0401] If the keyword corresponding to the unique identifier is not inserted into the sentence decoding network, the right node of the traversed outgoing arc is connected to the left node of the replacement grammar slot through a directed arc, and the pronunciation information of the traversed outgoing arc is stored on the directed arc;
[0402] as well as,
[0403] Traversing each incoming arc of the end node of the vertical keyword network, for each traversed incoming arc, determining whether the keyword corresponding to the unique identifier has been inserted into the sentence decoding network based on the unique identifier on the incoming arc and the set of networked keyword information;
[0404] If the keyword corresponding to the unique identifier is not inserted into the sentence decoding network, the left node of the traversed input arc is connected to the right node of the replacement grammar slot through a directed arc, and the pronunciation information of the traversed input arc is stored on the directed arc.
[0405] As an optional embodiment, the speech recognition decoding network is constructed by connecting the right node of each outgoing arc of the starting node of the vertical keyword network to the left node of the replacement grammar slot through a directed arc, and connecting the left node of each incoming arc of the ending node of the vertical keyword network to the right node of the replacement grammar slot through a directed arc, further comprising:
[0406] When the keyword in the vertical keyword network is inserted into the sentence decoding network, the unique identifier of the keyword and the left and right node numbers of the directed arc where the unique identifier is located in the sentence decoding network are stored in the networked keyword information set accordingly.
[0407] As an optional embodiment, the speech recognition decoding network is constructed by connecting the right node of each outgoing arc of the starting node of the vertical keyword network to the left node of the replacement grammar slot through a directed arc, and connecting the left node of each incoming arc of the ending node of the vertical keyword network to the right node of the replacement grammar slot through a directed arc, further comprising:
[0408] Traversing each unique identifier in the set of keyword information that has entered the network;
[0409] If the traversed unique identifier is not the unique identifier of any keyword in the vertical category keyword set under the scene to which the voice to be recognized belongs, the directed arc between the left and right node numbers corresponding to the unique identifier is disconnected.
[0410] As an optional implementation, the above-mentioned speech recognition device further includes:
[0411] A result correction unit is used to correct the first speech recognition result according to the second speech recognition result.
[0412] As an optional implementation manner, correcting the first speech recognition result according to the second speech recognition result includes:
[0413] Using the reference text content in the second speech recognition result, correct the non-vertical keyword content in the first speech recognition result to obtain a corrected first speech recognition result;
[0414] Among them, the reference text content is the text content in the second voice recognition result that matches the non-vertical keyword content in the first voice recognition result.
[0415] As an optional implementation, using the reference text content in the second speech recognition result to correct the non-vertical keyword content in the first speech recognition result to obtain a corrected first speech recognition result, including:
[0416] Determining vertical category keyword content and non-vertical category keyword content from the first speech recognition result, and determining text content corresponding to the non-vertical category keyword content in the first speech recognition result from the second speech recognition result as reference text content;
[0417] Determining corrected non-vertical keyword content based on the reference text content in the second speech recognition result and the non-vertical keyword content in the first speech recognition result;
[0418] The corrected non-vertical keyword content and the vertical keyword content are combined to obtain a corrected first speech recognition result.
[0419] As an optional implementation, determining text content corresponding to the non-vertical keyword content in the first speech recognition result from the second speech recognition result as reference text content includes:
[0420] Determine an edit distance matrix between the first speech recognition result and the second speech recognition result according to an edit distance algorithm;
[0421] Based on the edit distance matrix and the non-vertical keyword content in the first speech recognition result, the text content corresponding to the non-vertical keyword content in the first speech recognition result is determined from the second speech recognition result as the reference text content.
[0422] As an optional implementation, determining the corrected non-vertical keyword content based on the reference text content in the second speech recognition result and the non-vertical keyword content in the first speech recognition result includes:
[0423] Determine the target text content in the second voice recognition result or the non-vertical keyword content in the first voice recognition result as the corrected non-vertical keyword content based on the character difference between the reference text content in the second voice recognition result and the non-vertical keyword content in the first voice recognition result;
[0424] Among them, the target text content in the second voice recognition result refers to the text content in the second voice recognition result that corresponds to the position of the non-vertical keyword content in the first voice recognition result.
[0425] As an optional implementation, based on the character difference between the reference text content in the second voice recognition result and the non-vertical keyword content in the first voice recognition result, determining the target text content in the second voice recognition result or the non-vertical keyword content in the first voice recognition result as the corrected non-vertical keyword content includes:
[0426] Comparing the reference text content in the second speech recognition result with the non-vertical keyword content in the first speech recognition result to determine whether the reference text content in the second speech recognition result is the same as the non-vertical keyword content in the first speech recognition result;
[0427] If they are the same, the target text content in the second speech recognition result is determined as the corrected non-vertical keyword content;
[0428] If they are different, determining whether the second speech recognition result has more characters than the non-vertical keyword content in the first speech recognition result, and whether the difference in the number of characters between the two does not exceed a set threshold;
[0429] If the second speech recognition result has more characters than the non-vertical keyword content in the first speech recognition result, and the difference in the number of characters between the two does not exceed a set threshold, the target text content in the second speech recognition result is determined to be the corrected non-vertical keyword content;
[0430] If the number of characters in the second voice recognition result is less than that in the non-vertical keyword content in the first voice recognition result, and / or the difference in the number of characters between the two exceeds a set threshold, the non-vertical keyword content in the first voice recognition result is determined as the corrected non-vertical keyword content.
[0431] As an optional implementation manner, determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result includes:
[0432] Determining a degree of match between the first speech recognition result and the second speech recognition result by comparing the first speech recognition result and the second speech recognition result, and determining a confidence level of the first speech recognition result based on the degree of match between the first speech recognition result and the second speech recognition result;
[0433] When the confidence level of the first speech recognition result is greater than a preset confidence threshold, selecting a final speech recognition result from the first speech recognition result and the second speech recognition result based on the acoustic score of the first speech recognition result and the acoustic score of the second speech recognition result;
[0434] When the confidence of the first speech recognition result is not greater than a preset confidence threshold, the first speech recognition result is acoustically scored, and a final speech recognition result is selected from the first speech recognition result and the second speech recognition result based on the acoustic score of the first speech recognition result after the stimulation and the acoustic score of the second speech recognition result.
[0435] As an optional implementation manner, determining the confidence level of the first speech recognition result based on the matching degree between the first speech recognition result and the second speech recognition result includes:
[0436] Determining whether a degree of matching between the first speech recognition result and the second speech recognition result is greater than a set matching degree threshold;
[0437] If the match score is greater than a set matching threshold, the confidence level of the first speech recognition result is calculated based on the acoustic scores of each frame of the first speech recognition result;
[0438] If it is not greater than the set matching degree threshold, a decoding network is constructed using the vertical category keyword content in the first speech recognition result and the second speech recognition result, the decoding network is used to re-decode the acoustic state sequence, and the decoding result is used to update the first speech recognition result;
[0439] The confidence level of the first speech recognition result is calculated and determined based on the acoustic scores of the updated frames of the first speech recognition result.
[0440] As an optional implementation, when the acoustic scores of the first speech recognition result and the second speech recognition result are the same, the first speech recognition result and the second speech recognition result are taken together as the final speech recognition result.
[0441] As an optional implementation manner, performing acoustic score incentive on the first speech recognition result includes:
[0442] determining an acoustic excitation coefficient based on at least the vertical category keyword content and the non-vertical category keyword content in the first speech recognition result;
[0443] Using the acoustic excitation coefficient, updating the acoustic score of the vertical category keyword content in the first speech recognition result;
[0444] The acoustic score of the first speech recognition result is recalculated and determined based on the updated acoustic scores of the vertical category keyword content in the first speech recognition result and the acoustic scores of the non-vertical category keyword content in the first speech recognition result.
[0445] As an optional implementation manner, determining the acoustic excitation coefficient based on at least the vertical category keyword content and the non-vertical category keyword content in the first speech recognition result includes:
[0446] The acoustic excitation coefficient is calculated and determined based on the acoustic score excitation prior coefficient in the scenario to which the speech to be recognized belongs, the number of characters and phonemes of the vertical category keywords in the first speech recognition result, and the total number of characters and phonemes in the first speech recognition result.
[0447] As an optional implementation manner, determining the acoustic excitation coefficient based on at least the vertical category keyword content and the non-vertical category keyword content in the first speech recognition result includes:
[0448] Calculate the score confidence of the vertical category keyword content in the first speech recognition result based on the number of phonemes and acoustic scores of the vertical category keyword content in the first speech recognition result, and the number of phonemes and acoustic scores of the non-vertical category keyword content in the first speech data result;
[0449] The acoustic excitation coefficient is determined based on at least the confidence score of the vertical keyword content in the first speech recognition result.
[0450] As an optional implementation manner, determining the acoustic excitation coefficient based on at least the confidence score of the vertical category keyword content in the first speech recognition result includes:
[0451] The acoustic excitation coefficient is determined based on the score confidence of the vertical keyword content in the first speech recognition result, and the relationship between the predetermined acoustic excitation coefficient and the recognition effect and recognition false triggering.
[0452] As an optional implementation, performing language model excitation on the third speech recognition result includes:
[0453] Performing path expansion on the third speech recognition result based on a set of vertical category keywords under the scene to be recognized and a category label corresponding to the scene; the category label is determined by clustering the speech recognition scenes;
[0454] Determining the language model scores of the third speech recognition result and the extension path of the third speech recognition result respectively based on the recognition results of the training corpus by the clustering language model corresponding to the category label; wherein the clustering language model is obtained by performing speech recognition training on the target corpus, and the vertical category keywords in the target corpus are all replaced by the category label;
[0455] The language score of the third speech recognition result after language model excitation is determined according to the language model score of the third speech recognition result and the language model score of the extended path of the third speech recognition result.
[0456] As an optional implementation, the path expansion of the third speech recognition result according to the vertical category keyword set under the scene to which the speech to be recognized belongs and the category label corresponding to the scene includes:
[0457] Comparing the vertical category keywords in the third speech recognition result with the vertical category keywords in the vertical category keyword set for the scene to which the speech to be recognized belongs;
[0458] If the vertical category keyword in the third speech recognition result matches any vertical category keyword in the vertical category keyword set, a new path is extended between the left and right nodes of the slot where the vertical category keyword in the third speech recognition result is located, and the category label corresponding to the scene to which the speech to be recognized belongs is stored on the new path.
[0459] Specifically, for the specific working contents of each unit in the embodiments of the above-mentioned speech recognition devices, please refer to the processing contents of the corresponding steps of the above-mentioned speech recognition method, which will not be repeated here.
[0460] Another embodiment of the present application also provides a voice recognition device, see Figure 17 As shown, the device includes:
[0461] Memory 200 and processor 210;
[0462] The memory 200 is connected to the processor 210 and is used to store programs;
[0463] The processor 210 is configured to implement the speech recognition method disclosed in any of the above embodiments by running the program stored in the memory 200 .
[0464] Specifically, the above-mentioned speech recognition device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0465] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are interconnected via a bus.
[0466] A bus may include a pathway that transfers information between components of a computer system.
[0467] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, or the like, or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. Alternatively, it can be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware components.
[0468] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0469] The memory 200 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, which includes computer operating instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash memory, etc.
[0470] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0471] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speakers, etc.
[0472] The communication interface 220 may include any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0473] The processor 2102 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of the speech recognition method provided in the embodiment of the present application.
[0474] Another embodiment of the present application further provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the various steps of the speech recognition method provided in any of the above embodiments.
[0475] Specifically, the specific working contents of each part of the above-mentioned speech recognition device, as well as the specific processing contents of the computer program on the above-mentioned storage medium when being executed by the processor, can be found in the contents of the various embodiments of the above-mentioned speech recognition method, and will not be repeated here.
[0476] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0477] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and relevant parts can be referred to the partial description of the method embodiments.
[0478] The steps in the methods of each embodiment of the present application can be adjusted in sequence, merged, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0479] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be merged, divided, and deleted according to actual needs.
[0480] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or submodules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0481] The modules or submodules described as separate components may or may not be physically separate, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules may be selected to achieve the purpose of this embodiment according to actual needs.
[0482] In addition, each functional module or submodule in each embodiment of the present application may be integrated into a processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into a single module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or software functional modules or submodules.
[0483] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0484] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, software units executed by a processor, or a combination of the two. The software units may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0485] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise," "include," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a set of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0486] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.< / sth> < / phone> < / sth> < / phone> < / phone> < / sth> < / phone> < / phone> < / sth> < / sth> < / sth> < / phone> < / phone> < / sth> < / phone> < / xxxx>
Claims
1. A speech recognition method, characterized in that: include: Obtaining the acoustic state sequence of the speech to be recognized; Based on the vertical category keyword set and sentence decoding network of the scene to be recognized, a speech recognition decoding network is constructed, wherein the sentence decoding network is constructed by at least performing sentence induction processing on the text corpus of the scene to be recognized; Decoding the acoustic state sequence using the speech recognition decoding network to obtain a speech recognition result; It also includes: performing language model excitation on the speech recognition result so that the language model information is integrated into the speech recognition result; the language model excitation includes: matching the speech recognition result with the vertical category keyword under the scene to which the speech to be recognized belongs, and if the match is successful, performing path expansion on the speech recognition result; scoring the expanded speech recognition result based on the clustering language model to complete the excitation of the speech recognition result on the language model.
2. The method according to claim 1, characterized in that Based on the vertical keyword set and sentence decoding network in the scene to which the speech to be recognized belongs, a speech recognition decoding network is constructed, including: The vertical keyword set under the scene to which the speech to be recognized belongs is transmitted to the cloud server, so that the cloud server builds a speech recognition decoding network based on the vertical keyword set under the scene to which the speech to be recognized belongs and the sentence decoding network.
3. The method according to claim 1, characterized in that The speech recognition result is used as the first speech recognition result; The method further comprises: Decoding the acoustic state sequence using a general speech recognition model to obtain a second speech recognition result; A final speech recognition result is determined at least from the first speech recognition result and the second speech recognition result.
4. The method according to claim 3, characterized in that The method further comprises: Decoding the acoustic state sequence using a pre-trained scenario-customized model to obtain a third speech recognition result; wherein the scenario-customized model is obtained by performing speech recognition training on speech in the scenario to which the speech to be recognized belongs; Determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result includes: A final speech recognition result is determined from the first speech recognition result, the second speech recognition result, and the third speech recognition result.
5. The method according to claim 4, characterized in that Determining a final speech recognition result from the first speech recognition result, the second speech recognition result, and the third speech recognition result includes: Performing language model excitation on the first speech recognition result, the second speech recognition result, and the third speech recognition result respectively; According to the language scores of the first speech recognition result, the second speech recognition result and the third speech recognition result after stimulation, a final speech recognition result is determined from the first speech recognition result, the second speech recognition result and the third speech recognition result.
6. The method according to claim 4, characterized in that Determining a final speech recognition result from the first speech recognition result, the second speech recognition result, and the third speech recognition result includes: Performing acoustic score excitation on the first speech recognition result, and performing language model excitation on the third speech recognition result; Determining a candidate speech recognition result from the first speech recognition result and the second speech recognition result according to the acoustic score of the first speech recognition result after acoustic score excitation and the acoustic score of the second speech recognition result; Performing language model stimulation on the candidate speech recognition results; A final speech recognition result is determined from the candidate speech recognition results and the third speech recognition result based on the language scores of the candidate speech recognition results after language model excitation and the language scores of the third speech recognition results after language model excitation.
7. A speech recognition method, characterized in that: include: Obtaining the acoustic state sequence of the speech to be recognized; The acoustic state sequence is decoded using a speech recognition decoding network to obtain a first speech recognition result, and the acoustic state sequence is decoded using a general speech recognition model to obtain a second speech recognition result; the speech recognition decoding network is constructed based on a vertical category keyword set and a sentence decoding network in the scene to which the speech to be recognized belongs; Performing acoustic score incentive on the first speech recognition result; Determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result after stimulation; It also includes: performing language model excitation on the speech recognition result so that the language model information is integrated into the speech recognition result; the language model excitation includes: matching the speech recognition result with the vertical category keyword under the scene to which the speech to be recognized belongs, and if the match is successful, performing path expansion on the speech recognition result; scoring the expanded speech recognition result based on the clustering language model to complete the excitation of the speech recognition result on the language model.
8. The method according to claim 7, characterized in that The method further comprises: Decoding the acoustic state sequence using a pre-trained scenario-customized model to obtain a third speech recognition result; wherein the scenario-customized model is obtained by performing speech recognition training on speech in the scenario to which the speech to be recognized belongs; The step of determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result after stimulation includes: A final speech recognition result is determined from the first speech recognition result after stimulation, the second speech recognition result, and the third speech recognition result.
9. The method according to claim 8, characterized in that Determining a final speech recognition result from the stimulated first speech recognition result, the second speech recognition result, and the third speech recognition result includes: Determining a candidate speech recognition result from the first speech recognition result and the second speech recognition result according to the acoustic score of the first speech recognition result after acoustic score excitation and the acoustic score of the second speech recognition result; Performing language model excitation on the candidate speech recognition result and the third speech recognition result respectively; A final speech recognition result is determined from the candidate speech recognition results and the third speech recognition result based on the language scores of the candidate speech recognition results after language model excitation and the language scores of the third speech recognition results after language model excitation.
10. The method according to any one of claims 1 to 9, characterized in that The sentence decoding network for the scene to be recognized speech is constructed by the following process: By performing sentence induction and grammar slot definition processing on the corpus data in the scene to which the speech to be recognized belongs, a text sentence network is constructed; wherein the text sentence network includes ordinary grammar slots corresponding to non-vertical category keywords and replacement grammar slots corresponding to vertical category keywords, and the replacement grammar slots store placeholders corresponding to vertical category keywords; Segmenting the entries in the common grammar slots of the text sentence network and expanding the word nodes according to the segmentation results to obtain a word-level sentence decoding network; Each word in the ordinary grammar slot of the word-level sentence decoding network is replaced with the corresponding pronunciation, and the pronunciation node is expanded according to the pronunciation corresponding to the word to obtain a pronunciation-level sentence decoding network. The pronunciation-level sentence decoding network serves as the sentence decoding network for the scenario to which the speech to be recognized belongs.
11. The method according to any one of claims 1 to 9, characterized in that Based on the vertical keyword set and sentence decoding network in the scene to which the speech to be recognized belongs, a speech recognition decoding network is constructed, including: Obtaining a pre-built sentence decoding network for the scenario to which the speech to be recognized belongs; Building a vertical keyword network based on the vertical keyword set in the scene to which the speech to be recognized belongs; Insert the vertical keyword network into the sentence decoding network to obtain a speech recognition decoding network.
12. The method according to claim 11, characterized in that The step of constructing a vertical keyword network based on vertical keywords in a vertical keyword set in the scene to which the speech to be recognized belongs includes: Based on each vertical keyword in the vertical keyword set under the scene to which the speech to be recognized belongs, a word-level vertical keyword network is constructed; Each word in the word-level vertical keyword network is replaced with the corresponding pronunciation, and the pronunciation nodes are expanded according to the pronunciation corresponding to the word to obtain the pronunciation-level vertical keyword network.
13. The method according to claim 11, characterized in that The vertical keyword network and the sentence decoding network are both composed of nodes and directed arcs connecting the nodes, and pronunciation information or placeholders are stored on the directed arcs between the nodes; Inserting the vertical keyword network into the sentence decoding network to obtain a speech recognition decoding network includes: The vertical keyword network is connected to the left and right nodes of the replacement grammar slot of the sentence decoding network through directed arcs to construct a speech recognition decoding network.
14. The method according to claim 13, characterized in that The first arc and the last arc of each keyword in the vertical keyword network respectively store a unique identifier corresponding to the keyword; When the keyword in the vertical keyword network is inserted into the sentence decoding network, the unique identifier of the keyword and the left and right node numbers of the directed arc where the unique identifier is located in the sentence decoding network are stored correspondingly in the keyword information set that has been connected to the network; wherein, the keyword information set that has been connected to the network correspondingly stores the unique identifier of the keyword that has been inserted into the sentence decoding network and the left and right node numbers of the directed arc where the unique identifier is located in the sentence decoding network.
15. The method according to claim 14, characterized in that Also includes: Traversing each unique identifier in the set of keyword information that has entered the network; If the traversed unique identifier is not the unique identifier of any keyword in the vertical category keyword set under the scene to which the voice to be recognized belongs, the directed arc between the left and right node numbers corresponding to the unique identifier is disconnected.
16. The method according to any one of claims 3 to 9, characterized in that The method further comprises: Using the reference text content in the second speech recognition result, correct the non-vertical keyword content in the first speech recognition result to obtain a corrected first speech recognition result; Among them, the reference text content is the text content in the second voice recognition result that matches the non-vertical keyword content in the first voice recognition result.
17. The method according to claim 16, characterized in that Using the reference text content in the second speech recognition result, correcting the non-vertical keyword content in the first speech recognition result to obtain a corrected first speech recognition result, including: Determining vertical category keyword content and non-vertical category keyword content from the first speech recognition result, and determining text content corresponding to the non-vertical category keyword content in the first speech recognition result from the second speech recognition result as reference text content; Determining corrected non-vertical keyword content based on the reference text content in the second speech recognition result and the non-vertical keyword content in the first speech recognition result; The corrected non-vertical keyword content and the vertical keyword content are combined to obtain a corrected first speech recognition result.
18. The method according to claim 17, characterized in that Determining, from the second speech recognition result, text content corresponding to the non-vertical keyword content in the first speech recognition result as reference text content, including: Determine an edit distance matrix between the first speech recognition result and the second speech recognition result according to an edit distance algorithm; Based on the edit distance matrix and the non-vertical keyword content in the first speech recognition result, the text content corresponding to the non-vertical keyword content in the first speech recognition result is determined from the second speech recognition result as the reference text content.
19. The method according to claim 17, wherein The determining, based on the reference text content in the second speech recognition result and the non-vertical keyword content in the first speech recognition result, corrected non-vertical keyword content includes: Determining whether the reference text content in the second speech recognition result is the same as the non-vertical keyword content in the first speech recognition result; If they are the same, the target text content in the second speech recognition result is determined as the corrected non-vertical keyword content; If they are different, determining whether the second speech recognition result has more characters than the non-vertical keyword content in the first speech recognition result, and whether the difference in the number of characters between the two does not exceed a set threshold; If the second speech recognition result has more characters than the non-vertical keyword content in the first speech recognition result, and the difference in the number of characters between the two does not exceed a set threshold, the target text content in the second speech recognition result is determined to be the corrected non-vertical keyword content; If the number of characters in the second voice recognition result is less than that in the non-vertical keyword content in the first voice recognition result, and / or the difference in the number of characters between the two exceeds a set threshold, the non-vertical keyword content in the first voice recognition result is determined as the corrected non-vertical keyword content; Among them, the target text content in the second voice recognition result refers to the text content in the second voice recognition result that corresponds to the position of the non-vertical keyword content in the first voice recognition result.
20. The method according to claim 3, wherein Determining a final speech recognition result from at least the first speech recognition result and the second speech recognition result includes: Determining whether the confidence level of the first speech recognition result is greater than a preset confidence level threshold; When the confidence level of the first speech recognition result is greater than a preset confidence threshold, selecting a final speech recognition result from the first speech recognition result and the second speech recognition result based on the acoustic score of the first speech recognition result and the acoustic score of the second speech recognition result; When the confidence of the first speech recognition result is not greater than a preset confidence threshold, the first speech recognition result is acoustically scored, and a final speech recognition result is selected from the first speech recognition result and the second speech recognition result based on the acoustic score of the first speech recognition result after the stimulation and the acoustic score of the second speech recognition result.
21. The method according to claim 7 or 20, characterized in that When the acoustic scores of the first speech recognition result and the second speech recognition result are the same, the first speech recognition result and the second speech recognition result are taken together as the final speech recognition result.
22. The method according to claim 6, 4 or 20, characterized in that Performing acoustic score incentive on the first speech recognition result, comprising: determining an acoustic excitation coefficient based on at least the vertical category keyword content and the non-vertical category keyword content in the first speech recognition result; Using the acoustic excitation coefficient, updating the acoustic score of the vertical category keyword content in the first speech recognition result; The acoustic score of the first speech recognition result is recalculated and determined based on the updated acoustic scores of the vertical category keyword content in the first speech recognition result and the acoustic scores of the non-vertical category keyword content in the first speech recognition result.
23. The method according to claim 5, 6 or 9, characterized in that Performing language model excitation on the third speech recognition result, comprising: Performing path expansion on the third speech recognition result based on a set of vertical category keywords under the scene to be recognized and a category label corresponding to the scene; the category label is determined by clustering the speech recognition scenes; Determining the language model scores of the third speech recognition result and the extension path of the third speech recognition result respectively based on the recognition results of the training corpus by the clustering language model corresponding to the category label; wherein the clustering language model is obtained by performing speech recognition training on the target corpus, and the vertical category keywords in the target corpus are all replaced by the category label; The language score of the third speech recognition result after language model excitation is determined according to the language model score of the third speech recognition result and the language model score of the extended path of the third speech recognition result.
24. The method according to claim 23, wherein The performing path expansion on the third speech recognition result according to the vertical category keyword set under the scene to which the speech to be recognized belongs and the category label corresponding to the scene includes: Comparing the vertical category keywords in the third speech recognition result with the vertical category keywords in the vertical category keyword set for the scene to which the speech to be recognized belongs; If the vertical category keyword in the third speech recognition result matches any vertical category keyword in the vertical category keyword set, a new path is extended between the left and right nodes of the slot where the vertical category keyword in the third speech recognition result is located, and the category label corresponding to the scene to which the speech to be recognized belongs is stored on the new path.
25. A speech recognition device, characterized in that: include: An acoustic recognition unit, configured to obtain an acoustic state sequence of a speech to be recognized; A network construction unit, configured to construct a speech recognition decoding network based on a set of vertically categorized keywords and a sentence decoding network for the scene to which the speech to be recognized belongs, wherein the sentence decoding network is constructed at least by performing sentence induction processing on a text corpus for the scene to which the speech to be recognized belongs; A decoding processing unit, configured to decode the acoustic state sequence using the speech recognition decoding network to obtain a speech recognition result; The device is also used to: perform language model excitation on the speech recognition result so that language model information is integrated into the speech recognition result; the language model excitation includes: matching the speech recognition result with the vertical category keyword under the scene to which the speech to be recognized belongs, and if the match is successful, performing path expansion on the speech recognition result; scoring the expanded speech recognition result based on the clustered language model to complete the excitation of the speech recognition result on the language model.
26. A speech recognition device, characterized in that: include: An acoustic recognition unit, configured to obtain an acoustic state sequence of a speech to be recognized; A multidimensional decoding unit, configured to decode the acoustic state sequence using a speech recognition decoding network to obtain a first speech recognition result, and to decode the acoustic state sequence using a universal speech recognition model to obtain a second speech recognition result; the speech recognition decoding network is constructed based on a set of vertical keywords and a sentence decoding network for the scene to which the speech to be recognized belongs; an acoustic excitation unit, configured to perform acoustic score excitation on the first speech recognition result; a decision processing unit, configured to determine a final speech recognition result from at least the first speech recognition result and the second speech recognition result after stimulation; The device is further configured to: perform language model excitation on the speech recognition result so that the speech recognition result is integrated with language model information; the language model excitation includes: matching the speech recognition result with a vertical category keyword under the scene to which the speech to be recognized belongs, and if the match is successful, performing path expansion on the speech recognition result; The clustering-based language model scores the expanded speech recognition results to complete the motivation of the speech recognition results on the language model.
27. A speech recognition device, characterized in that include: memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the speech recognition method according to any one of claims 1 to 24 by running the program stored in the memory.
28. A storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the speech recognition method according to any one of claims 1 to 24 is implemented.
Citation Information
Patent Citations
Voice signal processing method and apparatus
CN105845133A
Decoding network generation method, apparatus and device and readable storage medium
CN109087645A
Method and device for acquiring text information, equipment and storage medium
CN113515945A