Speech Recognition Method, Related Devices and Readable Storage Medium
By building a decoding network including the main decoding network and the hot word decoding network, the problem of inefficient hot word recognition in the prior art is solved, and the efficiency of hot word recognition is improved.
Patent Information
- Application Number
- CN202111479624.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-06
AI Technical Summary
Existing speech recognition technology is inefficient in identifying hot words and requires two decoding, resulting in low efficiency in recognition of hot words.
Build a decoding network that includes the main decoding network and the hot word decoding network. Through the decoding process, use the hot word decoding network to stimulate the voice signal to improve the recognition efficiency.
Through a decoding process, effective incentives for hot words can be achieved, improving the recognition efficiency of hot words.
Smart Images

Figure CN114155836B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology. More specifically, it relates to a speech recognition method, related devices, and readable storage media. Background Art
[0002] With the development of artificial intelligence, speech recognition has penetrated into all aspects of people's lives. At present, general speech recognition has reached a very high level, but its recognition effect in special vocabulary, technical terms, proper nouns, etc. still needs to be further improved. These words are often high-frequency words used by users, that is, hot words. Hot words have user characteristics, and users have a very low tolerance for errors in recognizing hot words. Therefore, improving the recognition effect of hot words is highly anticipated by users. In order to improve the recognition effect of hot words, the hot words of users can be obtained, and during speech recognition, hot words can be used to assist in recognition.
[0003] Currently, when performing speech recognition, using hot words to assist in recognition specifically means using a decoding network to decode a speech signal to achieve the excitation of hot words and improve the recall rate of hot words in the recognition result, thereby improving the recognition effect of hot words. However, this method requires the decoding network to perform decoding twice, resulting in low efficiency in recognizing hot words.
[0004] Therefore, how to improve the recognition efficiency of hot words has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, this application proposes a speech recognition method, related devices, and readable storage media. The specific solutions are as follows:
[0006] A speech recognition method, the method includes:
[0007] Obtain a speech signal to be recognized;
[0008] Obtain a pre-constructed decoding network, where the decoding network includes a main decoding network and a hot word decoding network;
[0009] Use the decoding network to decode the speech signal. During the decoding process, use the hot word decoding network to perform hot word excitation on the speech signal to obtain a corresponding speech recognition text.
[0010] Optionally, the construction method of the decoding network includes:
[0011] Construct a main decoding network, where at least one slot is included in the main decoding network, and each slot is located between two nodes;
[0012] Obtain a first hot word list, where the first hot word list includes one or more first hot words;
[0013] Based on the first hot word list, a hot word decoding network is constructed, wherein the hot word decoding network includes a plurality of branches, the number of the branches is the same as the number of the first hot words in the first hot word list, and the arcs at the head and the tail of each branch are silent arcs;
[0014] For each slot in the main decoding network, the hot word decoding network is inserted into the slot to generate a decoding network.
[0015] Optionally, inserting the hot word decoding network into the slot includes:
[0016] For each branch in the hot word decoding network, remove the silent arcs at the head and tail of each branch in the hot word decoding network;
[0017] The first phoneme of the triphones corresponding to the real arc at the head of the branch is set as the last phoneme of the triphones corresponding to the first real arc before the slot in the main decoding network, and the last phoneme of the triphones corresponding to the real arc at the tail of the branch is set as the first phoneme of the triphones corresponding to the first real arc after the slot in the main decoding network;
[0018] The real arc at the head of the branch is connected to the first node before the slot in the main decoding network, and the real arc at the tail of the branch is connected to the first node after the slot in the main decoding network.
[0019] Optionally, the decoding of the speech signal by the decoding network, and during the decoding process, performing hot word excitation on the speech signal by the hot word decoding network to obtain a corresponding speech recognition text, comprises:
[0020] Obtaining the incentive score of the corresponding slot of the hot word decoding network;
[0021] For each speech signal frame in the speech signal, the speech signal frame is decoded according to the decoding network, and during the decoding process, the score of the decoding token in the hot word decoding network is incentivized according to the incentive score of the corresponding slot of the hot word decoding network;
[0022] After the last frame of speech signal is decoded, the decoding token with the maximum score is selected and backtracked to obtain the speech recognition text.
[0023] Optionally, the speech signal frame is decoded according to the decoding network, and during the decoding process, the score of the decoding token in the hot word decoding network is incentivized according to the incentive score of the corresponding slot of the hot word decoding network, including:
[0024] Determine the currently active decoding token;
[0025] For each currently active decoded token, pass the currently active decoded token through the decoding network. After all the active decoded tokens have been passed, all the decoded tokens corresponding to the current speech signal frame are obtained, where the scores of the decoded tokens passed through the hotword decoding network include the excitation scores of the corresponding slots of the hotword decoding network;
[0026] Determine the active decoded tokens corresponding to the current speech signal frame from all the decoded tokens corresponding to the current speech signal frame, and use them as the currently active decoded tokens.
[0027] Optionally, when all the decoded tokens corresponding to the current speech signal frame are in the main decoding network, the step of determining the active decoded tokens corresponding to the current speech signal frame from all the decoded tokens corresponding to the current speech signal frame includes:
[0028] Determine a preset number of decoded tokens with the top-ranked scores from all the decoded tokens corresponding to the current speech signal frame, and use them as the active decoded tokens corresponding to the current speech signal frame.
[0029] Optionally, when the decoded tokens corresponding to the current speech signal frame include the decoded tokens in the hotword decoding network, the step of determining the active decoded tokens corresponding to the current speech signal frame from all the decoded tokens corresponding to the current speech signal frame includes:
[0030] Determine a first set of decoded tokens and a second set of decoded tokens from all the decoded tokens corresponding to the current speech signal frame; determine the decoded tokens in the first set of decoded tokens and the second set of decoded tokens as the active decoded tokens corresponding to the current speech signal frame;
[0031] where the first set of decoded tokens includes a preset number of decoded tokens with the top-ranked scores among the decoded tokens in the main decoding network;
[0032] The second set of decoded tokens includes a preset number of decoded tokens with the top-ranked scores among the decoded tokens in the hotword decoding network.
[0033] Optionally, after using the decoding network to decode the speech signal and using the hotword decoding network to perform hotword excitation on the speech signal during the decoding process to obtain the corresponding speech recognition text, the method further includes:
[0034] Obtain a pre-determined grammar rule, where the grammar rule is used to indicate the syntactic information of each second hotword in the second hotword list;
[0035] Optimize the sentence patterns of the second hot words included in the speech recognition text according to the grammar rules.
[0036] Optionally, the main decoding network is a weighted finite state transducer (WFST) network; the hot word decoding network is a finite state automaton (FSA) network.
[0037] A speech recognition device, the device includes:
[0038] A speech signal acquisition unit, configured to acquire a speech signal to be recognized;
[0039] A decoding network acquisition unit, configured to acquire a pre-constructed decoding network, the decoding network including a main decoding network and a hot word decoding network;
[0040] An identification unit, configured to decode the speech signal by using the decoding network, and during the decoding process, perform hot word excitation on the speech signal by using the hot word decoding network to obtain a corresponding speech recognition text.
[0041] Optionally, the device includes a decoding network construction unit, and the decoding network construction unit includes:
[0042] A main decoding network construction unit, configured to construct a main decoding network, where at least one slot is included in the main decoding network, and each slot is located between two nodes;
[0043] A first hot word list acquisition unit, configured to acquire a first hot word list, where one or more first hot words are included in the first hot word list;
[0044] A hot word decoding network construction unit, configured to construct a hot word decoding network based on the first hot word list, where a plurality of branches are included in the hot word decoding network, the number of branches is the same as the number of first hot words in the first hot word list, and the arcs at the head and tail of each branch are silent arcs;
[0045] A decoding network generation unit, configured to insert the hot word decoding network into each slot in the main decoding network to generate a decoding network.
[0046] Optionally, the decoding network generation unit includes:
[0047] A removal unit, configured to remove the silent arcs at the head and tail of each branch in the hot word decoding network for each branch in the hot word decoding network;
[0048] A setting unit for setting the first phoneme in the three - phonemes corresponding to the real arc at the head of the branch as the last phoneme in the three - phonemes corresponding to the first real arc before the slot in the main decoding network, and setting the last phoneme in the three - phonemes corresponding to the real arc at the tail of the branch as the first phoneme in the three - phonemes corresponding to the first real arc after the slot in the main decoding network;
[0049] A connection unit for connecting the real arc at the head of the branch to the first node before the slot in the main decoding network, and connecting the real arc at the tail of the branch to the first node after the slot in the main decoding network.
[0050] Optionally, the recognition unit includes:
[0051] An excitation score acquisition unit for acquiring the excitation score of the corresponding slot of the hot - word decoding network;
[0052] An excitation unit for, for each speech signal frame in the speech signal, decoding the speech signal frame according to the decoding network, and during the decoding process, exciting the score of the decoding token in the hot - word decoding network according to the excitation score of the corresponding slot of the hot - word decoding network;
[0053] A backtracking unit for, after completing the decoding of the last speech signal frame, selecting the decoding token with the maximum score and backtracking to obtain the speech recognition text.
[0054] Optionally, the excitation unit includes:
[0055] A token determination unit for determining the current active decoding token;
[0056] An all - decoding - token determination unit for, for each current active decoding token, transmitting the current active decoding token in the decoding network. After all active decoding tokens are transmitted, all decoding tokens corresponding to the current speech signal frame are obtained, where the score of the decoding token transmitted in the hot - word decoding network includes the excitation score of the corresponding slot of the hot - word decoding network;
[0057] An active - decoding - token determination unit for determining the active decoding token corresponding to the current speech signal frame from all decoding tokens corresponding to the current speech signal frame as the current active decoding token.
[0058] Optionally, when all decoding tokens corresponding to the current speech signal frame are in the main decoding network, the active - decoding - token determination unit is specifically configured to:
[0059] Determine a preset number of decoding tokens with top scores from all the decoding tokens corresponding to the current speech signal frame as the active decoding tokens corresponding to the current speech signal frame.
[0060] Optionally, when the decoding tokens corresponding to the current speech signal frame include the decoding tokens of the hotword decoding network, the active decoding token determination unit is specifically configured to:
[0061] Determine a first decoding token set and a second decoding token set from all the decoding tokens corresponding to the current speech signal frame; determine the decoding tokens in the first decoding token set and the second decoding token set as the active decoding tokens corresponding to the current speech signal frame;
[0062] Wherein, the first decoding token set includes a preset number of decoding tokens with top scores among the decoding tokens of the main decoding network;
[0063] The second decoding token set includes a preset number of decoding tokens with top scores among the decoding tokens of the hotword decoding network.
[0064] Optionally, the device further includes:
[0065] A grammar rule acquisition unit, configured to, after decoding the speech signal by using the decoding network and performing hotword excitation on the speech signal by using the hotword decoding network during the decoding process to obtain a corresponding speech recognition text, acquire a preset grammar rule, where the grammar rule is used to indicate the sentence pattern information of each second hotword in the second hotword list;
[0066] An optimization unit, configured to optimize the sentence patterns of the second hotwords included in the speech recognition text according to the grammar rule.
[0067] Optionally, the main decoding network is a weighted finite state transducer (WFST) network; the hotword decoding network is a finite state automaton (FSA) network.
[0068] A speech recognition device includes a memory and a processor;
[0069] The memory is configured to store a program;
[0070] The processor is configured to execute the program to implement each step of the speech recognition method as described above.
[0071] A readable storage medium stores a computer program, and when the computer program is executed by a processor, each step of the speech recognition method as described above is implemented.
[0072] With the above technical solution, the present application discloses a voice recognition method, related devices and a readable storage medium. In this solution, by pre-constructing a decoding network, the decoding network includes a main decoding network and a hotword decoding network inserted in the main decoding network. After obtaining the voice signal to be recognized, the decoding network is used to decode the voice signal, and during the decoding process, the hotword decoding network is used to perform hotword excitation on the voice signal to obtain the corresponding voice recognition text. Based on this solution, only one decoding process is required for the voice signal to achieve hotword excitation. Therefore, this solution can improve the recognition efficiency of hotwords. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0074] Figure 1 is a schematic flowchart of the voice recognition method disclosed in the embodiment of the present application;
[0075] Figure 2 is a schematic diagram of a decoding network structure disclosed in the embodiment of the present application;
[0076] Figure 3 is a schematic diagram of a main decoding network structure disclosed in the embodiment of the present application;
[0077] Figure 4 is a schematic diagram of a hotword decoding network structure disclosed in the embodiment of the present application;
[0078] Figure 5 is a schematic diagram of a decoding network structure disclosed in the embodiment of the present application;
[0079] Figure 6 is a schematic diagram of the setting of the real arc at the head of the hotword decoding network disclosed in the embodiment of the present application;
[0080] Figure 7 is a schematic diagram of the structure of the voice recognition device disclosed in the embodiment of the present application;
[0081] Figure 8 is a hardware structure block diagram of a voice recognition device disclosed in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0083] Next, the speech recognition method provided by the present application will be introduced through the following embodiments.
[0084] Referring to Figure 1 , Figure 1 which is a schematic flowchart of the speech recognition method disclosed in the embodiments of the present application. The method may include:
[0085] Step S101: Obtain a speech signal to be recognized.
[0086] In the present application, the speech signal to be recognized may be a speech signal in any scenario. The present application does not make any limitation thereto. Additionally, the speech signal to be recognized may be an original speech signal collected by a speech acquisition device, such as a microphone, or a speech signal obtained after preprocessing the original speech signal (such as speech enhancement, etc.). The present application does not make any limitation thereto either.
[0087] Step S102: Obtain a pre-constructed decoding network, where the decoding network includes a main decoding network and a hotword decoding network.
[0088] Since the currently common decoding network is a WFST (Weighted Finite-State Transducers) network constructed based on a trained acoustic model, language model, and pronunciation dictionary, and the label on the FSA (Finite-State Automation) edge only has input and no weight, which is more suitable for the hotword scenario. Therefore, as an implementable manner, in the present application, the main decoding network is a WFST network; the hotword decoding network is an FSA network.
[0089] It should be noted that the main decoding network and the hotword decoding network in the present application need to be processed specifically during construction to integrate the hotword decoding network into the main decoding network to obtain a complete decoding network. The specific processing method will be described in detail through subsequent embodiments, and this embodiment will not be elaborated further.
[0090] Step S103: Use the decoding network to decode the speech signal. During the decoding process, use the hotword decoding network to perform hotword excitation on the speech signal to obtain the corresponding speech recognition text.
[0091] In the present application, the excitation score of the hot word decoding network can be preset. During the decoding process, the speech signal is subjected to hot word excitation based on the excitation score of the hot word decoding network. In this case, the recall rate of hot words in the speech recognition text will be improved. The specific implementation manner of using the hot word decoding network to perform hot word excitation on the speech signal will be described in detail in the following embodiments and will not be elaborated in this embodiment.
[0092] In this embodiment, a speech recognition method is disclosed. In this method, a decoding network is pre-constructed. The decoding network includes a main decoding network and a hot word decoding network. After obtaining the speech signal to be recognized, the decoding network is used to decode the speech signal, and during the decoding process, the hot word decoding network is used to perform hot word excitation on the speech signal to obtain the corresponding speech recognition text. Based on this solution, only one decoding process is required for the speech signal to achieve hot word excitation. Therefore, this method can improve the recognition efficiency of hot words.
[0093] It should be noted that for the sake of understanding, in the following embodiments, the main decoding network is a WFST network and the hot word decoding network is an FSA network as an example to elaborate on the entire solution in detail. However, based on the idea of the present application, the decoding network obtained from two networks in other forms is also within the protection scope of the present application.
[0094] In another embodiment of the present application, the construction method of the decoding network is introduced in detail. The method may include the following steps:
[0095] Step S201: Construct a main decoding network. The main decoding network includes at least one slot, and each slot is located between two nodes.
[0096] In the present application, at least one slot can be reserved during the process of packing the main decoding network model resources, and the rest is the same as the normal model resource packing process. When constructing the main decoding network, the speech recognition engine can read the main decoding network model resources and reconstruct the main decoding network therefrom. The main decoding network includes at least one slot, and each slot is located between two nodes.
[0097] Step S202: Obtain a first hot word list, where the first hot word list includes one or more first hot words.
[0098] In the field of speech recognition, hot words are generally divided into two types. One is a global hot word, which usually does not appear following a fixed sentence pattern and can appear in any part of a sentence. The other is a local hot word, which often appears following a fixed sentence pattern. In the present application, the first hot word list can be a global hot word list.
[0099] Step S203: Based on the first hot word list, construct a hot word decoding network.
[0100] In this application, the hot word decoding network includes multiple branches, and the number of branches is the same as the number of the first hot words in the first hot word list. The arcs at the head and tail of each branch are silent arcs.
[0101] Step S204: For each slot in the main decoding network, insert the hot word decoding network into the slot to generate a decoding network.
[0102] Refer to Figure 2 , Figure 2 which is a schematic diagram of a decoding network structure disclosed in an embodiment of this application. As Figure 2 shown, there is a slot between node 5 and node 6 in the main decoding network, that is, $reserve shown in the figure. The hot word decoding network is inserted into this slot. The hot word decoding network includes 50 branches, that is, FSA1, FSA2,..., FSA50 shown in the figure. Each branch head and tail includes an arc, and this arc is a silent arc (sil arc).
[0103] In another embodiment of this application, the implementation manner of inserting the hot word decoding network into the slot is introduced in detail, and this manner may include the following steps:
[0104] Step S301: For each branch in the hot word decoding network, remove the silent arcs at the head and tail of each branch in the hot word decoding network.
[0105] Step S302: Set the first phoneme in the triphone corresponding to the real arc at the branch head as the last phoneme in the triphone corresponding to the first real arc before the slot in the main decoding network, and set the last phoneme in the triphone corresponding to the real arc at the branch tail as the first phoneme in the triphone corresponding to the first real arc after the slot in the main decoding network.
[0106] It should be noted that in this application, the modeling granularity of the main decoding network is the same as that of the hot word decoding network. Since the modeling granularity of the currently commonly used acoustic model is triphone, in this application, the modeling granularity of the main decoding network and the hot word decoding network is both triphone. Then each real arc in the main decoding network corresponds to a triphone, and each branch in the hot word decoding network corresponds to the triphone of a first hot word.
[0107] It should be noted that when the first phoneme in the triphone corresponding to the real arc at the head of the branch is set to the last phoneme in the triphone corresponding to the first real arc before the slot in the main decoding network, and the last phoneme in the triphone corresponding to the real arc at the tail of the branch is set to the first phoneme in the triphone corresponding to the first real arc after the slot in the main decoding network, it is necessary to adjust the number of nodes or arcs in the branch, such as adding nodes and / or arcs, or deleting nodes and / or arcs.
[0108] Step S303: Connect the real arc at the head of the branch to the first node before the slot in the main decoding network, and connect the real arc at the tail of the branch to the first node after the slot in the main decoding network.
[0109] For ease of understanding, an example of constructing a decoding network is given in this application, as follows:
[0110] Reference Figure 3 , Figure 3 Schematic diagram of a main decoding network structure disclosed in an embodiment of the present application. The main decoding network is constructed based on the sentence "ring XXX on". The main decoding network includes two slots, namely $reserve1 and $reserve2 shown in the figure.
[0111] Reference Figure 4 , Figure 4 This is a schematic diagram of a hot word decoding network structure disclosed in an embodiment of the present application. The hot word decoding network is constructed based on the two first hot words "rhylee" and "luca".
[0112] Reference Figure 5 , Figure 5 A schematic diagram of a decoding network structure disclosed in an embodiment of the present application is shown in FIG. Figure 4 The hotword decoding network shown is inserted Figure 3 The $reserve1 and $reserve2 in the main decoding network shown in FIG. Figure 5 As shown, in Figure 4 The hotword decoding network shown is inserted Figure 3 When $reserve2 is used in the main decoding network shown, two arcs and one node need to be added in front of each entry node and behind each exit node of the hot word decoding network (not shown in the figure).
[0113] Reference Figure 6 , Figure 6 This is a schematic diagram of a hot word decoding network head real arc setting disclosed in an embodiment of the present application, Figure 6 The part shown is Figure 5 Detailed diagram of node 24 and its surrounding nodes, combined withFigure 5 and Figure 6 As shown in Figure 6 , node 24 is a newly added node, whose purpose is to connect node 6 of the main decoding network and node 13_1 of the hotword decoding network. Since the incoming arc triphone of node 6 is "En_r-En_ih+En_ng", and the outgoing arc of node 13_1 is "En_ih-En_r+En_ay", node 24 needs to newly add an incoming arc "En_ih-En_ng+En_ih" and an outgoing arc "En_ng-En_ih+En_r".
[0114] Since the decoding network of the present application has structural adjustments compared with the existing decoding network, and the existing decoding method cannot achieve a good hotword excitation effect, therefore, the existing decoding method is also improved in the present application.
[0115] In another embodiment of the present application, the specific implementation manner of decoding the voice signal by using the decoding network and performing hotword excitation on the voice signal by using the hotword decoding network during the decoding process to obtain the corresponding speech recognition text is introduced in detail. The method may include the following steps:
[0116] Step S401: Obtain the excitation score of the corresponding slot of the hotword decoding network.
[0117] In the present application, when constructing the main decoding network, the excitation score can be preset for each slot in the main decoding network. The excitation scores of different slots can be the same or different.
[0118] Step S402: For each voice signal frame in the voice signal, decode the voice signal frame according to the decoding network, and during the decoding process, excite the score of the decoding token in the hotword decoding network according to the excitation score of the corresponding slot of the hotword decoding network.
[0119] As an implementable manner, the specific implementation manner of decoding the voice signal frame according to the decoding network and exciting the score of the decoding token in the hotword decoding network according to the excitation score of the corresponding slot of the hotword decoding network may include the following steps:
[0120] Step S4021: Determine the current active decoding token.
[0121] In the present application, when the voice signal frame is the first voice signal frame, the current active decoding token is the preset original decoding token, and when the voice signal frame is a non-first voice signal frame, the current active decoding token is the active decoding token corresponding to the previous voice signal frame adjacent to the voice signal frame.
[0122] Step S4022: For each current active decoding token, pass the current active decoding token through the decoding network. After all the active decoding tokens have been passed, all the decoding tokens corresponding to the current speech signal frame are obtained. Among them, the scores of the decoding tokens that have been passed through the hotword decoding network include the excitation scores of the corresponding slots of the hotword decoding network.
[0123] Step S4023: Determine the active decoding tokens corresponding to the current speech signal frame from all the decoding tokens corresponding to the current speech signal frame as the current active decoding tokens.
[0124] As an implementable manner, when all the decoding tokens corresponding to the current speech signal frame are in the main decoding network, determining the active decoding tokens corresponding to the current speech signal frame from all the decoding tokens corresponding to the current speech signal frame includes:
[0125] Determine a preset number of decoding tokens with the top-ranked scores from all the decoding tokens corresponding to the current speech signal frame as the active decoding tokens corresponding to the current speech signal frame.
[0126] As another implementable manner, when the decoding tokens corresponding to the current speech signal frame include the decoding tokens in the hotword decoding network, determining the active decoding tokens corresponding to the current speech signal frame from all the decoding tokens corresponding to the current speech signal frame includes:
[0127] Determine a first set of decoding tokens and a second set of decoding tokens from all the decoding tokens corresponding to the current speech signal frame; determine the decoding tokens in the first set of decoding tokens and the second set of decoding tokens as the active decoding tokens corresponding to the current speech signal frame;
[0128] Among them, the first set of decoding tokens includes a preset number of decoding tokens with the top-ranked scores among the decoding tokens in the main decoding network;
[0129] The second set of decoding tokens includes a preset number of decoding tokens with the top-ranked scores among the decoding tokens in the hotword decoding network.
[0130] It should be noted that the preset number can be set according to the scenario requirements, and the present application does not make any limitations.
[0131] Step S403: After completing the decoding of the last speech signal frame, select the decoding token with the maximum score and backtrack to obtain the speech recognition text.
[0132] As pointed out in the foregoing, in the field of speech recognition, hot words are generally divided into two types. One is the global hot word, which usually does not appear following a fixed sentence pattern and can appear in any part of a sentence. The other is the local hot word, which often appears following a fixed sentence pattern. After completing the recognition of the global hot word using the above solution, it is often necessary to recognize the local hot word. Therefore, in this application, a local hot word recognition solution is also provided, which will be described in detail through the following embodiments:
[0133] In another embodiment of this application, after using the decoding network to decode the speech signal and obtaining the corresponding speech recognition text by using the hot word decoding network to perform hot word excitation on the speech signal during the decoding process, the method further includes:
[0134] Step S501: Obtain a pre-determined grammar rule, where the grammar rule is used to indicate the sentence pattern information of each second hot word in the second hot word list;
[0135] In this application, the second hot word list may be a local hot word list. The specific content of the grammar rule can be set based on the scenario requirements, and this application does not make any limitations.
[0136] Step S502: Optimize the sentence pattern of the second hot word included in the speech recognition text according to the grammar rule.
[0137] In this application, when optimizing the sentence pattern of the second hot word included in the speech recognition text according to the grammar rule, it can be implemented based on the way of dictionary tree search. In the traditional solution, for hot words with specific sentence pattern requirements, it is based on the matching method using the FSA network. The time complexity of dictionary tree search for strings is O(n), where n is the length of the string to be matched. Compared with the traditional solution of using the FSA network for hot word matching, the speed is faster.
[0138] The speech recognition device disclosed in the embodiments of this application will be described below. The speech recognition device described below can be correspondingly referred to the speech recognition method described above.
[0139] Referring to Figure 7 , Figure 7 is a schematic structural diagram of a speech recognition device disclosed in an embodiment of this application. As Figure 7 shown, the speech recognition device may include:
[0140] A speech signal acquisition unit 71, configured to acquire a speech signal to be recognized;
[0141] A decoding network acquisition unit 72, configured to acquire a pre-constructed decoding network, where the decoding network includes a main decoding network and a hot word decoding network;
[0142] An identification unit 73, configured to decode the speech signal by using the decoding network. During the decoding process, the hotword decoding network is used to perform hotword excitation on the speech signal to obtain a corresponding speech recognition text.
[0143] As an implementable manner, the apparatus includes a decoding network construction unit, and the decoding network construction unit includes:
[0144] A main decoding network construction unit, configured to construct a main decoding network, where at least one slot is included in the main decoding network, and each slot is located between two nodes;
[0145] A first hotword list acquisition unit, configured to acquire a first hotword list, where the first hotword list includes one or more first hotwords;
[0146] A hotword decoding network construction unit, configured to construct a hotword decoding network based on the first hotword list. The hotword decoding network includes multiple branches, and the number of branches is the same as the number of first hotwords in the first hotword list. The arcs at the head and tail of each branch are silent arcs;
[0147] A decoding network generation unit, configured to insert the hotword decoding network into each slot in the main decoding network to generate a decoding network.
[0148] As an implementable manner, the decoding network generation unit includes:
[0149] A removal unit, configured to remove the silent arcs at the head and tail of each branch in the hotword decoding network;
[0150] A setting unit, configured to set the first phoneme in the triphone corresponding to the real arc at the head of the branch as the last phoneme in the triphone corresponding to the first real arc before the slot in the main decoding network, and set the last phoneme in the triphone corresponding to the real arc at the tail of the branch as the first phoneme in the triphone corresponding to the first real arc after the slot in the main decoding network;
[0151] A connection unit, configured to connect the real arc at the head of the branch to the first node before the slot in the main decoding network, and connect the real arc at the tail of the branch to the first node after the slot in the main decoding network.
[0152] As an implementable manner, the identification unit includes:
[0153] An excitation score acquisition unit, configured to acquire the excitation score of the slot corresponding to the hotword decoding network;
[0154] An excitation unit, configured to decode each speech signal frame in the speech signal according to the decoding network. During the decoding process, the score of the decoding token in the hotword decoding network is excited according to the excitation score of the corresponding slot in the hotword decoding network.
[0155] A backtracking unit, configured to select the decoding token with the maximum score after completing the decoding of the last speech signal frame, and backtrack to obtain the speech recognition text.
[0156] As an implementable manner, the excitation unit includes:
[0157] A token determination unit, configured to determine the currently active decoding token.
[0158] An all-decoding-token determination unit, configured to transmit the currently active decoding token in the decoding network for each currently active decoding token. After all active decoding tokens are transmitted, all decoding tokens corresponding to the current speech signal frame are obtained, where the score of the decoding token transmitted in the hotword decoding network includes the excitation score of the corresponding slot in the hotword decoding network.
[0159] An active-decoding-token determination unit, configured to determine the active decoding token corresponding to the current speech signal frame from all decoding tokens corresponding to the current speech signal frame as the currently active decoding token.
[0160] As an implementable manner, when all decoding tokens corresponding to the current speech signal frame are in the main decoding network, the active-decoding-token determination unit is specifically configured to:
[0161] Determine a preset number of decoding tokens with the top-ranked scores from all decoding tokens corresponding to the current speech signal frame as the active decoding tokens corresponding to the current speech signal frame.
[0162] As an implementable manner, when the decoding tokens corresponding to the current speech signal frame include the decoding tokens in the hotword decoding network, the active-decoding-token determination unit is specifically configured to:
[0163] Determine a first decoding token set and a second decoding token set from all decoding tokens corresponding to the current speech signal frame; and determine the decoding tokens in the first decoding token set and the second decoding token set as the active decoding tokens corresponding to the current speech signal frame.
[0164] Wherein, the first decoding token set includes a preset number of decoding tokens with the top-ranked scores among the decoding tokens in the main decoding network.
[0165] The second decoding token set includes a preset number of decoding tokens with top scores among the decoding tokens in the hotword decoding network.
[0166] As an implementable manner, the device further includes:
[0167] A grammar rule acquisition unit, configured to, after decoding the speech signal by using the decoding network and performing hotword excitation on the speech signal by using the hotword decoding network during the decoding process to obtain a corresponding speech recognition text, acquire a pre-determined grammar rule, where the grammar rule is used to indicate the sentence pattern information of each second hotword in the second hotword list;
[0168] An optimization unit, configured to optimize the sentence patterns of the second hotwords included in the speech recognition text according to the grammar rule.
[0169] As an implementable manner, the main decoding network is a weighted finite state transducer (WFST) network; the hotword decoding network is a finite state automaton (FSA) network.
[0170] Referring to Figure 8 , Figure 8 is a hardware structure block diagram of the speech recognition device provided in the embodiment of the present application. Referring to Figure 8 ,the hardware structure of the speech recognition device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0171] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4;
[0172] The processor 1 may be a central processing unit (CPU), or a specific integrated circuit (ASIC) (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0173] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory;
[0174] Wherein, the memory stores a program, and the processor can call the program stored in the memory, and the program is used for:
[0175] Acquire a speech signal to be recognized;
[0176] Acquire a pre-constructed decoding network, where the decoding network includes a main decoding network and a hotword decoding network;
[0177] Decode the speech signal by using the decoding network. During the decoding process, use the hotword decoding network to perform hotword excitation on the speech signal to obtain the corresponding speech recognition text.
[0178] Optionally, the refinement function and expansion function of the program can be referred to the above description.
[0179] The embodiment of the present application also provides a readable storage medium, which can store a program suitable for execution by a processor. The program is used for:
[0180] Obtain a speech signal to be recognized;
[0181] Obtain a pre-constructed decoding network, where the decoding network includes a main decoding network and a hotword decoding network;
[0182] Decode the speech signal by using the decoding network. During the decoding process, use the hotword decoding network to perform hotword excitation on the speech signal to obtain the corresponding speech recognition text.
[0183] Optionally, the refinement function and expansion function of the program can be referred to the above description.
[0184] Finally, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0185] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0186] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A speech recognition method, characterized in that, The method includes: Obtaining a speech signal to be recognized; Obtaining a pre-constructed decoding network, where the decoding network includes a main decoding network and a hotword decoding network; the hotword decoding network is incorporated into the main decoding network; the main decoding network is a weighted finite state transducer (WFST) network; the hotword decoding network is a finite state automaton (FSA) network; Decoding the speech signal using the decoding network. During the decoding process, using the hotword decoding network to perform hotword excitation on the speech signal to obtain a corresponding speech recognition text. The step of using the hotword decoding network to perform hotword excitation on the speech signal includes performing hotword excitation on the speech signal based on a preset excitation score of the hotword decoding network; The method for constructing the decoding network includes: constructing a main decoding network, where at least one slot is included in the main decoding network, and each slot is located between two nodes; for each slot in the main decoding network, inserting the hotword decoding network into the slot to generate a decoding network.
2. The method according to claim 1, characterized in that, The method for constructing the decoding network further includes: Obtaining a first hotword list, where the first hotword list includes one or more first hotwords; Based on the first hotword list, constructing a hotword decoding network, where the hotword decoding network includes multiple branches, and the number of branches is the same as the number of first hotwords in the first hotword list. The arcs at the head and tail of each branch are silent arcs.
3. The method according to claim 1, characterized in that, The step of inserting the hotword decoding network into the slot includes: For each branch in the hotword decoding network, removing the silent arcs at the head and tail of each branch in the hotword decoding network; Setting the first phoneme in the triphone corresponding to the real arc at the head of the branch to the last phoneme in the triphone corresponding to the first real arc before the slot in the main decoding network, and setting the last phoneme in the triphone corresponding to the real arc at the tail of the branch to the first phoneme in the triphone corresponding to the first real arc after the slot in the main decoding network; Connecting the real arc at the head of the branch to the first node before the slot in the main decoding network, and connecting the real arc at the tail of the branch to the first node after the slot in the main decoding network.
4. The method according to claim 1, wherein The step of decoding the speech signal using the decoding network. During the decoding process, using the hotword decoding network to perform hotword excitation on the speech signal to obtain a corresponding speech recognition text includes: Obtaining the excitation score of the slot corresponding to the hotword decoding network; For each speech signal frame in the speech signal, decoding the speech signal frame according to the decoding network. During the decoding process, exciting the score of the decoding token in the hotword decoding network according to the excitation score of the slot corresponding to the hotword decoding network; After completing the decoding of the last speech signal frame, selecting the decoding token with the maximum score and backtracking to obtain the speech recognition text.
5. The method according to claim 4, characterized in that, The step of decoding the speech signal frame according to the decoding network. During the decoding process, exciting the score of the decoding token in the hotword decoding network according to the excitation score of the slot corresponding to the hotword decoding network includes: Determine the currently active decoding token; For each currently active decoding token, pass the currently active decoding token through the decoding network. After all active decoding tokens have been passed, all decoding tokens corresponding to the current speech signal frame are obtained. Among them, the scores of the decoding tokens passed through the hotword decoding network include the excitation scores of the corresponding slots of the hotword decoding network; From all the decoding tokens corresponding to the current speech signal frame, determine the active decoding tokens corresponding to the current speech signal frame as the currently active decoding tokens.
6. The method according to claim 5, characterized in that, When all the decoding tokens corresponding to the current speech signal frame are in the main decoding network, determining the active decoding tokens corresponding to the current speech signal frame from all the decoding tokens corresponding to the current speech signal frame includes: From all the decoding tokens corresponding to the current speech signal frame, determine a preset number of decoding tokens with the top-ranked scores as the active decoding tokens corresponding to the current speech signal frame.
7. The method according to claim 5, wherein When the decoding tokens corresponding to the current speech signal frame include the decoding tokens in the hotword decoding network, determining the active decoding tokens corresponding to the current speech signal frame from all the decoding tokens corresponding to the current speech signal frame includes: From all the decoding tokens corresponding to the current speech signal frame, determine a first decoding token set and a second decoding token set; determine the decoding tokens in the first decoding token set and the second decoding token set as the active decoding tokens corresponding to the current speech signal frame; Among them, the first decoding token set includes a preset number of decoding tokens with the top-ranked scores among the decoding tokens in the main decoding network; The second decoding token set includes a preset number of decoding tokens with the top-ranked scores among the decoding tokens in the hotword decoding network.
8. The method according to claim 1, wherein After using the decoding network to decode the speech signal and using the hotword decoding network to perform hotword excitation on the speech signal during the decoding process to obtain the corresponding speech recognition text, the method further includes: Obtain a pre-determined grammar rule, where the grammar rule is used to indicate the sentence pattern information of each second hotword in the second hotword list; Optimize the sentence patterns of the second hotwords included in the speech recognition text according to the grammar rule.
9. A voice recognition device, characterized in that, The device includes: A speech signal acquisition unit for acquiring the speech signal to be recognized; A decoding network acquisition unit for acquiring the decoding network pre-constructed by the decoding network construction unit. The decoding network includes a main decoding network and a hotword decoding network; the hotword decoding network is integrated into the main decoding network; the main decoding network is a weighted finite state transducer (WFST) network; the hotword decoding network is a finite state automaton (FSA) network; An identification unit, configured to decode the speech signal by using the decoding network. During the decoding process, the speech signal is subjected to hotword excitation by using the hotword decoding network to obtain a corresponding speech recognition text. The step of subjecting the speech signal to hotword excitation by using the hotword decoding network includes subjecting the speech signal to hotword excitation based on a preset excitation score of the hotword decoding network; The decoding network construction unit includes: A main decoding network construction unit, configured to construct a main decoding network, where at least one slot is included in the main decoding network, and each slot is located between two nodes; A decoding network generation unit, configured to insert the hotword decoding network into each slot in the main decoding network to generate a decoding network.
10. A voice recognition device, characterized in that, It includes a memory and a processor; The memory is configured to store a program; The processor is configured to execute the program to implement each step of the speech recognition method according to any one of claims 1 to 8.
11. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, each step of the speech recognition method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Method and system for automatically recognizing voice
CN103971686A
Speech recognition method and system based on self-adaptive hot word weight
CN111354347A
Voice recognition method, device, equipment, system and storage medium
CN113436614A