Speech Recognition Method, Apparatus, Storage Medium, and Electronic Device

By separating hot words from the language recognition model and building a hot word index model, the problem of insufficient flexibility in hot word updates in speech recognition is solved, and more accurate and reliable speech recognition results are achieved.

CN115240644BActive Publication Date: 2025-05-27NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210845333.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2025-05-27
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

The flexibility of hot words updates in existing speech recognition technologies is weak, and the hot words cannot be updated in real time, resulting in inaccurate recognition results.

Method used

By separating hot words from the language recognition model, a pre-constructed hot word index model is constructed, the recognition score and hot word score of the speech recognition model are obtained, and content decoding is performed based on these scores to achieve real-time update of hot words.

Benefits of technology

It improves the flexibility of hot word updates in the speech recognition process, enhances the accuracy and reliability of the recognition results, and can update hot word without changing the speech recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240644B_ABST
    Figure CN115240644B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech recognition method, apparatus, storage medium, and electronic device, relating to the field of artificial intelligence technology. The speech recognition method includes: obtaining recognition scores of each frame of speech in the speech to be recognized by a speech recognition model; determining hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model, where the hot word scores are used to represent whether each frame of speech contains a preset hot word; and performing content decoding on the speech to be recognized based on the recognition scores and the hot word scores of each frame of speech to obtain a content recognition result of the speech to be recognized. The technical problem of weak flexibility in hot word update during the current speech recognition process is solved, and the technical effect of improving the flexibility of hot word update is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] ASR (Automatic Speech Recognition) is a technology that converts human language into text and is widely used in various commercial products to improve the speed and convenience of people's communication, thereby enhancing the user experience. However, although some special words have only one pronunciation, they have different meanings in different scenarios. For example, the place names Hengshan and Heng Mountain, which will result in inaccurate speech recognition results.

[0003] Different users have different hot words. In speech recognition, the enhancement of hot words is mainly achieved by modifying the WFST (weighted finite-state transducer) static language model. However, this method requires adding hot words to the original language model to regenerate a new language model. This method has high accuracy, but it cannot update hot words in real time.

[0004] Therefore, the flexibility of hot word update in the current speech recognition process is weak. Summary of the Invention

[0005] The present disclosure provides a speech recognition method, apparatus, storage medium, and electronic device, thereby improving the flexibility of hot word update in the speech recognition process.

[0006] In a first aspect, an embodiment of the present disclosure provides a speech recognition method, including:

[0007] Obtaining the recognition scores of each frame of speech in the speech to be recognized by a speech recognition model;

[0008] Determining the hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model; wherein, the hot word score is used to represent whether each frame of speech contains a preset hot word;

[0009] Performing content decoding on the speech to be recognized based on the recognition scores and hot word scores of each frame of speech to obtain the content recognition result of the speech to be recognized.

[0010] In an optional embodiment of the present disclosure, before determining the hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model, the method further includes:

[0011] Performing word segmentation processing on each obtained preset hot word to obtain a plurality of hot word segments;

[0012] Using a hot word segment as an index node to construct a hot word index model including multiple hot word index paths.

[0013] In an optional embodiment of the present disclosure, determining the hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model includes:

[0014] Performing hot word indexing on each frame of speech in sequence along multiple hot word index paths in the hot word index model, and determining the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentation of each index node in the hot word index model;

[0015] Determining the hot word scores of each frame of speech according to the node scores of each frame of speech at each index node.

[0016] In an optional embodiment of the present disclosure, determining the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentation of each index node in the hot word index model includes:

[0017] Performing word segmentation processing on the current frame of speech to obtain multiple frame segmentations;

[0018] Indexing the multiple frame segmentations based on the hot word index model. If the current frame segmentation matches the hot word segmentation of the current index node, then determining the reward score of the current index node as the sum of the node score of the previous index node and a preset incremental score;

[0019] If the current frame segmentation does not match the hot word segmentation of the current index node, then determining the cumulative reward score of the historical index node as the penalty score of the current index node.

[0020] In an optional embodiment of the present disclosure, if the current frame segmentation does not match the hot word segmentation of the current index node, then determining the cumulative reward score of the historical index node as the penalty score of the current index node includes:

[0021] If the current index node is the end node of the current hot word index path, then determining the penalty score of the current index node as a preset penalty score.

[0022] In an optional embodiment of the present disclosure, performing content decoding on the speech to be recognized based on the recognition scores and hot word scores of each frame of speech to obtain the content recognition result of the speech to be recognized includes:

[0023] If the next frame segmentation matches the hot word segmentation of at least the next index node, then correcting the recognition score based on the reward score of the current frame segmentation to obtain the target score of the next frame segmentation at the next index node; wherein, the next frame segmentation refers to the segmentation that is located after the current frame segmentation and adjacent to the current frame segmentation in the word segmentation order; the next index node refers to the node that is located after the current index node and adjacent to the current index node along the index direction in the hot word index model;

[0024] Perform content decoding on the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

[0025] In an optional embodiment of the present disclosure, perform content decoding on the speech to be recognized based on the recognition scores of each frame of speech and the hotword scores to obtain the content recognition result of the speech to be recognized, including:

[0026] If the word segmentation of the next frame does not match the hotword segmentations of all the next index nodes, correct the recognition score based on the penalty score of the current frame's word segmentation to obtain the target score of the word segmentation of the next frame at the next index node;

[0027] Perform content decoding on the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

[0028] In an optional embodiment of the present disclosure, the method further includes:

[0029] Obtain the hotword to be updated;

[0030] If the hotword to be updated belongs to the training sample set of the speech recognition model, update the hotword index model based on the hotword to be updated.

[0031] In a second aspect, an embodiment of the present disclosure provides a speech recognition device, which includes:

[0032] An acquisition module, configured to acquire the recognition scores of each frame of speech in the speech to be recognized by the speech recognition model;

[0033] A determination module, configured to determine the hotword scores of each frame of speech in the speech to be recognized based on a pre-constructed hotword index model; wherein, the hotword score is used to represent whether each frame of speech contains a preset hotword;

[0034] A decoding module, configured to perform content decoding on the speech to be recognized based on the recognition scores and hotword scores of each frame of speech to obtain the content recognition result of the speech to be recognized.

[0035] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.

[0036] In a fourth aspect, an embodiment of the present disclosure provides an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the above method by executing the executable instructions.

[0037] The technical solution of the present disclosure has the following beneficial effects:

[0038] The above speech recognition method first obtains the recognition scores of each frame of speech in the speech to be recognized by the speech recognition model, as well as the hot word scores of each frame of speech in the speech to be recognized determined based on the pre-constructed hot word index model. Then, based on the recognition scores and hot word scores, content decoding is performed on the speech to be recognized, and a content recognition result closer to the speech to be recognized in line with the user's habits can be obtained, with higher accuracy and reliability of the recognition result. At the same time, in the embodiments of the present disclosure, the hot words and the language recognition model are independent of each other, and hot words can be added or updated without changing the speech recognition model, with higher flexibility, thereby solving the technical problem of weak flexibility in hot word updating during the current speech recognition process and achieving the technical effect of improving the flexibility of hot word updating.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0041] Figure 1 The flowchart of a speech recognition method in this exemplary embodiment is shown;

[0042] Figure 2 The flowchart of a speech recognition method in this exemplary embodiment is shown;

[0043] Figure 3 The schematic structural diagram of the hot word index model in a speech recognition method in this exemplary embodiment is shown;

[0044] Figure 4 The flowchart of a speech recognition method in this exemplary embodiment is shown;

[0045] Figure 5 The flowchart of a speech recognition method in this exemplary embodiment is shown;

[0046] Figure 6 The schematic structural diagram of the hot word index model in a speech recognition method in this exemplary embodiment is shown;

[0047] Figure 7 The flowchart of a speech recognition method in this exemplary embodiment is shown;

[0048] Figure 8The flowchart of a voice recognition method in this exemplary embodiment is shown;

[0049] Figure 9 The schematic structural diagram of a voice recognition device in this exemplary embodiment is shown;

[0050] Figure 10 The schematic structural diagram of an electronic device in this exemplary embodiment is shown. Detailed implementation manners

[0051] Now, the exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0052] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0053] The flowchart shown in the accompanying drawings is only an exemplary illustration and does not necessarily include all the steps. For example, some steps can be further decomposed, and some steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.

[0054] In related technologies, ASR (Automatic Speech Recognition) is a technology that converts human language into text and is widely used in various commercial products to improve the speed and convenience of people's communication, thereby enhancing the user experience. However, although some special words have only one pronunciation, they have different meanings in different scenarios. For example, the place names Mount Heng and Mount Hengshan, which may lead to inaccurate speech recognition results. Different users have different hot words. In speech recognition, the enhancement of hot words is mainly achieved by modifying the WFST (weighted finite-state transducer) static language model. However, this method requires adding hot words to the original language model to regenerate a new language model. This method has high accuracy but cannot update hot words in real time. Therefore, the flexibility of hot word update in the current speech recognition process is weak.

[0055] In view of the above problems, the embodiments of the present disclosure provide a speech recognition method that separates hot words from the language model (i.e., the speech recognition model in the embodiments of the present disclosure), enabling separate update of hot words, thereby improving the flexibility of hot word update in the current speech recognition process and enhancing the reliability of speech recognition. The following briefly introduces the application environment of the speech recognition method provided by the embodiments of the present disclosure:

[0056] The embodiments of the present disclosure are applied to a terminal device, which can be a local terminal device, such as any electronic device with a human-computer interaction interface like a mobile phone, a tablet, a computer, etc., or a client device in a cloud interaction system, such as a server, etc. The embodiments of the present disclosure do not make specific limitations. The speech recognition method provided by the embodiments of the present disclosure adopts an end-to-end recognition idea, that is, directly converting from speech to predicted text content, which is convenient, fast, and more efficient.

[0057] Taking the above terminal device as the execution subject and applying the speech recognition method to the above terminal device to perform text conversion on the speech to be recognized as an example for illustration. Please refer to Figure 1 The speech recognition method provided by the embodiments of the present disclosure includes the following steps 101 - step 103:

[0058] Step 101: The terminal device obtains the recognition scores of the speech recognition model for each frame of the speech to be recognized.

[0059] The speech recognition model refers to a neural network model used to convert the speech to be recognized into text content, such as a CTC model, an RNN Transducer model (RNN-T model for short), etc., which is not specifically limited in the embodiments of the present disclosure. The speech to be recognized refers to the speech generated by the user that needs to be recognized, and the recognition score refers to the output probability of the speech recognition model, or the value after a simple mathematical transformation of the output probability, such as taking its logarithm, etc., which is not specifically limited in the embodiments of the present disclosure.

[0060] Step 102: The terminal device determines a hot word score for each frame of speech to be recognized based on a pre-built hot word index model.

[0061] The preset hot word refers to a word that is more commonly used by users in a specific environment. For example, for "tou su", if the user is in the service industry or a general consumer, the hot word set for the voice may be "staying at a lodging", and if the user is a travel practitioner or a travel enthusiast, the hot word that may be set is "staying at a lodging". The hot word is set by the user independently, and is directly uploaded to the storage module of the terminal device, or sent to the terminal device through the user terminal directly used by the user, and the terminal device can obtain the hot word set by the user. The hot word can be a hot word set consisting of one or more hot words, and the embodiment of the present disclosure is not specifically limited. The hot word index model is constructed based on the hot words provided by the user. During the process of voice recognition, the hot word index model can be traversed to search one by one to determine whether the current frame voice contains the preset hot words. If the preset hot words are contained, the recognition result of the voice recognition model can be corrected according to the hot words. If the preset hot words are not contained, the recognition result is determined based on the output of the voice recognition model. Correspondingly, the hot word score is used to indicate whether each frame of speech contains a preset hot word. The corresponding relationship between the hot word score and the preset hot word can be specifically selected or set according to actual conditions, and no limitation is made here.

[0062] Step 103: The terminal device performs content decoding on the speech to be recognized based on the recognition score of each frame of speech and the hot word score to obtain a content recognition result of the speech to be recognized.

[0063] Content decoding refers to the process of converting the speech to be recognized into text according to the output result of the speech recognition model. In the embodiments of the present disclosure, the CTC Prefix Beam Search algorithm can be used for content decoding. Of course, other methods can also be used, which will not be enumerated here. As long as the purpose of obtaining the content recognition result of the speech to be recognized can be achieved. The recognition score is used to represent the output result of the speech recognition model, and the hot word score is used to represent whether there is a hot word in the speech to be recognized. A recognition result corresponding to the highest recognition score containing a preset hot word can be selected from the current recognition score through the hot word score as the target, that is, as the above-mentioned content recognition result.

[0064] The speech recognition method provided by the embodiments of the present disclosure first obtains the recognition scores of each frame of speech in the speech to be recognized by the speech recognition model, and the hot word scores of each frame of speech in the speech to be recognized determined based on the pre-constructed hot word index model, and then performs content decoding on the speech to be recognized based on the recognition score and the hot word score, so as to obtain a content recognition result that is closer to the speech to be recognized in line with the user's habits, and the accuracy and reliability of the recognition result are higher; at the same time, in the embodiments of the present disclosure, the hot word and the language recognition model are independent of each other, and hot words can be added or updated without changing the speech recognition model, with higher flexibility, thereby solving the technical problem of weak flexibility in hot word update in the current speech recognition process, and achieving the technical effect of improving the flexibility of hot word update.

[0065] Please refer to Figure 2 In an optional embodiment of the present disclosure, before step 102 above, where the terminal device determines the hot word scores of each frame of speech in the speech to be recognized based on the pre-constructed hot word index model, the speech recognition method further includes the following steps 201-202:

[0066] Step 201: The terminal device performs word segmentation processing on each obtained preset hot word to obtain a plurality of hot word segments.

[0067] The hot word segment is the smallest unit of the hot word index model. The hot word segment can be a Chinese character, an English word, or an English letter, etc., which is not specifically limited in the embodiments of the present disclosure. Word segmentation processing refers to splitting the original preset hot word into multiple hot word segments. For example, splitting the preset hot word "tea" into "t", "e", and "a", and so on.

[0068] Step 202: The terminal device uses a hot word segment as an index node to construct a hot word index model including multiple hot word index paths.

[0069] For example, please refer to Figure 3, a hot word index model with a prefix tree structure constructed based on the hot word segmentations "t", "A", "i", "o", "e", "a", "d", "n". A hot word segmentation serves as an index node, and each path starting from the bos root node is regarded as a hot word index path. For example, the first index path is the preset hot word "to", the second index path is the preset hot word "tea", the third index path is the preset hot word "ted", and so on. Figure 3 The hot words included also include the fourth index path "ten", the fifth index path "A", and the sixth index path "inn". It should be noted that a prefix tree is an ordered tree used to store associative arrays (i.e., the preset hot words in the embodiments of the present disclosure) for performing common phrase search prompts, etc.

[0070] In the embodiments of the present disclosure, after performing word segmentation processing on each obtained preset hot word to obtain multiple hot word segmentations, a hot word segmentation is used as an index node to construct a hot word index model containing multiple hot word index paths. When performing hot word search through the hot word index model constructed in this way, hot word search can be carried out orderly along each hot word index path, avoiding the occurrence of missing hot words, thereby improving the reliability and efficiency of hot word search and further improving the recognition efficiency of the speech recognition method provided by the embodiments of the present disclosure.

[0071] Please refer to Figure 4 , in an optional embodiment of the present disclosure, in step 102 above, the terminal device determines the hot word scores of each frame of speech in the speech to be recognized based on the pre-constructed hot word index model, including the following steps 401 - step 402:

[0072] Step 401, the terminal device sequentially performs hot word indexing on each frame of speech along multiple hot word index paths in the hot word index model, and determines the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentations of each index node in the hot word index model.

[0073] For example, please continue to refer to Figure 3, if the first-frame voice is "tea", and the corresponding hot-word segmentation is "t", "e", "a", and the bos root node stores the hot-word segmentation "t", "a", "i" of the next index node. The terminal device determines that the next index node contains the hot-word segmentation "t" in "tea", then it can search for the first hot-word segmentation "t" of the first-frame voice. During the search process, it is determined that the first hot-word segmentation "t" matches the first node "t" of the first row of index nodes in the hot-word index model, then the score of this index node is 1, and the node scores of the other two index nodes in the first row are 0. Those with a score of 0 will no longer continue to search this hot-word index path; and so on, the first row of index nodes "t" contains the hot-word segmentation "0" and "e" of the next index node, and the terminal device continues to search for the second hot-word segmentation "e" in the first-frame voice "tea". During the search process, it is determined that the second hot-word segmentation "e" matches the second node "e" of the second row of index nodes in the hot-word index model, then the score of this index node is increased by 1 point, and the score of this index node is 2, and the node score of the other index node in the second row is 0; and so on, the terminal device continues to search for the second hot-word segmentation "a" in the first-frame voice "tea", and the third hot-word segmentation "a" matches the first node "a" of the third row of index nodes in the hot-word index model, then the score of this index node is increased by 1 point, and the score of this index node is 3. The above is only an example, and other indexing methods and corresponding node scoring methods are not enumerated here. It should be noted that in this embodiment, the index node is a reward score, and the hot-word score under this index path is finally obtained through the superposition of each node. Of course, this reward score is only an example of the node score in this embodiment, and other situations, such as penalty-type scores, are not enumerated here, and can be specifically selected or set according to the actual situation.

[0074] Step 402: The terminal device determines the hot-word scores of each frame of voice according to the node scores of each frame of voice at each index node.

[0075] In step 401 above, the terminal device obtains the node scores of each index node, and then uses the node score of the last index node in the hot word index path of the finally indexed hot word as the hot word score of this frame of speech. For example, the node score of the last index node "a" of the first frame of speech "tea" above is 3, then the hot word score of the corresponding hot word index path is also 3, and the hot word score of this frame of speech is 3. The method of calculating the hot word score includes, but is not limited to, the above one. It can also score each index node separately, and then finally calculate the sum of the node scores to obtain the total score of different hot word index paths, and use this total score as the hot word score of this frame of speech. It should be noted that the hot word index path of this frame of speech includes, but is not limited to, one, and can be multiple. Correspondingly, there may be multiple hot word scores for this frame of speech.

[0076] The speech recognition method provided by the embodiments of the present disclosure first performs hot word indexing on each frame of speech along multiple hot word index paths in the hot word index model, and determines the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentation of each index node in the hot word index model. The hot word score of each frame of speech is determined according to the node scores of each frame of speech at each index node. The finally obtained hot word score includes each index node in this hot word index path. The hot word reliability determined by this hot word score is higher, which can further improve the accuracy and reliability of the speech recognition provided by the embodiments of the present disclosure.

[0077] Please refer to Figure 5 , in an optional embodiment of the present disclosure, step 402 above, the terminal device determines the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentation of each index node in the hot word index model, including the following steps 501-step 503:

[0078] Step 501, the terminal device performs word segmentation processing on the current frame of speech to obtain multiple frame segmentations.

[0079] The speech to be recognized is composed of multiple frames of speech. The speech content and length corresponding to one frame of speech are not the same. For the convenience of comparison in the hot word index model, the embodiments of the present disclosure first perform word segmentation processing on the current frame of speech, and convert the current frame of speech into different frame segmentations. The frame segmentation can be a character or the corresponding text. The embodiments of the present disclosure do not make specific limitations and can be determined according to the specific speech recognition steps.

[0080] Step 502, the terminal device indexes multiple frame segmentations based on the hot word index model. If the current frame segmentation matches the hot word segmentation of the current index node, the sum of the node score of the previous index node and the preset incremental score is determined as the reward score of the current index node.

[0081] Step 503: If the word segmentation of the current frame does not match the hot word segmentation of the current index node, the terminal device determines the cumulative reward score of the historical index node as the penalty score of the current index node.

[0082] Please refer to Figure 6 , corresponding to the above step 401, in the embodiment of the present disclosure, on the basis of the above step 401, the node score is further divided into a reward score (award) and a penalty score (backoff). Among them, the reward score of each index node can be calculated according to the following formulas (1) and (2):

[0083] A 1 = N (1)

[0084] A i = A i-1 * S (2)

[0085] In formulas (1) and (2), A i represents the reward score of the i-th index node, S represents the increment coefficient, which can be specifically set according to the actual situation, and A 1 represents that the initial reward score of the first index node is N, and N can be freely set according to the actual situation. For example, it can be 2.

[0086] The penalty score of each index node can be calculated according to the following formulas (3) and (4):

[0087] B 1 = N (3)

[0088] B i = B i-1 + A i (4)

[0089] In formulas (3) and (4), B i represents the penalty score of the i-th index node, A i represents the reward score of the i-th index node, and B 1 represents that the initial penalty score of the first index node is N, and N can be freely set according to the actual situation. For example, it can be 2.

[0090] That is to say, as long as the initial reward score and the matching degree between the word segmentation of each frame and the hot word segmentation in each index node are determined, the node score of the current frame of speech in each index node can be automatically calculated. Then, through the node score of the last index node, the hot word score in the hot word index path can be quickly obtained. It is simple and fast, which can greatly improve the determination efficiency and accuracy of the hot word score, and further improve the efficiency and reliability of the speech recognition provided by the embodiment of the present disclosure.

[0091] In an optional embodiment of the present disclosure, in step 503 above, if the current frame tokenization does not match the hot word tokenization of the current index node, the terminal device determines the cumulative reward score of the historical index node as the penalty score of the current index node, including the following step A:

[0092] Step A: If the current index node is the end node of the current hot word index path, the terminal device determines the penalty score of the current index node as the preset penalty score.

[0093] Please continue to refer to the above Figure 6 , each hot word index path corresponds to a preset hot word, and each preset hot word is configured with an end identifier at the last hot word. For example Figure 6 the end identifier in

[0094] Please refer to Figure 7 , in an optional embodiment of the present disclosure, in step 103 above, the terminal device decodes the content of the speech to be recognized based on the recognition score and the hot word score of each frame of speech, and obtains the content recognition result of the speech to be recognized, including the following steps 701-step 703:

[0095] Step 701: If the next frame tokenization matches the hot word tokenization of at least the next index node, the terminal device corrects the recognition score based on the reward score of the current frame tokenization to obtain the target score of the next frame tokenization at the next index node.

[0096] Among them, the next frame tokenization refers to the tokenization that is located after the current frame tokenization and adjacent to the current frame tokenization in the tokenization order; the next index node refers to the node that is located after the current index node and adjacent to the current index node along the index direction in the hot word index model. Each index node stores the hot word tokenization included in the next index node, so as to facilitate a preliminary judgment at the current index node. If the next node does not contain the hot word tokenization to be searched, then the search node can be ended at the current node, avoiding excessive calculations, saving computing resources, and improving the hot word search efficiency to a certain extent, and further improving the speech recognition efficiency of the embodiments of the present disclosure.

[0097] If the word segmentation of the next frame matches the hot word segmentation of at least the next index node, that is, the word segmentation of the next frame of the current frame voice can be found in the next index node at the next moment, the target score can be calculated according to the following formulas (5) and (6):

[0098]

[0099]

[0100] In formulas (5) and (6), represents the target score of the word segmentation of the next frame in the next index node, represents the target score of the word segmentation of the current frame in the current index node, represents the recognition score of the word segmentation of the next frame in the speech recognition model, a represents a word segmentation after preprocessing the frame word segmentation, *a represents the word segmentation sequence after preprocessing the frame word segmentation, *aa represents the padding frame word segmentation between two identical frame word segmentations a in the middle, *ab represents the padding frame word segmentation between the frame word segmentation a and the frame word segmentation b, and T represents the reward score of the next index node.

[0101] Step 702: If the word segmentation of the next frame does not match the hot word segmentation of all the next index nodes, the terminal device corrects the recognition score based on the penalty score of the current frame word segmentation to obtain the target score of the word segmentation of the next frame in the next index node.

[0102] If the word segmentation of the next frame does not match the hot word segmentation of at least the next index node, that is, the word segmentation of the next frame of the current frame voice cannot be found in the next index node at the next moment, the target score can be calculated according to the following formulas (7) and (8):

[0103]

[0104]

[0105] In formulas (7) and (8), represents the target score of the word segmentation of the next frame in the next index node, represents the target score of the word segmentation of the current frame in the current index node, represents the recognition score of the word segmentation of the next frame in the speech recognition model, a represents a word segmentation after preprocessing the frame word segmentation, *a represents the word segmentation sequence after preprocessing the frame word segmentation, *aa represents the padding frame word segmentation between two identical frame word segmentations a in the middle, *ab represents the padding frame word segmentation between the frame word segmentation a and the frame word segmentation b, and B represents the penalty score of the next index node.

[0106] Step 703: The terminal device decodes the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

[0107] After the terminal device obtains the target score according to the above step 701 or step 702, it can decode based on a pre-configured hotword decoder to obtain the content recognition result corresponding to the speech to be recognized. This recognition result is obtained by correcting the output result of the speech recognition model based on the hotwords pre-configured by the user, and is more in line with the actual application scenario of the user, with higher reliability. At the same time, the hotwords are independent of the speech recognition model, which is convenient for real-time updating or modification of the hotwords, with higher flexibility.

[0108] Please refer to Figure 8 , in an optional embodiment of the present disclosure, the above speech recognition method further includes the following steps 801 - step 802:

[0109] Step 801: The terminal device obtains the hotword to be updated.

[0110] The user can directly add the hotwords that need to be updated through the terminal device, or upload the hotwords that need to be updated through their corresponding user terminals. The embodiments of the present disclosure do not make specific limitations.

[0111] Step 802: If the hotword to be updated belongs to the training sample set of the speech recognition model, the terminal device updates the hotword index model based on the hotword to be updated.

[0112] This training sample set refers to the sample set used when constructing the speech recognition model or pre-training the speech recognition model. After the terminal device obtains the hotword to be updated, it first retrieves in the training sample set. If each word segment in the hotword to be updated belongs to the sample vocabulary in the training sample set, the terminal device expands on the existing hotword index model, further adding the hotword index path of the hotword to be updated to achieve the update of the hotword index model.

[0113] Furthermore, the terminal device can also determine the paths with an index frequency of 0 or lower than a certain threshold within a certain period based on the index frequencies of each hotword index path for deletion, thereby reducing the capacity of the hotword index model, saving resources, improving the hotword index efficiency, and further improving the speech recognition efficiency of the embodiments of the present disclosure.

[0114] To implement the above speech recognition method, an embodiment of the present disclosure provides a speech recognition device 900. Figure 9 The schematic architecture diagram of the speech recognition device 900 is shown. The speech recognition device 900 includes: an acquisition module 910, a determination module 920, and a decoding module 930, where:

[0115] The obtaining module 910 is configured to obtain the recognition scores of each frame of speech in the speech to be recognized by the speech recognition model;

[0116] The determining module 920 is configured to determine the hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model; wherein, the hot word score is used to represent whether each frame of speech contains a preset hot word;

[0117] The decoding module 930 is configured to perform content decoding on the speech to be recognized based on the recognition scores and hot word scores of each frame of speech, and obtain the content recognition result of the speech to be recognized.

[0118] In an optional embodiment, the determining module 920 is further configured to perform word segmentation processing on each obtained preset hot word to obtain a plurality of hot word segments; use a hot word segment as an index node to construct a hot word index model including multiple hot word index paths.

[0119] In an optional embodiment, the determining module 920 is specifically configured to sequentially perform hot word indexing on each frame of speech along multiple hot word index paths in the hot word index model, and determine the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segments of each index node in the hot word index model; determine the hot word scores of each frame of speech according to the node scores of each frame of speech at each index node.

[0120] In an optional embodiment, the determining module 920 is specifically configured to perform word segmentation processing on the current frame of speech to obtain a plurality of frame segments; perform indexing on the plurality of frame segments based on the hot word index model. If the current frame segment matches the hot word segment of the current index node, determine the sum of the node score of the previous index node and a preset increasing score as the reward score of the current index node; if the current frame segment does not match the hot word segment of the current index node, determine the cumulative reward score of the historical index node as the penalty score of the current index node.

[0121] In an optional embodiment, the determining module 920 is specifically configured to, if the current index node is the end node of the current hot word index path, determine the penalty score of the current index node as a preset penalty score.

[0122] In an alternative embodiment, the decoding module 930 is specifically configured to, if the next-frame word segmentation matches the hot-word segmentations of at least the next index node, correct the recognition score based on the reward score of the current-frame word segmentation to obtain the target score of the next-frame word segmentation at the next index node; wherein, the next-frame word segmentation refers to the word segmentation that is adjacent to and after the current-frame word segmentation in the word segmentation order; the next index node refers to the node that is adjacent to and after the current index node along the index direction in the hot-word index model; perform content decoding on the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

[0123] In an alternative embodiment, the decoding module 930 is specifically configured to, if the next-frame word segmentation does not match the hot-word segmentations of all the next index nodes, correct the recognition score based on the penalty score of the current-frame word segmentation to obtain the target score of the next-frame word segmentation at the next index node; perform content decoding on the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

[0124] In an alternative embodiment, the obtaining module 910 is further configured to obtain the hot-word to be updated; if the hot-word to be updated belongs to the training sample set of the speech recognition model, update the hot-word index model based on the hot-word to be updated.

[0125] The exemplary embodiments of the present disclosure also provide a computer-readable storage medium, which can be implemented in the form of a program product, including program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above. In one embodiment, the program product can be implemented as a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0126] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0128] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0129] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider). In the embodiments of the present disclosure, when the program code stored in the computer-readable storage medium is executed, any step in the above voice recognition method may be implemented.

[0130] Please refer to Figure 10 , the exemplary embodiments of the present disclosure also provide an electronic device 1000, which may be a background server of an information platform. The following will be described with reference to Figure 10 this electronic device 1000. It should be understood thatFigure 10 The illustrated electronic device 1000 is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present disclosure.

[0131] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include, but are not limited to: at least one processing unit 1010, at least one storage unit 1020, and a bus 1030 that connects different system components (including the storage unit 1020 and the processing unit 1010).

[0132] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section of this specification above. For example, the processing unit 1010 can execute the method steps as Figure 1 shown, etc.

[0133] The storage unit 1020 may include a volatile storage unit, such as a random access storage unit (RAM) 1021 and / or a cache storage unit 1022, and may further include a read-only storage unit (ROM) 1023.

[0134] The storage unit 1020 may also include a program / utilities 1024 having a set (at least one) of program modules 1025. Such program modules 1025 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0135] The bus 1030 may include a data bus, an address bus, and a control bus.

[0136] The electronic device 1000 may also communicate with one or more external devices 2000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and such communication may be carried out through an input / output (I / O) interface 1040. The electronic device 1000 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1050. As shown in the figure, the network adapter 1050 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0137] In the embodiments of the present disclosure, when the program code stored in the electronic device is executed, any step in the above voice recognition method can be implemented.

[0138] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0139] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here. After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0140] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only defined by the appended claims.

Claims

1. A speech recognition method, characterized in that, it includes: Obtaining the recognition scores of each frame of speech in the speech to be recognized by the speech recognition model; Determining the hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model; wherein, the hot word scores are used to characterize whether each frame of speech contains a preset hot word; Performing content decoding on the speech to be recognized based on the recognition scores and the hot word scores of each frame of speech to obtain the content recognition result of the speech to be recognized; The hot word index model contains index nodes, and each index node is configured with an end flag and a penalty score; the index nodes include hot word segmentations obtained by performing word segmentation on the preset hot words; the end flag and the penalty score are used to determine whether to stop hot word indexing; the penalty score is determined as the penalty score of the current index node by taking the cumulative reward score of the historical index node when the current frame segmentation does not match the hot word segmentation of the current index node, and the current frame segmentation is the frame segmentation obtained by performing word segmentation on the current frame of speech; if the current index node is an end node, then the penalty score of the current index node is determined as a preset penalty score to stop the hot word indexing based on the preset penalty score.

2. The speech recognition method according to claim 1, characterized in that, Before determining the hot word scores of each frame of speech in the speech to be recognized based on the pre-constructed hot word index model, the method further includes: Performing word segmentation on each obtained preset hot word to obtain a plurality of hot word segmentations; Taking one of the hot word segmentations as an index node to construct the hot word index model including multiple hot word index paths.

3. The speech recognition method according to claim 2, characterized in that, The determining the hot word scores of each frame of speech in the speech to be recognized based on the pre-constructed hot word index model includes: Performing hot word indexing on each frame of speech in sequence along the multiple hot word index paths in the hot word index model, and determining the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentations of each index node in the hot word index model; the node scores include the penalty score; Determining the hot word scores of each frame of speech according to the node scores of each frame of speech at each index node.

4. The speech recognition method according to claim 3, characterized in that, The node scores further include the reward score; the determining the node scores of each frame of speech at each index node according to the matching degree between each frame of speech and the hot word segmentations of each index node in the hot word index model includes: Performing word segmentation on the current frame of speech to obtain a plurality of frame segmentations; Indexing the multiple frame segmentations based on the hot word index model, if the current frame segmentation matches the hot word segmentation of the current index node, then determining the sum of the node score of the previous index node and a preset incremental score as the reward score of the current index node.

5. The speech recognition method according to claim 4, characterized in that, Performing content decoding on the speech to be recognized based on the recognition scores and hot word scores of each frame of speech to obtain a content recognition result of the speech to be recognized, including: If the next-frame word segmentation matches the hot word segmentations of at least the next index node, correcting the recognition score based on the reward score of the current-frame word segmentation to obtain a target score of the next-frame word segmentation at the next index node; wherein, the next-frame word segmentation refers to the word segmentation that is after the current-frame word segmentation and adjacent to the current-frame word segmentation in the word segmentation order; the next index node refers to the node that is after the current index node and adjacent to the current index node along the index direction in the hot word index model; Performing content decoding on the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

6. The speech recognition method according to claim 4, characterized in that Performing content decoding on the speech to be recognized based on the recognition scores and hot word scores of each frame of speech to obtain a content recognition result of the speech to be recognized, including: If the next-frame word segmentation does not match the hot word segmentations of all the next index nodes, correcting the recognition score based on the penalty score of the current-frame word segmentation to obtain a target score of the next-frame word segmentation at the next index node; Performing content decoding on the speech to be recognized based on the target score to obtain the content recognition result of the speech to be recognized.

7. The speech recognition method according to claim 1, characterized in that The method further includes: Obtaining a hot word to be updated; If the hot word to be updated belongs to the training sample set of the speech recognition model, updating the hot word index model based on the hot word to be updated.

8. A speech recognition device, characterized in that The device includes: An acquisition module, configured to acquire recognition scores of each frame of speech in the speech to be recognized by a speech recognition model; A determination module, configured to determine hot word scores of each frame of speech in the speech to be recognized based on a pre-constructed hot word index model; wherein, the hot word score is used to represent whether each frame of speech contains a preset hot word; A decoding module, configured to perform content decoding on the speech to be recognized based on the recognition scores and hot word scores of each frame of speech to obtain a content recognition result of the speech to be recognized; The hot word index model includes index nodes, and each index node is configured with an end flag and a penalty score; the index node includes a hot word segmentation obtained by performing word segmentation on the preset hot word; the end flag and the penalty score are used to determine whether to stop hot word indexing; the penalty score is determined as the penalty score of the current index node by taking the cumulative reward score of the historical index node when the current-frame word segmentation does not match the hot word segmentation of the current index node, and the current-frame word segmentation is a frame segmentation obtained by performing word segmentation on the current frame of speech; if the current index node is an end node, determining the penalty score of the current index node as a preset penalty score to stop the hot word indexing based on the preset penalty score.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the method according to any one of claims 1 to 7 by executing the executable instructions.