Speech Recognition Method, System, Electronic Device, and Storage Medium

By constructing hot word trie diagrams and real-time decoding technology, the inefficiency problem of traditional hot word strengthening methods is solved, efficient and accurate recognition of custom hot words is achieved, and the hot word recognition performance of the speech recognition model is improved.

CN116229967BActive Publication Date: 2025-07-11AISPEECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310145726.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-07-11
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

The traditional hot word strengthening method requires a large amount of manual collection of corpus and recompilation of decoding networks, resulting in high time consumption and the hot word paths in the end-to-end speech recognition model are difficult to fully include.

Method used

A hot word trie diagram is constructed, combined with an encoder and a decoder, decode the voice signal in real time, determine the potential hot word embedding through the hot word trie diagram, update the score of the candidate bundle search path, and select the highest scoring path as the recognition result.

Benefits of technology

It realizes efficient and accurate recognition of customized hot words, reduces the time for manual collection of corpus and network compilation, and improves the performance of hot words recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229967B_ABST
    Figure CN116229967B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a speech recognition method, system, electronic device, and storage medium. The method includes: constructing a hot word trie graph according to a custom hot word input by a user; in response to the input of a speech signal, sending the speech signal to a hot word-aware speech recognition model in real time; determining hidden layer features of the speech signal through an encoder; a decoder performing real-time decoding on the hidden layer features to obtain multiple candidate beam search paths, determining potential hot word embeddings of each candidate beam search path through the hot word trie graph, and the decoder updating the scores of the multiple candidate beam search paths based on the potential hot word embeddings until an end symbol is decoded, and selecting the candidate beam search path with the highest score to determine the recognition result of the speech signal. In the embodiment of the present invention, a hot word trie graph is constructed using a custom hot word, and during model decoding, the corresponding hot word embedding is determined through the hot word trie graph, and the potential hot word embedding is used to enhance the accurate recognition of the hot word by the hot word-aware speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice, and particularly to a voice recognition method, system, electronic device and storage medium. Background Art

[0002] With the rapid development of end-to-end voice technology, voice products based on a large amount of corpus have met daily usage. However, in practical applications, with the continuous development of social networks, new hot topics and hot words are constantly emerging. At the same time, users themselves also have different personalized words, such as names of people and places. These hot words usually have a very low frequency of occurrence in the original corpus for training the language model, so that the recognition effect is not good when recognizing user-defined words. Long speech real-time transcription and hot word enhancement of similar products have become an important requirement. Long speech refers to the scenario of continuously performing voice recognition, such as meeting transcription, audio and video subtitles, etc. Real-time transcription refers to streaming the recognition of voice results while speaking. In order to improve the recognition effect, hot word enhancement processing is usually performed. Among them, hot word enhancement means that the user gives a list of hot words, and the language model parameters are re-estimated by collecting hot word corpus, and a new decoding network is generated. Or at the time of decoding, the decoding path containing the hot word is excited to increase the cumulative historical path probability of the path where the hot word is located; improve the recognition effect of the hot word.

[0003] In the process of implementing the present invention, the inventor found that there are at least the following problems in the related art:

[0004] Traditional hot word enhancement usually enhances hot words at the language model layer. Collecting corpus requires a lot of manual work and is not timely, and recompiling the decoding network consumes a lot of time.

[0005] Based on the end-to-end speech recognition model, the acoustic model and the language model are coupled together to obtain the best fusion performance. The size of the beam is usually much smaller than the traditional WFST (Weighted Finite State Transducer) method, and the path containing the hot word rarely exists in the beam. Summary of the Invention

[0006] In order to at least solve the problems in the prior art that collecting corpus requires a lot of manual work and recompiling the decoding network consumes a lot of time, and the beam search path is difficult to contain hot words. In a first aspect, an embodiment of the present invention provides a voice recognition method, including:

[0007] Construct a hot word trie graph according to the user input custom hot words, wherein the leaf nodes of the hot word trie graph include the hot word lists corresponding to each character in the custom hot words;

[0008] In response to the input of a voice signal, the voice signal is sent to a hot-word aware speech recognition model in real time, where the hot-word aware speech recognition model includes: an encoder and a decoder based on the hot-word trie graph;

[0009] Determine the hidden layer features of the voice signal through the encoder;

[0010] The decoder decodes the hidden layer features in real time to obtain multiple candidate beam search paths, determines the potential hot-word embeddings of each candidate beam search path through the hot-word trie graph, and the decoder updates the scores of the multiple candidate beam search paths based on the potential hot-word embeddings until an end symbol is decoded, and selects the candidate beam search path with the highest score to determine the recognition result of the voice signal.

[0011] In a second aspect, an embodiment of the present invention provides a voice recognition system, including:

[0012] A graph determination program module for constructing a hot-word trie graph according to a custom hot word input by a user, where the leaf nodes of the hot-word trie graph include a hot-word list corresponding to each character in the custom hot word;

[0013] A data transmission program module for, in response to the input of a voice signal, sending the voice signal to a hot-word aware speech recognition model in real time, where the hot-word aware speech recognition model includes: an encoder and a decoder based on the hot-word trie graph;

[0014] An encoding program module for determining the hidden layer features of the voice signal through the encoder;

[0015] A recognition program module for the decoder to decode the hidden layer features in real time to obtain multiple candidate beam search paths, determine the potential hot-word embeddings of each candidate beam search path through the hot-word trie graph, and the decoder updates the scores of the multiple candidate beam search paths based on the potential hot-word embeddings until an end symbol is decoded, and selects the candidate beam search path with the highest score to determine the recognition result of the voice signal.

[0016] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the voice recognition method according to any embodiment of the present invention.

[0017] Fourthly, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, the steps of the speech recognition method according to any embodiment of the present invention are implemented.

[0018] The beneficial effects of the embodiments of the present invention are as follows: a custom hot word is used to construct a hot word trie graph. When decoding each step of the hot word-aware speech recognition model, the corresponding hot word embedding is determined through the hot word trie graph, and the potential hot word embedding is used to enhance the accurate recognition of hot words by the hot word-aware speech recognition model. With the update of the custom hot word, only the hot word trie graph needs to be updated, and efficient and accurate recognition of the custom hot word can be achieved during recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of a speech recognition method provided by an embodiment of the present invention;

[0021] Figure 2 is a schematic diagram of the construction of a hot word trie graph of a speech recognition method provided by an embodiment of the present invention;

[0022] Figure 3 is a model structure diagram of a speech recognition method provided by an embodiment of the present invention;

[0023] Figure 4 is a test data diagram of a speech recognition method provided by an embodiment of the present invention;

[0024] Figure 5 is a schematic structural diagram of a speech recognition system provided by an embodiment of the present invention;

[0025] Figure 6 is a schematic structural diagram of an embodiment of an electronic device for speech recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] As Figure 1 shown is a flowchart of a speech recognition method provided by an embodiment of the present invention, including the following steps:

[0028] S11: Construct a hotword trie graph according to the custom hotwords input by the user. Among them, the leaf nodes of the hotword trie graph include the hotword lists corresponding to each character in the custom hotwords;

[0029] S12: In response to the input of the voice signal, send the voice signal to the hotword-aware speech recognition model in real time. Among them, the hotword-aware speech recognition model includes: an encoder and a decoder based on the hotword trie graph;

[0030] S13: Determine the hidden layer features of the voice signal through the encoder;

[0031] S14: The decoder decodes the hidden layer features in real time to obtain multiple candidate beam search paths, determines the potential hotword embeddings of each candidate beam search path through the hotword trie graph, and the decoder updates the scores of the multiple candidate beam search paths based on the potential hotword embeddings until the end symbol is decoded, and selects the candidate beam search path with the highest score to determine the recognition result of the voice signal.

[0032] In this embodiment, considering that the end-to-end speech recognition couples the acoustic model and the language model at the model layer and it is difficult to optimize them separately, this method optimizes both at the model layer and the decoding layer to achieve the best performance of hotword recognition. This method can be adapted to intelligent devices. For example, taking a smart speaker as an example, users interact with the smart speaker at any time in daily life.

[0033] For step S11, in daily interactions with users, custom hotwords input by users will be collected. For example, custom hotwords specifically input by users; words that the smart speaker cannot recognize during the interaction with users can also be collected, the explanations of these words can be queried through the Internet, and they can be determined as custom hotwords. Custom hotwords can also be set to be collected specifically by the cloud, and the smart speaker regularly downloads these custom hotwords from the cloud to obtain custom hotwords.

[0034] After obtaining the custom hot words, use the custom hot words to construct a hot word trie graph. As an implementation, constructing the hot word trie graph according to the custom hot words input by the user includes:

[0035] Decompose the custom hot words to obtain multiple hot characters;

[0036] Convert the multiple hot characters into an audio sequence, construct the nodes of the hot word trie graph through the audio sequence, and determine the hot word list corresponding to the hot characters in the leaf nodes.

[0037] In this real-time method, for a given custom hot word, first convert it into a pinyin sequence, and then construct a Trie graph according to the pinyin sequence. For example, the hot words are: iFlytek, Bi Dao, Chisheng Information, Bichi Technology. The corresponding pinyin sequences are: iFlytek: si1 bi4 chi2; Bi Dao: bi4 dao3; Chisheng Information: chi2 sheng1 xin4 xi1; Bichi Technology: bi4 chi2 ke1 ji4. As Figure 2 Shown is the constructed hot word trie graph, where the edge represents the matchable pinyin, and the dotted node represents the matched hot word. The characters in the brackets in the node represent potential hot characters. For example, the hot word list of the node in the lower left corner of the figure is: chi: [#hot word#, sheng, ke]. As the user uses it, or with the update of the latest hot words in the cloud, the updated custom hot words can be obtained periodically to further optimize the hot word trie graph.

[0038] For step S12, after pre-constructing the hot word trie graph in step S11, with the hot word trie graph, hot word-aware speech recognition can be performed. Collect the voice signal input by the user through VAD (Voice Activity Detection). Specifically, the received voice signal can be sampled as discrete energy values and stored in the data buffer, and the buffered voice signal is gradually and real-time input into the hot word-aware speech recognition model, such as Figure 3 Shown is the structural schematic diagram of the hot word-aware speech recognition model. For example, the voice input by the user is "Welcome Bi Dao to visit your company Bichi Technology".

[0039] For step S13, use the encoder to extract the hidden layer features by frame for the buffered data. Usually, the frame length is 25 ms and the frame shift is 10 ms. After performing the discrete Fourier transform on the data and windowing, 40-dimensional FBANK features are obtained. That is, each frame is 40-dimensional features, and there are 100 frames per second. Store the extracted hidden layer features in the feature buffer for the decoder to decode.

[0040] For step S14, the decoder is used to decode the hidden layer features extracted in step S13 to obtain multiple candidate beam search paths. Each candidate beam search path represents a different recognition prediction result, which is composed of the score probabilities of different predicted words. The potential hot word embeddings of each candidate beam search path are determined through the hot word trie graph. Specifically, by walking the candidate beam search path in the hot word trie graph, a list of potential hot words that need to be strengthened for each word in the speech signal can be obtained. If there is no walkable edge, walk null until a walkable edge is found. If retreating to a node where there is no null edge, return to the root node.

[0041] For example, the list of potential hot words corresponding to each word in "Welcome, Mr. Bi, to your company, Bichi Technology" is: Huan: [Si, Bi, Bi, Chi]; Ying: [Si, Bi, Bi, Chi]; Bi: [Dao, Chi]; Dao: [#Hot Word#]; Guang: [Si, Bi, Bi, Chi]; Lin: [Si, Bi, Bi, Chi]; Gui: [Si, Bi, Bi, Chi]; Si: [Bi]; Bi: [Chi, Dao, Chi]; Chi: [#Hot Word#, Sheng, Ke]; Ke: [Ji]; Ji: [#Hot Word#]. Convert the potential hot words into a zero-one vector of character count dimension, where the characters existing in the hot word list are 1, otherwise 0. The obtained zero-one vector of character count dimension is [0,...1,...0], that is, the corresponding potential hot word embedding is obtained.

[0042] At each step of decoding, for each candidate path in the beam search, walk one step in the hot word trie graph with the character decoded in the previous step to obtain the potential hot word embedding. The potential hot word embedding and the character decoded in the previous step are jointly input into the decoding network to obtain a hidden layer feature as a query vector (query), and an attention operation is performed with the hidden layer output of the encoder as the key and value to obtain an output vector, and the score of the next candidate hot word is recalculated. i continues to decode until the EOS (end of sentence, end symbol. For example, when the VAD detection ends, or when it is detected that the user's sentence ends, the EOS end symbol is generated) is decoded. Then, the scores of each candidate beam search path are comprehensively determined using the scores of each candidate hot word, and the candidate beam search path with the highest score is selected to determine the recognition result of the speech signal. The recognition result updated with the hot word is "Welcome, Mr. Bi, to your company, Bichi Technology".

[0043] As an implementation, the hot word-aware speech recognition model is obtained by training with a training text annotated with hot words, including:

[0044] Using a semantic analyzer to parse the training text and extract multiple nouns contained in multiple training sentences in the training text;

[0045] Randomly select any noun in each training sentence as a hot word, and use the hot word to construct a hot word trie graph;

[0046] After adding hot word tags to the hot words in each training sentence, input them into the hot word-aware speech recognition model, where the hot word tags are used to strengthen the recognition of the hot words by the hot word-aware speech recognition model;

[0047] Determine the hidden layer features of the training sentence through the encoder;

[0048] The decoder performs real-time decoding on the hidden layer features to obtain multiple candidate beam search paths, input the candidate beam search paths into the hot word trie graph, query the corresponding hot word list, and determine the potential hot word embeddings corresponding to each candidate character in the candidate beam search paths through the hot word list;

[0049] The decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings, and selects the candidate beam search path with the highest score to determine the predicted recognition result of the training sentence;

[0050] Train the hot word-aware speech recognition model based on the predicted recognition result and the loss function of the hot words with added hot word tags until the hot words with added hot word tags are recognized in the predicted recognition result.

[0051] In this embodiment, the hot word-aware speech recognition model is obtained by training with training texts labeled with hot words. Specifically, a crawling algorithm can be used to crawl the speech content of different users from the forum as the training text. After obtaining the training text, use a semantic analyzer to parse the training text and extract multiple nouns contained in multiple training sentences in the training text; for each training sentence, randomly select one of the nouns as a hot word, and add the #hot word# tag after its annotation to strengthen the network's ability to recognize hot words through this tag, and use the randomly selected hot words to construct a hot word trie graph.

[0052] Input the training sentences with hot word tags into the hot word-aware speech recognition model, and utilize the encoder and decoder in the hot word-aware speech recognition model. The encoding and decoding processes are as shown in steps S13 - S14, which will not be elaborated here. Among them, the hot words marked during the training process can additionally stimulate the characters in the potential hot word list to obtain better hot word recognition effects. That is, when calculating the score of the next candidate character, the characters in the potential hot word list obtain additional scores. At this time, the decoder needs to maintain the total additional score obtained currently, and deduct the additional score when walking the null edge or returning to the root node and the current node is not a dotted node, so as to obtain the predicted recognition results of each training sentence. Finally, train the hot word-aware speech recognition model based on the predicted recognition results and the loss function of the hot words with hot word tags until the hot words with hot word tags are recognized in the predicted recognition results.

[0053] It can be seen from this embodiment that by using the custom hot words to construct the hot word trie graph, at each step of decoding of the hot word-aware speech recognition model, the corresponding hot word embedding is determined through the hot word trie graph, and the potential hot word embedding is used to strengthen the accurate recognition of hot words by the hot word-aware speech recognition model. With the update of the custom hot words, only the hot word trie graph needs to be updated, and efficient and accurate recognition of the custom hot words can be achieved during recognition.

[0054] As Figure 4 shown, it is the experimental test result of this method. It can be seen that this method does not require additional collection of training corpus. Only a hot word list needs to be provided, which greatly improves the recognition performance of the recognition model for hot words while having little impact on the performance of general speech recognition. The user experience of speech recognition is further improved.

[0055] As Figure 5 shown is the structural schematic diagram of a speech recognition system provided by an embodiment of the present invention. This system can execute the speech recognition method described in any of the above embodiments and is configured in a terminal.

[0056] A speech recognition system 10 provided in this embodiment includes: a graph determination program module 11, a data transmission program module 12, an encoding program module 13, and a recognition program module 14.

[0057] Among them, the graph determination program module 11 is used to construct a hot word trie graph according to the custom hot words input by the user. Among them, the leaf nodes of the hot word trie graph include the hot word lists corresponding to each character in the custom hot words; the data transmission program module 12 is used to respond to the input of the voice signal and send the voice signal to the hot word-aware speech recognition model in real time. Among them, the hot word-aware speech recognition model includes: an encoder and a decoder based on the hot word trie graph; the encoding program module 13 is used to determine the hidden layer features of the voice signal through the encoder; the recognition program module 14 is used for the decoder to perform real-time decoding on the hidden layer features to obtain multiple candidate beam search paths, determine the potential hot word embeddings of each candidate beam search path through the hot word trie graph, and the decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings until the end symbol is decoded, and selects the candidate beam search path with the highest score to determine the recognition result of the voice signal.

[0058] An embodiment of the present invention also provides a non-volatile computer storage medium, and the computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the speech recognition method in any of the above method embodiments;

[0059] As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:

[0060] Construct a hot word trie graph according to the custom hot words input by the user. Among them, the leaf nodes of the hot word trie graph include the hot word lists corresponding to each character in the custom hot words;

[0061] Respond to the input of the voice signal and send the voice signal to the hot word-aware speech recognition model in real time. Among them, the hot word-aware speech recognition model includes: an encoder and a decoder based on the hot word trie graph;

[0062] Determine the hidden layer features of the voice signal through the encoder;

[0063] The decoder performs real-time decoding on the hidden layer features to obtain multiple candidate beam search paths, determines the potential hot word embeddings of each candidate beam search path through the hot word trie graph, and the decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings until the end symbol is decoded, and selects the candidate beam search path with the highest score to determine the recognition result of the voice signal.

[0064] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium and, when executed by a processor, perform the speech recognition method in any of the above method embodiments.

[0065] Figure 6 It is a schematic diagram of the hardware structure of an electronic device for the speech recognition method provided in another embodiment of the present application. As Figure 6 shown, the device includes:

[0066] One or more processors 610 and a memory 620. Figure 6 Taking one processor 610 as an example. The device for the speech recognition method may further include: an input device 630 and an output device 640.

[0067] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected through a bus or other means. Figure 6 Taking the connection through a bus as an example.

[0068] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the speech recognition method in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the speech recognition method in the above method embodiments.

[0069] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data, etc. In addition, the memory 620 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the mobile device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0070] The input device 630 can receive input digital or character information. The output device 640 may include a display device such as a display screen.

[0071] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, perform the speech recognition method in any of the above method embodiments.

[0072] The above product can execute the method provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present application.

[0073] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0074] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the speech recognition method according to any embodiment of the present invention.

[0075] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:

[0076] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0077] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as tablet computers.

[0078] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0079] (4) Other electronic devices with data processing functions.

[0080] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the said element.

[0081] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice recognition method, comprising: Constructing a hot word trie graph according to custom hot words input by a user, wherein leaf nodes of the hot word trie graph include hot word lists corresponding to each character in the custom hot words; In response to the input of a voice signal, sending the voice signal to a hot word-aware speech recognition model in real time, wherein the hot word-aware speech recognition model includes: an encoder and a decoder based on the hot word trie graph; Determining hidden layer features of the voice signal through the encoder; The decoder decodes the hidden layer features in real time to obtain multiple candidate beam search paths, and determines potential hot word embeddings of each candidate beam search path through the hot word trie graph. Walking the candidate beam search paths in the hot word trie graph to obtain a list of potentially reinforced hot words for each character in the voice signal. The decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings, including: at each step of decoding during walking, for each candidate path in the beam search, walking one step in the hot word trie graph with the character decoded in the previous step to obtain a potential hot word embedding, and using the potential hot word embedding and the character decoded in the previous step to obtain a hidden layer feature as a query vector to recalculate the score of the next candidate character until the end symbol is decoded, and selecting the candidate beam search path with the highest score to determine the recognition result of the voice signal.

2. The method according to claim 1, wherein The decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings, including: Decoding the potential hot word embedding to obtain a hot word query vector; Performing attention mechanism processing on the hot word query vector and the hidden layer features determined by the encoder as keys and values to determine the scores of each candidate character in the multiple candidate beam search paths; Determining the scores of the multiple candidate beam search paths according to the scores of each candidate character.

3. The method according to claim 1, wherein The constructing a hot word trie graph according to custom hot words input by a user includes: Decomposing the custom hot words to obtain multiple hot words; Converting the multiple hot words into an audio sequence, constructing nodes of the hot word trie graph through the audio sequence, and determining hot word lists corresponding to the hot words in the leaf nodes.

4. The method according to claim 1, wherein, The hot word-aware speech recognition model is obtained by training with training texts labeled with hot words, including: Using a semantic analyzer to parse the training texts, and extracting multiple nouns contained in multiple training sentences in the training texts; Randomly selecting any noun in each training sentence as a hot word, and constructing a hot word trie graph using the hot word; Adding a hot word tag to the hot word in each training sentence and inputting it into the hot word-aware speech recognition model, wherein the hot word tag is used to enhance the recognition of the hot word by the hot word-aware speech recognition model; Determining hidden layer features of the training sentence through the encoder; The decoder decodes the hidden layer features in real time to obtain multiple candidate beam search paths, inputs the candidate beam search paths into the hot word trie graph, queries the corresponding hot word list, and determines the potential hot word embeddings corresponding to the candidate characters in the candidate beam search paths through the hot word list; The decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings, and selects the candidate beam search path with the highest score to determine the predicted recognition result of the training statement; Based on the predicted recognition result and the loss function of the hot word with the hot word mark added, the hot word-aware speech recognition model is trained until the hot word with the hot word mark added is recognized in the predicted recognition result.

5. A speech recognition system, comprising: A graph determination program module for constructing a hot word trie graph according to a custom hot word input by a user, wherein the leaf nodes of the hot word trie graph include a hot word list corresponding to each character in the custom hot word; A data transmission program module for, in response to the input of a voice signal, sending the voice signal to the hot word-aware speech recognition model in real time, wherein the hot word-aware speech recognition model includes: an encoder and a decoder based on the hot word trie graph; An encoding program module for determining the hidden layer features of the voice signal through the encoder; A recognition program module for the decoder to decode the hidden layer features in real time to obtain multiple candidate beam search paths, and determine the potential hot word embeddings of each candidate beam search path through the hot word trie graph, wherein the candidate beam search paths are walked in the hot word trie graph to obtain a list of potentially reinforced hot words for each character in the voice signal, and the decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings, including: at each step of decoding during walking, for each candidate path in the beam search, walk one step of the character decoded in the previous step in the hot word trie graph to obtain a potential hot word embedding, and use the potential hot word embedding and the character decoded in the previous step to obtain the hidden layer features as a query vector to recalculate the score of the next candidate character until the end symbol is decoded, and select the candidate beam search path with the highest score to determine the recognition result of the voice signal.

6. The system according to claim 5, wherein, The recognition program module is used for: Decoding the potential hot word embedding to obtain a hot word query vector; Performing attention mechanism processing on the hot word query vector and the hidden layer features determined by the encoder as keys and values to determine the scores of the candidate characters in the multiple candidate beam search paths; Determining the scores of the multiple candidate beam search paths according to the scores of the candidate characters; 7. The system according to claim 5, wherein The graph determination program module is used for: Decomposing the custom hot word to obtain multiple hot words; Converting the multiple hot words into an audio sequence, constructing the nodes of the hot word trie graph through the audio sequence, and determining the hot word list corresponding to the hot words in the leaf nodes; 8. The system according to claim 5, wherein The system further includes a training program module for: Parse the training text using a semantic analyzer to extract multiple nouns contained in multiple training statements in the training text; Randomly select any noun in each training statement as a hot word, and use the hot word to construct a hot word trie graph; Add a hot word mark to the hot word in each training statement and input it into the hot word-aware speech recognition model, where the hot word mark is used to strengthen the recognition of the hot word by the hot word-aware speech recognition model; Determine the hidden layer features of the training statement through an encoder; The decoder performs real-time decoding on the hidden layer features to obtain multiple candidate beam search paths, inputs the candidate beam search paths into the hot word trie graph, queries the corresponding hot word list, and determines the potential hot word embeddings corresponding to each candidate character in the candidate beam search paths through the hot word list; The decoder updates the scores of the multiple candidate beam search paths based on the potential hot word embeddings, and selects the candidate beam search path with the highest score to determine the predicted recognition result of the training statement; Train the hot word-aware speech recognition model based on the predicted recognition result and the loss function of the hot word with the hot word mark added until the hot word with the hot word mark added is recognized in the predicted recognition result.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method according to any one of claims 1-4.

10. A storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Speech recognition method, device and apparatus

    CN109523991A

  • Putonghua and Cantonese hybrid speech recognition model training method and system

    CN111816160A