A keyword recognition method and related device

The method enhances keyword recognition by using frame-based decoding and lattice graph adjustments to improve accuracy and reduce latency in noisy environments.

CN115762521BActive Publication Date: 2025-07-15ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211238834.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-07-15
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

Existing keyword recognition systems face challenges in accurately identifying keywords in noisy environments and suffer from delayed response times due to voice activity detection (VAD) inaccuracies and increased computational load from multiple neural networks, affecting response speed and accuracy.

Method used

A method that involves real-time frame decoding of audio signals, determining keyword boundaries based on initial N-second segments, and adjusting candidate word sequence scores to enhance accuracy and speed by using a lattice graph to identify the most probable keyword sequence.

Benefits of technology

This approach reduces latency and improves keyword recognition accuracy by using frame-based decoding and lattice graph adjustments, ensuring faster and more reliable keyword identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762521B_ABST
    Figure CN115762521B_ABST
Patent Text Reader

Abstract

The present application discloses a keyword recognition method and related device, which relates to the technical field of speech recognition. In the present application, the keyword recognition system adopts a streaming processing mode to receive and decode the speech signal emitted by the target object. Among them, when the keyword recognition system recognizes a preset keyword during the first decoding process, the keyword recognition system will perform secondary decoding on the first decoding result that is no more than N seconds and contains the keyword based on the speech word graph obtained during the first decoding process. By adopting this method, the cut-off point of the keyword speech is determined according to the decoding result obtained from the first decoding, which reduces the latency of keyword recognition. Secondly, by only obtaining the decoding result with a duration of N seconds forward from the current moment, the decoding duration is reduced, and the speed of secondary decoding is accelerated. At the same time, secondary decoding is performed on the first decoding result and confidence judgment is performed on the second decoding result, which improves the accuracy of keyword recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and particularly to a keyword recognition method and related devices. Background Art

[0002] As an important branch of speech recognition technology, keyword recognition is widely used in human-computer interaction and smart home. For example, in some smart home scenarios, users will use wake-up words to wake up smart devices, and then speak out the voice commands they hope the devices to execute, instructing the devices to complete specific actions, such as "turn on the air conditioner". In this process, both voice wake-up and command word recognition involved are important applications of keyword recognition technology.

[0003] Keyword speech recognition systems usually combine streaming processing and a Voice Activity Detection (VAD) module. By using VAD to judge the start and end points of speech, better recognition results can be obtained. The drawback of this method is that the endpoint detection effect of VAD will affect the performance of speech recognition. For example, in some noisy environments, VAD can hardly accurately judge the start and end points of speech, which will lead to a sharp decline in keyword recognition effect. In addition, VAD may introduce additional latency when judging the end point of speech, which will reduce the response speed of the keyword recognition system. Especially when decoding a long speech with keywords in the middle of a sentence, if waiting for VAD to judge the end of speech before decoding, the response speed of the entire system will be greatly reduced.

[0004] In response to the above problems, there are currently two effective solutions as follows:

[0005] Method 1: Recognize speech in a streaming processing manner, and at the same time judge whether to output the corresponding recognition result according to the confidence of the recognition result during the streaming process, without waiting for the end of the voice command.

[0006] However, in this method, judging whether to output the recognition result in real time according to the confidence, for a speech with correct pronunciation, the confidence may be relatively low in the initial pronunciation stage. At this time, this method may discard the recognition result with low confidence, reducing the accuracy of keyword recognition.

[0007] Method 2: Use a first neural network to extract multi-dimensional acoustic features of speech, use a second neural network to extract first acoustic information from the multi-dimensional acoustic features, then use an attention mechanism to obtain second acoustic information from the first acoustic information, and finally obtain the probability distribution of phonemes based on the two acoustic information and perform decoding according to the probability distribution to obtain the recognition result.

[0008] However, the forward inference of multiple neural network acoustic models and the fusion process of posterior probabilities in this method will greatly increase the computational complexity of the decoding process, reduce the decoding speed of this algorithm, and thus reduce the response speed of keyword recognition.

[0009] In view of this, a new keyword recognition method needs to be proposed to improve the response speed and accuracy of keyword recognition. Summary of the Invention

[0010] This application provides a keyword recognition method and related device to improve the response speed and accuracy of keyword recognition.

[0011] In a first aspect, an embodiment of this application provides a keyword recognition method, and the method includes:

[0012] Perform real-time frame-by-frame decoding on the acquired speech signal to be recognized. When the duration of the first decoded speech data is not less than N seconds, use the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result;

[0013] If the first decoding result contains a preset keyword, continue to perform frame-by-frame decoding on the speech to be recognized until non-keyword appears in the second decoded speech data. Then, obtain the speech word graph generated during the decoding of the second speech data within N seconds before the second current moment. The speech word graph includes: multiple candidate character sequences whose similarity to the keyword is higher than a set threshold;

[0014] According to a preset ratio, adjust the proportion of the evaluation index of each candidate character included in each candidate character sequence in the speech word graph respectively, and based on the adjustment result, obtain the target evaluation value of each candidate character sequence respectively;

[0015] Based on the obtained target evaluation values, decode the speech word graph to obtain a second decoding result. When the second decoding result contains the keyword, output the second decoding result as the recognition result of the speech signal to be recognized.

[0016] In a second aspect, an embodiment of this application further provides a keyword recognition device, and the device includes:

[0017] A first decoding module, configured to perform real-time frame-by-frame decoding on the acquired speech signal to be recognized. When the duration of the first decoded speech data is not less than N seconds, use the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result;

[0018] The second decoding module, if the preset keyword is included in the first decoding result, continues to perform frame-by-frame decoding on the speech to be recognized until a non-keyword appears in the decoded second speech data, and obtains the speech word graph generated during the decoding process of the second speech data within N seconds forward from the second current moment. The speech word graph includes: a plurality of candidate character sequences whose similarity to the keyword is higher than the set threshold;

[0019] The adjustment module adjusts the proportion of the evaluation index of each candidate character included in each candidate character sequence in the speech word graph according to the preset ratio, and based on the adjustment result, obtains the target evaluation value of each candidate character sequence respectively;

[0020] The output module decodes the speech word graph based on the obtained target evaluation values to obtain the second decoding result, and when the keyword is included in the second decoding result, outputs the second decoding result as the recognition result of the speech signal to be recognized.

[0021] Optionally, when taking the data part obtained within N seconds forward from the first current moment in the first speech data as the first decoding result, the first decoding module is further configured to:

[0022] If the duration of the decoded first speech data is less than N seconds and includes the preset keyword, directly use the first speech data as the first decoding result.

[0023] Optionally, the first decoding module is further configured to:

[0024] If the preset keyword is not included in the first decoding result, discard the decoded first speech data and re-perform frame-by-frame decoding on the subsequent real-time collected speech signal to be recognized.

[0025] Optionally, when adjusting the proportion of the evaluation index of each candidate character included in each candidate character sequence according to the preset ratio and obtaining the target evaluation value of each candidate character sequence based on the adjustment result, the adjustment module is configured to:

[0026] For each candidate character sequence, perform the following operations respectively:

[0027] For each candidate character included in a candidate character sequence, perform the following operations respectively: based on the preset ratio, adjust the weights of the evaluation indexes associated with a candidate character, and obtain the character evaluation value of a candidate character based on the evaluation indexes and the corresponding weights;

[0028] Based on the character evaluation values of the candidate characters, obtain the target evaluation value of a candidate character sequence.

[0029] Optionally, if the keyword included in the second decoding result is composed of a candidate character sequence in the speech word graph, then before outputting the second decoding result as a segment of recognition result, the output module is further configured to:

[0030] Obtain the confidence of the second decoding result based on the target evaluation value corresponding to the candidate character sequence that composes the keyword;

[0031] If the confidence of the second decoding result is greater than a preset confidence threshold, then output the second decoding result as a segment of recognition result.

[0032] Optionally, the output module is further configured to:

[0033] If the confidence of the second decoding result is not greater than the preset confidence threshold, then discard the second decoding result and re-perform frame-by-frame decoding on the subsequent to-be-recognized speech signals collected in real time.

[0034] Optionally, when performing frame-by-frame decoding on the to-be-recognized speech signal, the first decoding module and the second decoding module are configured to:

[0035] For each speech frame included in the to-be-recognized speech signal, respectively perform the following operations:

[0036] Extract the acoustic features of a speech frame, send the acoustic features into a pre-trained acoustic model, and obtain the phoneme probability distribution of a speech frame;

[0037] Based on the phoneme probability distribution, obtain the decoding result of a speech frame.

[0038] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any item of the first aspect is implemented.

[0039] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first aspect are implemented.

[0040] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is called by a computer, the computer is enabled to execute the method described in the first aspect.

[0041] In the embodiment of the present application, the keyword recognition system adopts a streaming processing method to receive and decode the voice signal emitted by the target object. Among them, when the keyword recognition system recognizes a preset keyword during the first decoding process, or when the decoded voice duration exceeds N seconds and contains a keyword, the keyword recognition system will perform secondary decoding on the first decoding result that is no more than N seconds and contains a keyword based on the voice word graph obtained during the first decoding process.

[0042] By adopting this method, the cut-off point of the keyword voice is determined according to the decoding result obtained from the first decoding, reducing the latency of keyword recognition. Secondly, by only obtaining the decoding result of the duration of N seconds forward from the current moment, the decoding duration is reduced, accelerating the speed of secondary decoding. At the same time, secondary decoding of the first decoding result and confidence judgment of the second decoding result improve the accuracy of keyword recognition. Brief Description of the Drawings

[0043] Figure 1 It is a schematic diagram of the system architecture in the embodiment of the present application;

[0044] Figure 2 It is a detailed flowchart of keyword recognition under the system architecture in the embodiment of the present application;

[0045] Figure 3 It is a schematic diagram of a voice word graph provided in the embodiment of the present application;

[0046] Figure 4 It is a schematic diagram of the path score corresponding to "shutdown" in a voice word graph provided in the embodiment of the present application;

[0047] Figure 5 It is a schematic diagram of the path score corresponding to "close" in a voice word graph provided in the embodiment of the present application;

[0048] Figure 6 It is the first sub-diagram of the detailed flowchart of keyword recognition under the system architecture in the embodiment of the present application;

[0049] Figure 7 It is a schematic diagram of the path score corresponding to "shutdown" after weight adjustment provided in the embodiment of the present application;

[0050] Figure 8 It is a schematic diagram of the path score corresponding to "close" after weight adjustment provided in the embodiment of the present application;

[0051] Figure 9 It is a schematic diagram of the optimal path in a voice word graph provided in the embodiment of the present application;

[0052] Figure 10 It is the second sub-diagram of the detailed flowchart of keyword recognition under the system architecture in the embodiment of the present application;

[0053] Figure 11 A schematic diagram of a speech word graph in a specific application scenario provided in an embodiment of the present application;

[0054] Figure 12 This is a schematic diagram of the structure of a keyword recognition device in an embodiment of the present application;

[0055] Figure 13 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.

[0057] The following is an introduction to some concepts involved in the embodiments of the present application.

[0058] (1) Keywords: Some specific words preset in the keyword recognition system, generally command words, used to instruct the device to complete specific actions.

[0059] (2) Acoustic characteristics: physical quantities that characterize the acoustic characteristics of speech, including all acoustic manifestations of the four elements of sound (sound length, sound intensity, pitch, and sound quality), such as the strong frequency concentration area, resonance peak frequency, resonance peak bandwidth, etc. that characterize the sound quality, and the duration, fundamental frequency, and average speech power of the super-sound quality components.

[0060] (3) Phoneme: It is the smallest unit of speech divided according to the natural properties of speech. It is analyzed based on the pronunciation action in the syllable. One action constitutes a phoneme. For example, the Chinese syllable 啊 (ā) has only one phoneme, 爱 (ài) has two phonemes, 代 (dài) has three phonemes, etc.

[0061] (4) Lattice graph: It is essentially a directed acyclic graph. Each word graph contains a start node and an end node. Each node is connected by a directed arc. Each node or arc stores a possible recognition result of a speech frame, as well as the acoustic probability, language probability and other scoring information corresponding to the recognition result.

[0062] (5) Confidence: It represents the credibility of the recognition result. Generally, one or more methods such as acoustic score, image score, image confusion degree, language model fallback probability, etc. can be selected to calculate the confidence of the recognition result.

[0063] The preferred embodiments of the present application will be introduced in detail below in conjunction with the accompanying drawings.

[0064] Refer to Figure 1 As shown, in the embodiments of the present application, there are three main parts: a target object 100, an intelligent device 101, and a keyword recognition system 103. Among them, the keyword recognition system 103 is configured in the intelligent device 101. The intelligent device 101 can be a terminal device or a server device. The terminal device can include, but is not limited to, intelligent assistants, intelligent speakers, air conditioners, TVs, and other electrical appliances. The server device can include, but is not limited to, cloud servers, distributed servers, blockchain servers, independent physical servers, etc.; the target object 100 is used to generate a voice signal to be recognized, and the keyword recognition system 103 is used to perform real-time frame-by-frame decoding on the obtained voice signal to be recognized;

[0065] As an embodiment, the voice signal to be recognized can be collected in real time. For example, when the target object 100 emits a voice signal, the keyword recognition system 103 adopts a streaming processing method to receive the voice signal in real time as the voice signal to be recognized and perform real-time frame-by-frame decoding on it, and controls the intelligent device 101 to complete a specific action corresponding to the keyword according to the result of the real-time frame-by-frame decoding.

[0066] As an embodiment, the voice signal to be recognized can also be pre-collected offline voice. During the process of inputting the offline voice into the keyword recognition system 103, the keyword recognition system 103 takes the offline voice received in real time as the voice signal to be recognized and adopts a streaming processing method to perform real-time frame-by-frame decoding on the voice signal to be recognized received in real time.

[0067] Based on the above system architecture, refer to Figure 2 As shown, in the embodiments of the present application, the detailed process of the keyword recognition system decoding the voice signal collected in real time is as follows:

[0068] Step 201: Perform real-time frame-by-frame decoding on the obtained voice signal to be recognized. When the duration of the first decoded voice data is not less than N seconds, the data part obtained within N seconds before the first current moment in the first voice data is used as the first decoding result.

[0069] Specifically, in the embodiments of the present application, the keyword recognition system takes the voice signal received in real time as the voice signal to be recognized obtained in real time in Step 201 and performs real-time frame-by-frame decoding on it. If the duration of the first decoded voice data is less than N seconds and contains a preset keyword, the first voice data is directly used as the first decoding result, where the first voice data refers to the decoding result corresponding to all the currently decoded voice frames in the voice signal to be recognized.

[0070] For example, after the current speech frame is decoded, it is determined whether the duration of the first decoded speech data is less than 3 seconds. If so, it is further determined whether the first decoded speech data contains a preset keyword (such as "play music", "shutdown", etc.). If it contains, the frame-by-frame decoding is stopped, and the first speech data is used as the first decoding result. Otherwise, the frame-by-frame decoding continues until the decoded speech data contains the preset keyword or the duration of the decoded speech data is not less than 3 seconds. At this time, the moment corresponding to the last decoded speech frame is used as the first current moment, and the data part obtained from the first current moment backward for 3 seconds in the first speech data is used as the first decoding result, where the value of N is preset according to actual experience or set according to the length of the keyword. The longer the keyword length, the larger the value of N; conversely, the smaller the value of N.

[0071] Based on the description in step 201, by only obtaining the first decoding result for the duration of N seconds backward from the current moment, the decoding duration is reduced and the decoding speed is increased.

[0072] Step 202: If the first decoding result contains a preset keyword, continue to perform frame-by-frame decoding on the speech to be recognized until a non-keyword appears in the decoded second speech data. At this time, obtain the speech word graph generated during the decoding of the second speech data for the duration of N seconds backward from the second current moment. The speech word graph contains: multiple candidate character sequences with a similarity higher than a set threshold to the keyword.

[0073] For example, if the first decoding result contains a preset keyword (such as "shutdown"), continue to perform frame-by-frame decoding on the speech to be recognized until a non-keyword appears in the decoded second speech data. At this time, the speech frame where the non-keyword appears is used as the end point of the second speech data, and the moment corresponding to the speech frame where the non-keyword appears is used as the second current moment. Then obtain the Lattice graph generated during the decoding of the second speech data for the duration of 3 seconds backward from the second current moment. The Lattice graph contains multiple candidate character sequences with a similarity higher than the set threshold to the preset keyword. Here, the second speech data refers to the decoding result corresponding to all speech frames decoded from the first current moment to the second current moment.

[0074] Optionally, refer to Figure 3As shown in the figure, in the Lattice diagram of the embodiment of the present application, there are two candidate word sequences, "shutdown" and "turn off", which are respectively stored in the corresponding nodes. Combining an input node and an output node, each candidate word sequence forms a path. The similarity between each candidate word sequence and the preset keyword is reflected by the comprehensive score of the graph score and the acoustic score of each candidate word sequence. The path with the highest comprehensive score is the optimal path, and the keyword recognition system considers the candidate word sequence corresponding to the optimal path as the keyword. When the comprehensive score of the graph score and the acoustic score of the recognition result differs from the comprehensive score of the optimal path in the current Lattice diagram by no more than 10, the keyword recognition system saves the recognition result as a candidate word sequence in the Lattice diagram.

[0075] For example, referring to Figure 4 and Figure 5 As shown, in the candidate word sequence "shutdown", the graph score of the candidate word "machine" is 4 and the acoustic score is 50. In the candidate word sequence "turn off", the graph score of the candidate word "close" is 5 and the acoustic score is 40. The graph score of the same candidate word "off" is 6 and the acoustic score is 50. Then the comprehensive score of the candidate word sequence "shutdown" is 110, and the comprehensive score of the candidate word sequence "turn off" is 101.

[0076] In addition, in some embodiments, if the first decoding result does not contain the preset keyword, the decoded first speech data is discarded, and the subsequent real-time collected speech signal to be recognized is decoded frame by frame again.

[0077] Based on step 202, the end point of the keyword speech is determined according to the decoding result obtained from the first decoding, reducing the latency of keyword recognition.

[0078] Step 203: According to a preset ratio, adjust the proportion of the evaluation index of each candidate word included in each candidate word sequence in the speech word graph respectively, and based on the adjustment result, obtain the target evaluation value of each candidate word sequence respectively.

[0079] Specifically, in the embodiment of the present application, the target evaluation value of each candidate word sequence is obtained in the following manner:

[0080] Referring to Figure 6 As shown: For each candidate word sequence, the following operations are performed respectively:

[0081] Step 2031: For each candidate word included in a candidate word sequence, the following operations are performed respectively: Based on a preset ratio, adjust the weights of the evaluation indexes associated with a candidate word, and based on each evaluation index and the corresponding weight, obtain the word evaluation value of a candidate word.

[0082] For example, the comprehensive score of each candidate word sequence is used as the target evaluation value of each candidate word sequence. The evaluation index of each candidate word sequence includes the image score and the acoustic score. Assuming that the initial weight corresponding to each evaluation index is 1, refer to Figure 7 As shown, for the candidate character "机" contained in the candidate character sequence "关机", the weight of its acoustic score is adjusted based on a ratio of 0.5, and the weight of the image score is kept unchanged, then the final character evaluation value of the candidate character "机" is: 0.5×50+4=29.

[0083] Similarly, see Figure 8 As shown, for the candidate character "闭" contained in the candidate character sequence "闭", the weight of its acoustic score is adjusted based on a ratio of 0.5, and the weight of the image score is kept unchanged, then the final character evaluation value of the candidate character "闭" is: 0.5×40+5=25.

[0084] against Figure 7 and Figure 8 For the same candidate character “关” in , adjust the weight of its acoustic score based on the ratio of 0.5, and keep the weight of the image score unchanged. Then the final character evaluation value of the candidate character “闭” is: 0.5×50+6=31.

[0085] Step 2032: Based on the character evaluation value of each candidate character, a target evaluation value of a candidate character sequence is obtained.

[0086] For example, the target evaluation value of each candidate word sequence is the sum of the word evaluation values corresponding to each candidate word contained in each candidate word sequence. Therefore, the target evaluation value of the candidate word sequence "shutdown" is 31+29=60, and the target evaluation value of the candidate word sequence "close" is 31+25=56.

[0087] Step 204: Based on the obtained target evaluation values, the speech word graph is decoded to obtain a second decoding result, and when the second decoding result contains a keyword, the second decoding result is output as a recognition result of the speech signal to be recognized.

[0088] Specifically, in an embodiment of the present application, the keyword contained in the second decoding result is composed of a candidate word sequence in the speech word graph, and when the keyword is included in the second decoding result, before outputting the second decoding result as a recognition result, the keyword recognition system also needs to obtain the confidence of the second decoding result based on the target evaluation value of the candidate word sequence corresponding to the keyword. If the confidence of the second decoding result is greater than a preset confidence threshold, the second decoding result is output as a recognition result.

[0089] For example, see Figure 9As shown in the figure, the target evaluation value 60 of the candidate word sequence "shutdown" is the maximum value in the Lattice diagram. Then, the keyword recognition system determines the candidate word "shutdown" as a keyword. Before outputting the candidate word sequence "shutdown" as the recognition result, the keyword recognition system will also perform a confidence judgment on it. Assuming that the target evaluation value of the candidate word sequence "shutdown" is directly used as its confidence level, and the preset confidence threshold is 58. Since 60 > 58, the keyword recognition system believes that this recognition result is credible, and then outputs the candidate word sequence "shutdown" as a recognition result segment.

[0090] Optionally, when the keyword recognition system performs a confidence judgment on the second decoding result, in addition to using the target evaluation value composed of the graph score and the acoustic score for the judgment, one or more methods such as the graph confusion degree and the language model backoff probability can also be added as the basis for the confidence judgment.

[0091] For example, refer to Figure 9 As shown in the figure, assume that the preset confidence threshold is 65, and there is no candidate word sequence with a target evaluation value greater than 65 in the current Lattice diagram. Then, at this time, the graph confusion degree can be further used to determine whether the graph contains keywords. Since there are only two paths in the Lattice diagram at this time, indicating a relatively low graph confusion degree (this method is an existing technology and will not be described in detail here), the keyword recognition system directly outputs the candidate word sequence "shutdown" with the highest target evaluation value in the Lattice diagram as a recognition result segment, and instructs the intelligent device to complete the shutdown.

[0092] On the other hand, if the confidence level of the second decoding result is not greater than the preset confidence threshold, the second decoding result is discarded, and the subsequent real-time collected speech signal to be recognized is decoded frame by frame again.

[0093] Based on step 204, the secondary decoding of the first decoding result and the confidence judgment of the second decoding result improve the accuracy of keyword recognition.

[0094] Further, in the embodiment of the present application, the keyword recognition system decodes the speech signal to be recognized frame by frame in the following manner:

[0095] The speech signal to be recognized is generally a discrete-time signal. When the keyword recognition system receives the speech signal to be recognized, it needs to perform processing such as frame division, windowing, and pre-emphasis on it.

[0096] Refer to Figure 10 As shown in the figure: For each speech frame included in the speech signal to be recognized, the following operations are respectively performed:

[0097] Step 2041: Extract the acoustic features of a speech frame, send the acoustic features into a pre-trained acoustic model, and obtain the phoneme probability distribution of a speech frame.

[0098] Specifically, the acoustic features can be one type, such as Mel-Frequency Cepstral Coefficients (MFCC) features, Filter bank (FBANK) features, pitch features, and Identity Vector (i-vector) features, etc., or can be a fusion of multiple acoustic features, such as MFCC + i-vector features. The keyword recognition system inputs the extracted acoustic features into a pre-trained acoustic model to obtain the phoneme probability distribution corresponding to the speech frame, that is, the acoustic score.

[0099] Step 2042: Obtain the decoding result of a speech frame based on the phoneme probability distribution.

[0100] Specifically, input the phoneme probability distribution and the pre-trained decoding graph into the decoder, and the text content corresponding to the speech frame can be recognized. Among them, in the pre-trained decoding graph, there are multiple candidate word sequences similar to multiple preset keywords.

[0101] The above embodiments will be further described in detail through a specific application scenario below.

[0102] If the first decoding result contains the preset keyword "open", the keyword recognition system continues to decode the speech to be recognized frame by frame until a non-keyword appears in the decoded second speech data. The speech frame where the non-keyword appears is used as the end point of the second speech data, and the Lattice graph generated during the decoding of the second speech data within 3 seconds before the current speech frame is obtained. Refer to Figure 11 As shown, the Lattice graph contains three paths composed of three candidate word sequences similar to the keyword, namely "open", "clock in", and "bandwidth". Among them, based on the adjustment results of the word evaluation values corresponding to each candidate word included in each candidate word sequence, the target evaluation values corresponding to each candidate word sequence are: "open": 30 + 25 = 55; "clock in": 30 + 20 = 50; "bandwidth": 25 + 10 = 35. The optimal path is the path corresponding to the candidate word sequence "open", then the keyword recognition system performs a confidence judgment on the candidate word sequence "open". Assume that the target evaluation value 55 corresponding to the candidate word sequence "open" is its corresponding confidence, and the preset confidence threshold is 65. Since 56 < 65, the keyword recognition system determines that the second decoding result does not contain the keyword and discards the second decoding result, and re-performs frame-by-frame decoding on the subsequent real-time collected speech signal to be recognized.

[0103] In addition, although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0104] Based on the same technical concept, referring to Figure 12 as shown, an embodiment of the present application further provides a keyword recognition device, and the device includes:

[0105] A first decoding module 1201, configured to perform real-time frame-by-frame decoding on the acquired speech signal to be recognized, and when the duration of the first decoded speech data is not less than N seconds, use the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result;

[0106] A second decoding module 1202, if the preset keyword is included in the first decoding result, continues to perform frame-by-frame decoding on the speech to be recognized until non-keyword appears in the second decoded speech data, and acquires the speech word graph generated during the decoding process of the second speech data within N seconds before the second current moment, where the speech word graph includes: multiple candidate character sequences whose similarity to the keyword is higher than the set threshold;

[0107] An adjustment module 1203, adjusts the proportion of the evaluation index of each candidate character included in each candidate character sequence in the speech word graph according to a preset ratio, and respectively obtains the target evaluation value of each candidate character sequence based on the adjustment result;

[0108] An output module 1204, decodes the speech word graph based on the obtained target evaluation values to obtain a second decoding result, and when the keyword is included in the second decoding result, outputs the second decoding result as the recognition result of the speech signal to be recognized.

[0109] Optionally, when using the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result, the first decoding module 1201 is further configured to:

[0110] If the duration of the first decoded speech data is less than N seconds and includes a preset keyword, directly use the first speech data as the first decoding result.

[0111] Optionally, the first decoding module 1201 is further configured to:

[0112] If the preset keyword is not included in the first decoding result, discard the first decoded speech data, and re-perform frame-by-frame decoding on the subsequent real-time acquired speech signal to be recognized.

[0113] Optionally, when adjusting the proportion of the evaluation index of each candidate character included in each candidate character sequence according to a preset ratio and obtaining the target evaluation value of each candidate character sequence based on the adjustment result, the adjustment module 1203 is configured to:

[0114] For each candidate character sequence, perform the following operations respectively:

[0115] For each candidate character included in a candidate character sequence, perform the following operations respectively: Based on a preset ratio, adjust the weights of the evaluation indexes associated with a candidate character, and obtain the character evaluation value of a candidate character based on each evaluation index and the corresponding weights;

[0116] Based on the character evaluation values of each candidate character, obtain the target evaluation value of a candidate character sequence.

[0117] Optionally, if the keyword included in the second decoding result is composed of a candidate character sequence in the speech word graph, then when the second decoding result includes the keyword, before outputting the second decoding result as a segment of recognition result, the output module 1204 is further configured to:

[0118] Based on the target evaluation value of the candidate character sequence corresponding to the keyword, obtain the confidence of the second decoding result;

[0119] If the confidence of the second decoding result is greater than a preset confidence threshold, then output the second decoding result as a segment of recognition result.

[0120] Optionally, the output module 1204 is further configured to:

[0121] If the confidence of the second decoding result is not greater than a preset confidence threshold, then discard the second decoding result and re-perform frame-by-frame decoding on the subsequent real-time collected speech signal to be recognized.

[0122] Optionally, when performing frame-by-frame decoding on the speech signal to be recognized, the first decoding module 1201 and the second decoding module 1202 are configured to:

[0123] For each speech frame included in the speech signal to be recognized, perform the following operations respectively:

[0124] Extract the acoustic features of a speech frame, send the acoustic features into a pre-trained acoustic model, and obtain the phoneme probability distribution of a speech frame;

[0125] Based on the phoneme probability distribution, obtain the decoding result of a speech frame.

[0126] Based on the same technical concept, an embodiment of the present application further provides an electronic device, and this electronic device can implement the method flow of keyword recognition provided in the above embodiments of the present application.

[0127] In one embodiment, the electronic device can be a server, a terminal device, or other electronic devices.

[0128] Refer to Figure 13 As shown, the electronic device may include:

[0129] At least one processor 1301 and a memory 1302 connected to the at least one processor 1301. In the embodiments of the present application, the specific connection medium between the processor 1301 and the memory 1302 is not limited. Figure 13 Here, it is taken as an example that the processor 1301 and the memory 1302 are connected through a bus 1300. The bus 1300 is Figure 13 represented by a thick line herein. The connection manners between other components are only for illustrative purposes and are not limiting. The bus 1300 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 13 it is only represented by a thick line herein, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 1301 can also be called a controller, and the name is not limited.

[0130] In the embodiments of the present application, the memory 1302 stores instructions executable by the at least one processor 1301. By executing the instructions stored in the memory 1302, the at least one processor 1301 can execute a keyword recognition method described above. The processor 1301 can implement Figure 12 the functions of each module in the device shown.

[0131] Among them, the processor 1301 is the control center of the device. It can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 1302 and calling the data stored in the memory 1302, various functions of the device and process data, so as to monitor the device as a whole.

[0132] In a possible design, the processor 1301 may include one or more processing units. The processor 1301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 1301. In some embodiments, the processor 1301 and the memory 1302 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.

[0133] The processor 1301 may be a general-purpose processor, such as a CPU, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of a keyword recognition method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0134] The memory 1302, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1302 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disc, and so on. The memory 1302 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1302 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0135] By designing and programming the processor 1301, the code corresponding to a keyword recognition method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 2 the steps of a keyword recognition method of the embodiment shown. How to design and program the processor 1301 is a well-known technology to those skilled in the art and will not be elaborated here.

[0136] Based on the same inventive concept, the embodiments of the present application also provide a storage medium that stores computer instructions, and when the computer instructions are run on a computer, the computer is enabled to execute a keyword recognition method described above.

[0137] In some possible embodiments, aspects of a keyword recognition method provided by the present application may also be implemented in the form of a program product, which includes program code that, when the program product runs on a device, causes the control device to execute the steps in a keyword recognition method according to various exemplary embodiments of the present application described above in this specification.

[0138] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by multiple units.

[0139] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0140] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0141] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the function specified in one or more of the processes Figure 1 steps or a plurality of processes and / or boxes Figure 1 boxes or a plurality of boxes.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes Figure 1 steps or a plurality of processes and / or boxes Figure 1 boxes or a plurality of boxes.

[0144] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A keyword recognition method, characterized in that, The method includes: Performing real-time frame-by-frame decoding on the obtained speech signal to be recognized until the duration of the first decoded speech data is not less than N seconds, and taking the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result; If the preset keyword is included in the first decoding result, continue to perform frame-by-frame decoding on the speech to be recognized until a non-keyword appears in the second decoded speech data, and obtain the speech word graph generated during the decoding of the second speech data within N seconds before the second current moment. The speech word graph includes: multiple candidate character sequences with a similarity higher than a set threshold to the keyword; For each candidate character sequence, perform the following operations respectively: For each candidate character included in a candidate character sequence, perform the following operations respectively: Based on a preset ratio, adjust the weights of the respective evaluation metrics associated with a candidate character, and obtain the character evaluation value of the candidate character based on the respective evaluation metrics and the corresponding weights, where the evaluation metrics include grapheme score and phonetic score; Based on the character evaluation values of the respective candidate characters, obtain the target evaluation value of a candidate character sequence; Based on the obtained target evaluation values, decode the speech word graph to obtain a second decoding result, and when the keyword is included in the second decoding result, output the second decoding result as the recognition result of the speech signal to be recognized.

2. The method according to claim 1, characterized in that, When taking the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result, it further includes: If the duration of the first decoded speech data is less than N seconds and includes the preset keyword, directly take the first speech data as the first decoding result.

3. The method according to claim 1, characterized in that, It further includes: If the preset keyword is not included in the first decoding result, discard the first decoded speech data and re-perform frame-by-frame decoding on the subsequent real-time collected speech signal to be recognized.

4. The method according to claim 1, wherein If the keyword included in the second decoding result is composed of a candidate character sequence in the speech word graph, before outputting the second decoding result as a segment of recognition result when the keyword is included in the second decoding result, it further includes: Based on the target evaluation value of the candidate character sequence corresponding to the keyword, obtain the confidence level of the second decoding result; If the confidence level of the second decoding result is greater than a preset confidence level threshold, output the second decoding result as a segment of recognition result.

5. The method according to claim 4, wherein It further includes: If the confidence level of the second decoding result is not greater than the preset confidence level threshold, discard the second decoding result and re-perform frame-by-frame decoding on the subsequent real-time collected speech signal to be recognized.

6. The method according to claim 1, wherein The frame-by-frame decoding of the speech signal to be recognized includes: For each speech frame included in the speech signal to be recognized, perform the following operations respectively: Extract the acoustic features of a speech frame, send the acoustic features into a pre-trained acoustic model, and obtain the phoneme probability distribution of the speech frame; Based on the phoneme probability distribution, obtain the decoding result of the speech frame.

7. A keyword recognition device, characterized in that, It includes: A first decoding module, configured to perform frame-by-frame decoding on a speech signal to be recognized collected in real time, and when the duration of the first decoded speech data is not less than N seconds, use the data part obtained within N seconds before the first current moment in the first speech data as the first decoding result; A second decoding module, if the preset keyword is included in the first decoding result, continues to perform frame-by-frame decoding on the speech to be recognized until a non-keyword appears in the second decoded speech data, and obtains a speech word graph generated during the decoding process of the second speech data within N seconds before the second current moment. The speech word graph includes: a plurality of candidate character sequences whose similarity to the keyword is higher than a set threshold; An adjustment module, which respectively performs the following operations for each candidate character sequence: For each candidate character included in a candidate character sequence, respectively perform the following operations: Based on a preset ratio, adjust the weights of the respective evaluation indicators associated with a candidate character, and based on the respective evaluation indicators and the corresponding weights, obtain a character evaluation value of the candidate character, where the evaluation indicators include graph score and acoustic score; Based on the character evaluation values of the respective candidate characters, obtain a target evaluation value of the candidate character sequence; An output module, based on the obtained target evaluation values, decodes the speech word graph to obtain a second decoding result, and when the keyword is included in the second decoding result, outputs the second decoding result as a recognition result segment.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1-6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method described in any one of claims 1-6 are implemented.

10. A computer program product, characterized in that, When the computer program product is called by a computer, the computer is caused to execute the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Voice activation method and system

    CN105374352A

  • Decoding method and device

    CN108932944A