An end-side keyword recognition method and system based on front and rear audio compensation and hierarchical determination
By re-feeding preamble audio and continuing to decode the silent audio at the end of the voice during edge-side streaming voice interaction, and combining dynamic thresholds and path determination, the problems of incomplete recognition of the start and end of voice, false wake-up and false commands are solved, thus improving recognition stability and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-14
AI Technical Summary
In existing edge-side streaming keyword recognition solutions, the beginning segment of speech is easily truncated due to segment boundaries, delayed judgment of speech activity detection, or insufficient first frame, and the ending sound is easily truncated due to premature end judgment, resulting in incomplete recognition; it is easy to miss detection in scenarios such as weak syllables, slight near sounds, and slight tone deviations; it is difficult to uniformly process keywords of different lengths with thresholds, resulting in a high probability of false wake-up and false command triggering; the echo of the prompt tone after successful recognition and residual audio pollute the subsequent recognition link, causing repeated triggering and state disorder.
By re-feeding preamble audio at the start of speech, continuing decoding to supplement the silent audio at the end, combining the length verification of the precise hit path and the matching depth judgment of the fuzzy path, dynamically adjusting the threshold, coordinating the processing of wake word and command word paths, resetting the speech activity detection state and cache after output, and supporting dynamic updates of the keyword set.
It improves the completeness of speech start and end recognition, reduces the probability of false wake-up, enhances the recognition adaptability of keywords of different lengths, reduces interference in the recognition process, and enhances the stability and robustness of the link.
Smart Images

Figure CN122392495A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech recognition and intelligent voice interaction technology, specifically relating to a keyword recognition method, system, electronic device, and computer-readable storage medium for edge-side streaming voice interaction. Background Technology
[0002] With the development of smart homes, consumer electronics, automotive devices, industrial terminals, and various low-power embedded devices, edge-side offline voice interaction is gradually becoming an important interaction method for smart hardware. In existing technologies, edge-side voice interaction solutions typically achieve voice triggering and control through preset wake words or command words, and use keyword detection or wake word detection technology to recognize continuous input voice.
[0003] Existing edge-side keyword recognition solutions commonly employ the following approaches: pre-define a keyword set, perform basic voice activity detection gating on the input speech, perform keyword decoding after speech detection, and determine whether to output a wake-up event or command event based on whether the decoding result matches a preset keyword. For some products, a command word recognition stage is also initiated after wake-up to support further voice control operations.
[0004] Existing technologies for implementing edge-side streaming keyword recognition have the following main drawbacks: 1. When speech transitions from a silence segment to a speech segment, the initial speech segment is easily truncated due to segment boundaries, delays in speech activity detection, or insufficient first frames, resulting in incomplete recognition of the first syllable or the first character. When speech transitions from a speech segment to a silence segment, the final sound is easily truncated due to premature termination detection, leading to missed keyword detection. 2. Existing solutions largely rely on precise keyword hits for trigger determination, which can easily lead to missed detections in scenarios such as weak syllables, slight similar sounds, slight tone deviations, accent differences, or insufficient sound pickup capabilities of the device. Simply relaxing the determination threshold to improve the hit rate would significantly increase the probability of false wake-ups and false command triggers, making it difficult to simultaneously achieve a good balance between recognition rate and false trigger rate. 3. When a fixed threshold is applied uniformly to keywords of different lengths and categories, short keywords are prone to misjudgment due to an overly lenient threshold, while long keywords may result in missed judgments due to an overly strict threshold. Furthermore, a lack of differentiated processing between wake words and command words will also affect the overall recognition performance. 4. In scenarios involving continuous interaction on the edge device, the prompt tone or feedback tone played after successful recognition can easily be picked up again by the microphone, creating echo interference. If the voice activity detection state, audio cache, and recognition stream state are not cleared in a timely manner, residual audio and historical states may contaminate subsequent recognition processes, causing repeated triggering, state disorder, or misjudgment. 5. Existing technologies often employ independent local optimization methods for requirements such as segmented compensation, recognition and judgment, state switching, false wake-up suppression, and dynamic keyword set updates. They lack a collaborative processing mechanism around the same end-side streaming recognition link, making it difficult to simultaneously ensure recognition stability, false trigger control, and robustness of continuous interaction in complex scenarios. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this invention proposes an end-side keyword recognition method and system based on front-to-back audio compensation and hierarchical determination, in order to improve keyword recognition stability, reduce false trigger probability, and enhance link robustness in continuous interaction scenarios in end-side streaming voice interaction scenarios. I. The technical problem to be solved by the invention
[0006] The present invention mainly solves the following technical problems: 1. In existing edge-side streaming keyword recognition schemes, when speech transitions from a silence segment to a speech segment, the initial audio is easily truncated due to segment boundaries, delays in speech activity detection, or insufficient first frames, resulting in incomplete recognition of the first syllable or the first character. 2. When speech transitions from a speech segment to a silence segment, the ending sound is easily truncated due to premature termination detection, resulting in missed keyword detection. 3. Existing solutions mostly rely on precise keyword hits for trigger determination, which can easily lead to missed detections in scenarios such as weak syllables, slight near-sounds, slight tone deviations, accent differences, or insufficient sound pickup capabilities. Simply relaxing the determination threshold can easily lead to an increase in false wake-ups and false command triggers. 4. When applying a fixed threshold to keywords of different lengths and categories, it is difficult to simultaneously achieve both high recognition rate and low false trigger rate. 5. In continuous interaction scenarios on the device side, the echo of the prompt tone after successful recognition, residual audio, and historical states can easily pollute subsequent recognition links, causing repeated triggering, state disorder, or misjudgment. 6. Existing technologies lack a collaborative processing mechanism around the same recognition link for requirements such as segment compensation, recognition and judgment, state reset, mode switching, and dynamic keyword set update. II. Technical Solution
[0007] A keyword recognition method for edge-side streaming voice interaction includes acquiring a continuous voice stream from the edge and performing voice activity detection, keyword decoding, hit verification, and result output processing on the continuous voice stream, characterized in that: Acquire a continuous audio stream from the device side, establish a keyword recognition stream based on the continuous audio stream, and identify whether the current audio is in a speech segment or a silent segment based on the speech activity detection result; When a speech segment is detected to be transitioning from a silence segment, at least a portion of the preamble audio cached during the silence segment is fed back into the keyword recognition stream; Keyword decoding is continuously performed within the speech segment, and the peak depth of context matching in the current speech segment is recorded based on the changes in context matching status during the decoding process. When the end trend of the speech segment is detected, the continuation decoding is performed and the end-of-segment silence audio is added to the keyword recognition stream; When a precise keyword hit result is obtained, the minimum pass matching length is calculated based on the keyword length information corresponding to the precise keyword, and the length of the precise keyword hit result is verified based on the comparison result between the current decoded output matching length and the minimum pass matching length. When no precise keyword hit result is obtained, a dynamic judgment threshold is calculated based on the context matching peak depth and the keyword length information corresponding to the candidate keyword, and a candidate fuzzy hit result is generated based on the comparison result between the context matching peak depth and the dynamic judgment threshold. When the exact keyword hit result passes the length check, or the candidate fuzzy hit result passes the check, the corresponding keyword recognition result is output; After outputting the keyword recognition results, reset the voice activity detection status and related cache.
[0008] Furthermore, the keyword recognition results include wake word recognition results and command word recognition results. The method also includes: distinguishing wake word paths and command word paths according to the current recognition mode, and adopting different matching length verification rules for different paths.
[0009] Furthermore, the preamble audio is historical audio that is continuously cached during the silence phase, and the step of feeding back at least a portion of the preamble audio cached during the silence phase into the keyword recognition stream includes: when a speech segment is detected to be entering from a silence segment, selecting a preamble audio of a predetermined length from the historical audio and injecting the preamble audio into the keyword recognition stream.
[0010] Furthermore, the step of continuing to perform continuation decoding when a speech segment ending trend is detected, and supplementing the keyword recognition stream with tail-end silence audio, includes: after the speech activity detection result is converted from a speech segment to a silence segment, continuing to perform keyword decoding within a predetermined hangover period, and supplementing the keyword recognition stream with tail-end silence audio and / or flush audio after the hangover period ends.
[0011] Furthermore, the keyword length information is preferably the total length of the token corresponding to the keyword.
[0012] Furthermore, the context matching peak depth is determined by continuously reading the context state level corresponding to the current optimal decoding path during multiple decoding processes of the current speech segment, and determining the maximum level read as the context matching peak depth of the current speech segment.
[0013] Furthermore, the precise keyword hit results and candidate fuzzy hit results are processed in a hierarchical manner. That is, length verification is performed first based on the precise keyword hit results, and candidate fuzzy hit results are generated only when no precise keyword hit results that pass the length verification are obtained in the current speech segment.
[0014] Furthermore, the method also includes: receiving wake word text or command word text configured at runtime, converting the wake word text or command word text into a keyword representation acceptable to the keyword recognition stream, and updating the keyword set and keyword recognition stream based on the conversion result.
[0015] Furthermore, the step of converting the wake word text or command word text into a keyword representation acceptable to the keyword recognition stream includes: converting Chinese character keywords into pinyin token sequences, and generating multiple pronunciation combinations when polyphonic characters exist, so as to write the multiple pronunciation combinations as multiple keyword representations of the same keyword into the keyword set.
[0016] Furthermore, the method also includes performing gain adjustment, noise threshold control, and / or target energy alignment processing on the input speech stream.
[0017] Furthermore, the method also includes: initiating a cooling-off period control after the keyword recognition result is output, and suppressing repeated triggering during the cooling-off period.
[0018] Furthermore, the method also includes: switching to command word recognition mode after successful wake word recognition, and reconstructing the keyword recognition stream based on the switched keyword set; and restoring to wake word recognition mode after command word recognition is completed.
[0019] Furthermore, the process of restoring the wake-up word recognition mode after the command word recognition is completed also includes: reconstructing the keyword recognition stream based on the restored keyword set, so that the restored wake-up word set becomes effective again in the current recognition link.
[0020] Furthermore, the present invention also provides a keyword recognition system for end-to-end streaming voice interaction to implement the above method, comprising: Audio receiving module, used to acquire continuous voice stream from the end side; The voice activity detection module is used to identify whether the current audio is in a speech segment or a silent segment based on the voice activity detection results; The preamble buffer and backfeed module is used to backfeed at least a portion of the preamble audio buffered during the silence phase to the keyword recognition stream when a speech segment is detected to be transitioned from a silence segment. The keyword recognition module is used to continuously perform keyword decoding within a speech segment; The depth tracking module is used to record the peak depth of context matching for the current speech segment based on changes in context matching state during decoding. The tail segment compensation module is used to continue performing sustained decoding when the end trend of the speech segment is detected, and to supplement the keyword recognition stream with tail segment silence audio; The precise hit verification module is used to calculate the minimum pass matching length based on the keyword length information corresponding to the precise keyword when a precise keyword hit result is obtained, and to perform length verification on the precise keyword hit result based on the comparison result between the current decoded output matching length and the minimum pass matching length. The fuzzy judgment module is used to calculate a dynamic judgment threshold based on the context matching peak depth and the keyword length information corresponding to the candidate keyword when no accurate keyword hit result is obtained, and to generate a candidate fuzzy hit result based on the comparison result between the context matching peak depth and the dynamic judgment threshold. The result output module is used to output the corresponding keyword recognition result when the exact keyword hit result passes the length check or the candidate fuzzy hit result passes the check. The state reset module is used to reset the speech activity detection state and related cache after the keyword recognition results are output.
[0021] Furthermore, the system also includes a dynamic keyword update module, which receives the wake word text or command word text configured at runtime, converts it into a keyword representation acceptable to the keyword recognition stream, and updates the keyword set and keyword recognition stream based on the conversion result.
[0022] The present invention also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the electronic device performs the above-described keyword recognition method.
[0023] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor performs the above-described keyword recognition method. III. Beneficial Effects
[0024] Compared with the prior art, the beneficial effects of the present invention are: 1. By re-injecting the preamble audio at the beginning of speech, the problem of missed initial syllables caused by the loss of the starting segment is reduced. 2. By continuing decoding after the speech ends and supplementing with silent audio at the end, keyword hit rate is improved in scenarios with a light ending sound. 3. By verifying the length of the exact hit path and making a secondary determination based on the matching depth of the inaccurate hit path, fault tolerance is improved while false wake-ups are suppressed. 4. By using dynamic threshold control based on keyword length information, the adaptability of keywords of different lengths in the unified recognition link is improved. 5. By resetting the speech activity detection state and related buffers after the recognition results are output, the interference of prompt echoes and residual speech on subsequent recognition processes is reduced. 6. Improve link stability in single-turn dialogues, long-sentence dialogues, and scenarios with one wake-up word and multiple branches by coordinating the processing of wake-up word paths and command word paths. 7. By supporting runtime updates of the keyword set, dynamic keyword configuration is more easily integrated with the edge-side streaming recognition process. 8. By sequentially performing preamble compensation, tail compensation, precise path length verification, fuzzy path peak depth determination, and post-hit state reset within the same speech segment, a closed-loop collaborative control mechanism is formed around a single speech segment, thereby improving the hit rate while avoiding the false triggering amplification problem caused by simply relaxing the threshold. Attached Figure Description
[0025] Figure 1 A schematic diagram of the end-side streaming keyword recognition system according to one embodiment of the present invention; Figure 2 This is the main flowchart of the end-side streaming keyword recognition method according to one embodiment of the present invention; Figure 3 This is a flowchart of the preamble audio re-feedback and tail segment compensation in one embodiment of the present invention. Figure 4 This is a flowchart of the accurate hit path verification process in one embodiment of the present invention; Figure 5 This is a flowchart of the fuzzy hit path determination in one embodiment of the present invention; Figure 6 This is a flowchart of the wake word and command word mode switching in an optional embodiment of the present invention; Figure 7 This is a flowchart illustrating the status reset process after successful identification in one embodiment of the present invention.
[0026] Among them, such as Figure 1As shown, 100 represents the input and preprocessing layer, 101 represents the audio input module, 102 represents the audio preprocessing module, and 103 represents the speech activity detection and segmentation control module; 200 represents the configuration and control layer, 201 represents the dynamic keyword update module, 202 represents the keyword set construction module, and 203 represents the mode switching module; 300 represents the output and feedback layer, and 301 represents the result output and state reset module; 400 represents the core recognition engine layer, 401 represents the keyword recognition module, 402 represents the accurate hit verification module, 403 represents the fuzzy hit judgment module, and 404 represents the pre- and post-segment compensation module. Detailed Implementation
[0027] The present invention will be further described in detail below with reference to specific embodiments. It should be understood that the following embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention. Without departing from the concept of the present invention, those skilled in the art can make adjustments or substitutions to the module division, processing order, parameter configuration method and implementation details, and such adjustments or substitutions should all fall within the scope of protection of the present invention.
[0028] In this invention, "keyword" is a higher-level concept, including wake-up words and command words; "keyword length information" is a higher-level concept, preferably the total length of the token corresponding to the keyword; "context matching peak depth" is the maximum matching depth corresponding to the context matching structure during the keyword recognition process.
[0029] In this invention, the "hangover cycle" refers to the extended period during which decoding continues to be performed after the speech activity detection result is converted from a speech segment to a silence segment in order to preserve the end-segment convergence capability; the "flush audio" refers to the supplementary audio used to trigger or facilitate the decoder to complete the end-segment state convergence. Example 1
[0030] A keyword recognition system for edge-side streaming voice interaction includes: 1. An input and preprocessing layer 100, comprising an audio input module 101, an audio preprocessing module 102, and a speech activity detection and segmentation control module 103; 2. Configuration and control layer 200, which includes a dynamic keyword update module 201, a keyword set construction module 202, and a mode switching module 203; 3. Output and feedback layer 300, wherein the output and feedback layer 300 includes a result output and state reset module 301; 4. Core recognition engine layer 400, which includes a keyword recognition module 401, a precise hit verification module 402, a fuzzy hit determination module 403, and a front-end and back-end compensation module 404.
[0031] The work process / implementation steps are as follows: 1. For example Figure 2 As shown, in step S202, the continuous speech stream on the end side is acquired; in step S203, speech activity detection is performed; in step S204, it is determined whether the speech segment has entered from the silence segment; if the determination result is yes, then in steps S205 and S206, the preamble audio extraction and re-feedback are performed.
[0032] 2. In steps S207 and S208, keyword decoding is continuously performed and the peak depth of context matching is recorded; in step S209, it is determined whether a speech segment end trend is detected. If the determination result is yes, then in steps S210 and S211, hangover decoding and end-segment mute or flush audio compensation are performed.
[0033] 3. In step S212, it is determined whether an accurate keyword hit result is obtained, and based on the determination result, it proceeds to the accurate hit path verification corresponding to step S213 or the fuzzy hit path determination corresponding to step S214. If the accurate hit path passes the length verification in step S215, or the fuzzy hit path passes the verification in step S217, then the keyword recognition result is output in step S216, and the voice activity detection state and related cache are reset in step S219. Then, in step S220, the current round of recognition ends or the next round of recognition begins. If the corresponding verification fails, the state of waiting to be recognized is returned in step S218.
[0034] The terminal device continuously receives a continuous audio stream from the microphone and sends it to the voice activity detection module and the keyword recognition link. The voice activity detection module determines whether the current input audio is in a speech segment or a silent segment. The keyword recognition link can employ an edge-side keyword detection engine, a streaming decoder, an online keyword recognizer, or other recognition structures capable of outputting keyword candidate results and matching status.
[0035] During the silence phase, the system buffers historical audio from the end of the silence phase to form a preamble buffer. The preamble buffer can be a historical audio segment of fixed duration or a historical audio segment with a fixed number of sampling points.
[0036] When the speech activity detection module detects that the input audio has transitioned from a silence segment to a speech segment, the system selects at least a portion of historical audio from the preamble buffer and feeds this portion of historical audio back into the keyword recognition stream. Subsequently, the system continues to input real-time audio from the current speech segment into the keyword recognition stream to perform subsequent keyword decoding. This method reduces the loss of speech start segments due to segment boundaries or delays in speech activity detection.
[0037] Within a speech segment, the system continuously performs keyword decoding and reads the context matching state corresponding to the current optimal decoding path after each decoding. If the context state depth corresponding to the current optimal decoding path is greater than the currently recorded value, the peak depth of the context matching for the current speech segment is updated. The peak depth of the context matching can be obtained by reading the context state level corresponding to the optimal hypothesis among the candidate hypotheses in the keyword recognition stream at the current moment.
[0038] When the speech activity detection module detects a transition from a speech segment to a silence segment in the input audio, the system does not immediately terminate the recognition of that speech segment. Instead, it first enters the continuation decoding stage. During this stage, the system continues to perform keyword decoding within a predetermined hangover period and may supplement the keyword recognition stream with trailing silence audio or flush audio to preserve the ability to converge at the end of the speech. After the continuation decoding is completed, the system determines whether to output a keyword recognition event based on the final recognition result of that speech segment, the exact hit verification result, and / or the fuzzy hit verification result.
[0039] In this embodiment, the precise hit path and the fuzzy hit path are not a simple substitution for parallel threshold relaxation, but rather a hierarchical judgment relationship performed sequentially for the same speech segment. The system prioritizes using the precise keyword hit result and its minimum passing match length to perform a stricter first-level judgment; only when the first-level judgment fails to obtain a passing result is the second-level judgment mechanism based on the context matching peak depth and keyword length information invoked, in order to reduce the increase in false triggers caused by directly relaxing the threshold overall.
[0040] After outputting the keyword recognition results, the system resets the speech activity detection state, preamble buffer, current speech segment statistics state, and keyword recognition stream state to prevent echoes, residual audio, or old recognition states from affecting the next round of recognition. Example 2
[0041] A preamble audio re-feedback compensation method includes: 1. Continuously buffer the leading audio during silent segments; 2. When a transition from a silence segment to a speech segment is detected, at least a portion of the preamble audio is re-fed; 3. The preamble audio and real-time speech after refeeding are continuously fed into the keyword recognition stream.
[0042] The work process / implementation steps are as follows: 1. For example Figure 3 As shown, in step S301, the preamble audio is continuously buffered in the silence segment, and in step S302, the transition from the silence segment to the speech segment is detected.
[0043] 2. In step S303, at least a portion of the preamble audio is re-fed, and in step S304, the normal speech segment is decoded.
[0044] 3. Compensate for information loss in the initial speech segment caused by segment boundary or speech activity detection delay by re-feeding back the preamble audio.
[0045] In some edge devices, due to delays in voice activity detection, low front-end signal amplitude, complex environmental noise, or soft user initial syllables, several frames of audio at the beginning of a speech segment may not be sent to the keyword recognition link in time, resulting in the first syllable or the first word being truncated. Therefore, this invention continuously buffers the last audio syllable during the silence phase.
[0046] The caching strategy may include: caching the most recent silent segment of audio in a sliding window manner; limiting the total cached amount by a predetermined length; selecting only the most recent cached audio segment for reloading when the speech begins; and performing gain adjustment or target energy alignment on the cached audio before reloading.
[0047] When the start of speech is detected, the system extracts a pre-defined length of preamble audio from the cache and continuously feeds this preamble audio along with the real-time speech into the keyword recognition stream. In this way, starting syllables that might have been lost due to delays in speech activity detection can be reintegrated into the keyword recognition process, thereby improving the recognition stability for scenarios with weak starting syllables and light initial characters. Example 3
[0048] A method for hangover and noise compensation in the tail section includes: 1. Upon detecting a speech termination trend, proceed to the hangover continuation decoding stage; 2. After the hangover phase ends, add a silent audio segment and / or a flush audio segment at the end; 3. The current speech segment recognition ends after the tail segment converges.
[0049] The work process / implementation steps are as follows: like Figure 3 As shown, a speech termination trend is detected in step S305, and hangover decoding is performed in step S306.
[0050] In step S307, add a trailing mute and / or flush audio.
[0051] In step S308, the current speech segment recognition is terminated to improve the stability of keyword recognition in the end-segment scenario.
[0052] When the voice activity detection result changes from a speech segment to a silence segment, the system enters the hangover phase. During the hangover phase, the system continues to feed the current end-segment audio or supplementary audio into the keyword recognition stream and performs subsequent decoding. After the hangover phase ends, the system can also supplement the keyword recognition stream with short silence audio or flush audio to facilitate the decoder's convergence to the end-segment token.
[0053] In one implementation, when the system detects a speech ending trend, it first continues to receive several frames of audio near the silence boundary to avoid cutting off too quickly at the end of the speech segment; then it supplements the keyword recognition stream with short-term tail silence and additional flush silence, thereby improving the recognition stability in tail sound scenarios.
[0054] Through this mechanism, the system can retain more effective recognition information when the user's voice is soft at the end, the end of the segment is too fast, or there is a lot of environmental noise interference, thereby improving the keyword recognition hit rate in the end segment scenario. Example 4
[0055] A length verification method for accurately hitting a path includes: 1. Obtain precise keyword candidates and their corresponding keyword length information; 2. Calculate the minimum acceptable matching length based on keyword categories; 3. Determine whether the exact hit result passes the verification based on the current decoding output matching length.
[0056] The work process / implementation steps are as follows: 1. For example Figure 4 As shown, in step S402, precise keyword candidates are obtained, and in step S403, the keyword length information corresponding to the keyword is read.
[0057] 2. In step S404, the minimum passing matching length is calculated based on the keyword category. In step S405, the matching length of the current decoding output is obtained. In step S406, it is determined whether the matching length has reached the minimum passing matching length.
[0058] 3. If the exact match is achieved, the exact match result is determined to pass the verification in step S407, and the keyword recognition result is output in step S409; if the exact match is not achieved, the exact match result is discarded in step S408, and the recognition result is not output in step S410.
[0059] When the system obtains a precise keyword match result, it does not directly output a recognition event. Instead, it further reads the keyword length information corresponding to that keyword and calculates the minimum matching length based on the keyword's category. The keyword category can include wake word categories and command word categories. Different minimum matching ratios can be configured for different keyword categories.
[0060] The system further obtains the actual number of matched tokens in the current decoding output and determines whether this number of tokens reaches the minimum acceptable matching length. If it does, the exact keyword hit result is considered to have passed the verification; if it does not, the hit is discarded and no recognition event is output.
[0061] In one example, if the total length of the target wake word token is 8 and the minimum matching ratio corresponding to the wake word path is 75%, then the minimum passing matching length is 6 after rounding up; if the current exact hit result only outputs 5 matching tokens, then the result fails the validation. Example 5
[0062] A method for secondary determination of matching depth under inaccurate hit path conditions, comprising: 1. Read the peak depth of context matching for the current speech segment when no precise keyword match results are obtained; 2. Traverse the candidate keyword set and calculate the dynamic judgment threshold for each candidate keyword; 3. Generate candidate fuzzy hit results based on the comparison results of the context matching peak depth and the dynamic judgment threshold, and perform verification.
[0063] The work process / implementation steps are as follows: 1. For example Figure 5 As shown, in step S502, it is confirmed that no accurate keyword hit result has been obtained, and in step S503, the peak depth of context matching of the current speech segment is read.
[0064] 2. In steps S504 to S507, traverse the candidate keyword set, read the candidate keyword length information, calculate the dynamic judgment threshold, and determine whether the peak depth reaches the dynamic judgment threshold.
[0065] 3. If the target is reached, then in steps S508 to S511, candidate fuzzy hit results are generated, keyword category verification is performed, and recognition results are output; if the target is not reached or the verification is not passed, then in steps S512 and S513, the traversal continues or the target is determined to be a miss.
[0066] In some speech segments, user pronunciation may contain weak syllables, slight near-homophones, slight tone deviations, or attenuation of the final consonant, causing the system to fail to obtain complete and accurate keyword candidates. If the system directly judges this as a miss, it will increase the probability of false negatives. To address this, this invention introduces a secondary judgment mechanism based on the peak depth of context matching when an accurate keyword match result is not obtained.
[0067] Throughout the entire speech segment, the system continuously tracks the context matching status corresponding to the current optimal candidate path during keyword decoding and updates the peak depth of context matching within the current speech segment. The peak depth reflects the maximum degree of matching between the current speech and a candidate keyword during the decoding process.
[0068] The peak depth of context matching reflects the maximum extent to which the current speech segment advances the context state of the candidate keywords during the decoding process. Even if the current speech segment does not form a complete and accurate keyword hit result, if the peak depth of context matching has reached a judgment level that is compatible with the length of the candidate keywords, it indicates that the speech segment has a high degree of proximity in temporal matching, and therefore can be used as a basis for judging the generation of candidate fuzzy hit results.
[0069] If the system has not obtained a precise keyword match result after the speech segment ends, it does not immediately determine it as a miss. Instead, it iterates through the candidate keyword set. For each candidate keyword, the system calculates a dynamic judgment threshold based on its keyword length information and determines whether the peak depth of context matching within the current speech segment reaches the dynamic judgment threshold. If it does, the candidate keyword is considered a candidate fuzzy match result.
[0070] In one implementation, the dynamic judgment threshold is determined at least based on the keyword length information of the candidate keywords, and can be adjusted according to the current recognition mode, keyword category, or recognition link status. For shorter keywords, the dynamic judgment threshold can be increased to suppress false judgments; for longer keywords, the dynamic judgment threshold can be decreased to avoid overly strict thresholds leading to missed judgments.
[0071] In one optional implementation, the system may perform keyword category verification on the candidate fuzzy hit results corresponding to the current recognition pattern to determine whether the result meets the pattern requirements in the current recognition chain. The system only outputs the final keyword recognition result if the verification passes.
[0072] In one example, if the total length of a candidate wake word's token is 8, its dynamic judgment threshold can be set to 6 after rounding up. When the peak depth of the context matching recorded in the current speech segment reaches 6, even if no precise keyword hit result is obtained, the candidate wake word can be used as a candidate fuzzy hit result. Example 6
[0073] A method for distinguishing between wake word paths and command word paths includes: 1. Enter wake word recognition mode in the initial state; 2. After successful wake-up word recognition, switch to command word recognition mode and reconstruct the keyword recognition stream; 3. When command recognition is complete or recovery conditions are met, restore to wake word recognition mode and rebuild keyword recognition stream again.
[0074] The work process / implementation steps are as follows: 1. For example Figure 6 As shown, in step S602, the system is powered on or initialized; in step S603, the system enters the wake-up word recognition mode; and in step S604, it is determined whether to output the wake-up word recognition result.
[0075] 2. If so, switch to command word recognition mode in step S605, reconstruct the keyword recognition stream based on the keyword set of the current mode in step S606, and then perform command word recognition in step S607.
[0076] 3. In step S608, determine whether command recognition is completed or whether the recovery condition is met. If it is met, restore the wake word recognition mode in step S609, and reconstruct the keyword recognition stream based on the keyword set of the current mode in step S610.
[0077] In edge voice interaction, wake-up words and command words have different functional roles. Wake-up words are typically used to activate the device and are therefore more sensitive to accidental triggering; command words usually execute specific operations after being woken up, and their recognition path emphasizes continuous interaction and diverse control. Therefore, this invention distinguishes between wake-up word paths and command word paths within the same keyword recognition path.
[0078] In one implementation, the system initially operates in wake word mode, performing primary recognition only on the wake word set. Upon successful wake word recognition, the system switches to command word mode and incorporates the command word set into the current keyword recognition stream. In command word mode, the system can apply a different minimum matching ratio to command words compared to the wake word setting.
[0079] After the command word recognition is completed, the system returns to the wake word recognition mode and reconstructs the keyword recognition stream based on the restored wake word set to prevent the keyword set in the command word mode from remaining in the subsequent recognition link. Example 7
[0080] A method for post-hit state reset and echo suppression includes: 1. Perform a state reset after outputting the keyword recognition results; 2. Reset the speech activity detection status, clear the preamble audio cache, clear the current speech segment statistics status, and reset the keyword recognition stream respectively; 3. Return to the pending recognition state to suppress subsequent interference from the echo of the prompt tone and residual audio.
[0081] The work process / implementation steps are as follows: 1. For example Figure 7 As shown, the keyword recognition result is output in step S701, and the state is reset in step S702.
[0082] 2. In steps S703 to S706, the following steps are executed respectively: resetting the speech activity detection state, clearing the preamble audio cache, clearing the current speech segment statistics state, and resetting the keyword recognition stream.
[0083] 3. In step S707, return to the state to be identified.
[0084] In edge devices, when a wake-up word or command word is successfully recognized, the system typically plays a prompt tone, a response tone, or performs other feedback operations. This feedback tone, after being played through the speaker, may be picked up again by the microphone and misinterpreted as a new speech segment by voice activity detection. If the system does not promptly clear its state after successful recognition, the feedback tone may lead to repeated triggering, segmentation anomalies, or state disorder.
[0085] Therefore, this invention immediately performs a state reset operation after the keyword recognition result is output. The state reset operation may include: resetting the speech activity detection model state; clearing the preceding audio buffer; clearing the peak matching depth, matching count, and other statistical states within the current speech segment; and resetting the keyword recognition stream. In this way, the system can quickly return to a controllable state after a triggered event, thereby reducing the contamination of the next round of recognition process by echoing prompts and residual speech. Example 8
[0086] A dynamic keyword update method includes: 1. During operation, receive wake word text or command word text input by the user; 2. Construct new keyword representations based on the text; 3. Update the keyword set and keyword recognition stream based on the construction results.
[0087] The work process / implementation steps are as follows: 1. During operation, the system receives wake-up word text or command word text input by the user. The keyword update process does not depend on retraining the complete acoustic model, nor on replacing the entire package model or upgrading the complete firmware.
[0088] 2. After receiving the keyword text, the system converts the Chinese keyword into a Pinyin token sequence. For keywords containing polyphonic characters, the system can generate multiple pronunciation combinations and write these multiple pronunciation combinations as multiple keyword representations of the same keyword into the keyword set.
[0089] 3. After the keyword set is updated, the system can reconstruct the keyword recognition stream based on the current recognition mode, so that the newly configured keywords take effect at runtime. Example 9
[0090] A mode switching and recognition stream reconstruction method, comprising: 1. After successful wake-word recognition, switch to command word recognition mode; 2. Reconstruct the keyword recognition stream based on the keyword set corresponding to the current pattern; 3. After the command word recognition is completed, restore to the wake word recognition mode and rebuild the keyword recognition stream again.
[0091] The work process / implementation steps are as follows: 1. After the wake word is successfully recognized, the system switches to the command word recognition mode and reconstructs the keyword recognition stream based on the command word set so that subsequent speech segments are mainly oriented towards command word recognition.
[0092] 2. After the command word recognition is completed, the system returns to the wake word recognition mode and reconstructs the keyword recognition stream based on the restored keyword set.
[0093] 3. In another optional implementation, the system can also receive a runtime-updated set of command words during the listening process, and reconstruct the keyword recognition stream in real time while it is currently in command word recognition mode, so that the updated command words take effect immediately. Example 10
[0094] An audio preprocessing method for weak sound pickup scenarios includes: 1. Perform gain adjustment, noise threshold control, and / or target energy alignment processing on the input audio; 2. Suppress or mute audio segments below the noise threshold; 3. Apply appropriate gain to the effective speech segments and align them with the target energy.
[0095] The work process / implementation steps are as follows: 1. In some devices, due to low microphone sensitivity, insufficient front-end gain, the user being far from the device, or attenuation in the audio link at the end panel level, the overall volume of the input voice may be low. In this case, weak starting syllables and weak ending syllables are more likely to be ignored by the voice activity detection or keyword recognition engine.
[0096] 2. Before the audio input keyword recognition link is processed, the system can perform gain adjustment, noise threshold control and / or target energy alignment on the input audio.
[0097] 3. For audio segments below the noise threshold, the system can reduce their priority in entering the keyword recognition stream or treat them as silent; for valid speech segments, the system can perform appropriate gain and target energy alignment to improve the stability of subsequent keyword recognition. Example 11
[0098] A parameter configuration scheme includes: 1. Configuration of preamble buffer length, silence buffer limit, hangover period, trailing silence, and flush audio; 2. Configuration of minimum pass-through length ratio threshold for precise hit path and dynamic determination threshold for fuzzy hit path; 3. Configuration of the cooling-off period after the recognition results are output.
[0099] The work process / implementation steps are as follows: 1. The preamble buffer length can be configured from 80ms to 1000ms, preferably from 300ms to 900ms; the maximum mute buffer length can be configured from 500ms to 1500ms, preferably from 800ms to 1200ms.
[0100] 2. The hangover period can be configured from 80ms to 800ms, preferably from 300ms to 650ms; the trailing mute can be configured from 100ms to 400ms; the flush audio can be configured from 80ms to 250ms; the minimum pass-match length for the exact hit path can be configured with a proportional threshold according to the keyword category, and the proportional threshold can take any value between 0.50 and 1.00, preferably from 0.65 to 0.85, wherein the proportional threshold for the wake word path can be higher than the proportional threshold for the command word path.
[0101] 3. For fuzzy hit paths, the dynamic judgment threshold can be monotonically adjusted according to the keyword length information to avoid the threshold being too loose for short keywords and too strict for long keywords; the cooling period after the recognition result is output can be configured to be 100ms to 3000ms, preferably 300ms to 1500ms. During the cooling period, repeated triggering of the same keyword or keywords of the same category can be suppressed. Example 12
[0102] A scheme for demonstrating the effects of a collaborative technology includes: 1. Compare this invention with a scheme that uniformly relaxes the trigger threshold; 2. Compare this invention with solutions that only add a leading buffer or a trailing hangover; 3. The present invention is compared with a scheme that performs a general cleanup operation only after the recognition result is output.
[0103] The work process / implementation steps are as follows: 1. In one comparative approach, the hit rate in weak pronunciation scenarios is improved simply by uniformly lowering the keyword trigger threshold. While this approach may improve recall in some slightly near-sound scenarios, it fails to distinguish between precise and inaccurate hit paths, nor does it incorporate keyword length information for stratified judgment. Consequently, it is prone to overly lenient treatment of short keywords or environmental noise segments, leading to a simultaneous increase in false wake-up rate.
[0104] 2. In another comparative scheme, segment boundary issues are improved solely by increasing the preamble buffer or the hangover duration at the end of the syllable. This scheme can alleviate the first syllable or last syllable truncation to some extent, but without subsequent precise path length verification and fuzzy path peak depth determination, it is still difficult to balance missed detection suppression and false trigger control in complex weak pronunciation scenarios. In yet another comparative scheme, a general cleanup operation is performed only after the recognition result is output. Although this scheme can reduce state pollution after a hit to some extent, it does not directly improve the missed detection problems caused by lost speech segment beginnings, insufficient end-segment convergence, and slight near-sounds.
[0105] 3. Unlike the aforementioned local optimization schemes, this invention integrates preamble audio refeedback, tail-segment hangover and silence compensation, accurate hit path length verification, fuzzy hit path peak depth determination, and post-hit state reset into the same end-side streaming recognition link, forming a continuous closed-loop processing around the same speech segment. This closed loop enables the system to compensate for information loss caused by segment boundaries, provide a secondary judgment channel for weak pronunciation scenarios without relaxing the overall trigger threshold, and promptly clear the state and restore the link after a hit, thereby simultaneously improving the hit rate, false trigger rate, and continuous interaction stability.
[0106] It should be noted that the above embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent substitutions, improvements, or variations made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A keyword recognition method for edge-side streaming voice interaction, characterized in that, include: Acquire a continuous audio stream from the device side, establish a keyword recognition stream based on the continuous audio stream, and identify whether the current audio is in a speech segment or a silent segment based on the speech activity detection result; When a speech segment is detected to be transitioning from a silence segment, at least a portion of the preamble audio cached during the silence segment is fed back into the keyword recognition stream; Keyword decoding is continuously performed within the speech segment, and the peak depth of context matching in the current speech segment is recorded based on the changes in context matching status during the decoding process. When the end trend of the speech segment is detected, the continuation decoding is performed and the end-of-segment silence audio is added to the keyword recognition stream; When a precise keyword hit result is obtained, the minimum pass matching length is calculated based on the keyword length information corresponding to the precise keyword, and the length of the precise keyword hit result is verified based on the comparison result between the current decoded output matching length and the minimum pass matching length. When no precise keyword hit result is obtained, a dynamic judgment threshold is calculated based on the context matching peak depth and the keyword length information corresponding to the candidate keyword, and a candidate fuzzy hit result is generated based on the comparison result between the context matching peak depth and the dynamic judgment threshold. When the exact keyword hit result passes the length check, or the candidate fuzzy hit result passes the check, the corresponding keyword recognition result is output; After outputting the keyword recognition results, reset the voice activity detection status and related cache.
2. The keyword recognition method according to claim 1, characterized in that, The keyword recognition results include wake word recognition results and command word recognition results; the method further includes: distinguishing wake word paths and command word paths according to the current recognition mode, and adopting different matching length verification rules for different paths.
3. The keyword recognition method according to claim 1, characterized in that, The preamble audio is historical audio that is continuously cached during the silence phase; the step of feeding back at least a portion of the preamble audio cached during the silence phase into the keyword recognition stream includes: when a speech segment is detected to be entering from a silence segment, selecting a preamble audio of a predetermined length from the historical audio and injecting the preamble audio into the keyword recognition stream.
4. The keyword recognition method according to claim 1, characterized in that, The step of continuing to perform continuation decoding when a speech segment ending trend is detected, and supplementing the keyword recognition stream with tail-end silence audio, includes: after the speech activity detection result changes from a speech segment to a silence segment, continuing to perform keyword decoding within a predetermined hangover period, and supplementing the keyword recognition stream with tail-end silence audio and / or flush audio after the hangover period ends.
5. The keyword recognition method according to claim 1, characterized in that, The step of calculating the minimum pass matching length based on the keyword length information corresponding to the precise keyword includes: calculating the minimum pass matching length based on the total length of the token corresponding to the precise keyword and the category to which the precise keyword belongs.
6. The keyword recognition method according to claim 1, characterized in that, The step of calculating the dynamic judgment threshold based on the peak depth of context matching and the keyword length information corresponding to the candidate keyword includes: determining the dynamic judgment threshold of the candidate keyword based on the total length of the token corresponding to the candidate keyword.
7. The keyword recognition method according to claim 1, characterized in that, The process of generating candidate fuzzy hit results includes: traversing the candidate keyword set and determining the candidate keywords that satisfy the context matching peak depth reaching their corresponding dynamic judgment threshold as candidate fuzzy hit results.
8. The keyword recognition method according to claim 1, characterized in that, The context matching peak depth includes: during multiple decoding processes of the current speech segment, reading the context state level corresponding to the current optimal decoding path, and determining the maximum level read as the context matching peak depth.
9. The keyword recognition method according to claim 1, characterized in that, The verification of the candidate fuzzy hit results includes: performing keyword category verification corresponding to the current recognition mode on the candidate fuzzy hit results.
10. The keyword recognition method according to claim 1, characterized in that, The precise keyword hit results and candidate fuzzy hit results are processed in a hierarchical manner, including: prioritizing length verification based on the precise keyword hit results, and only generating candidate fuzzy hit results based on the cumulative context matching peak depth of the current speech segment when no precise keyword hit results that pass the length verification are obtained in the current speech segment.
11. The keyword recognition method according to claim 1, characterized in that, The resetting of the speech activity detection state and related cache includes: resetting the speech activity detection model state, clearing the leading audio cache, clearing the matching status statistics of the current speech segment, and / or resetting the keyword recognition stream.
12. The keyword recognition method according to claim 1, characterized in that, Also includes: Receive the wake word text or command word text configured at runtime, convert the wake word text or command word text into a keyword representation acceptable to the keyword recognition stream, and update the keyword set and keyword recognition stream based on the conversion result.
13. The keyword recognition method according to claim 12, characterized in that, The step of converting the wake word text or command word text into a keyword representation acceptable to the keyword recognition stream includes: converting Chinese character keywords into pinyin token sequences, and generating multiple pronunciation combinations when polyphonic characters exist, so as to write the multiple pronunciation combinations as multiple keyword representations of the same keyword into the keyword set.
14. The keyword recognition method according to claim 1, characterized in that, Also includes: Perform gain adjustment, noise threshold control, and / or target energy alignment processing on the input speech stream.
15. The keyword recognition method according to claim 1, characterized in that, Also includes: After the keyword recognition results are output, a cooling-off period control is initiated to suppress repeated triggering during the cooling-off period.
16. The keyword recognition method according to claim 1, characterized in that, Also includes: After the wake word is successfully recognized, the system switches to the command word recognition mode and reconstructs the keyword recognition stream based on the new keyword set. After the command word recognition is completed, the system returns to the wake word recognition mode and reconstructs the keyword recognition stream based on the restored keyword set.
17. A keyword recognition system for end-to-end streaming voice interaction, characterized in that, include: Audio receiving module, used to acquire continuous voice stream from the end side; The voice activity detection module is used to identify whether the current audio is in a speech segment or a silent segment based on the voice activity detection results; The preamble buffer and backfeed module is used to backfeed at least a portion of the preamble audio buffered during the silence phase to the keyword recognition stream when a speech segment is detected to be transitioned from a silence segment. The keyword recognition module is used to continuously perform keyword decoding within a speech segment; The depth tracking module is used to record the peak depth of context matching for the current speech segment based on changes in context matching state during decoding. The tail segment compensation module is used to continue performing sustained decoding when the end trend of the speech segment is detected, and to supplement the keyword recognition stream with tail segment silence audio; The precise hit verification module is used to calculate the minimum pass matching length based on the keyword length information corresponding to the precise keyword when a precise keyword hit result is obtained, and to perform length verification on the precise keyword hit result based on the comparison result between the current decoded output matching length and the minimum pass matching length. The fuzzy judgment module is used to calculate a dynamic judgment threshold based on the context matching peak depth and the keyword length information corresponding to the candidate keyword when no accurate keyword hit result is obtained, and to generate a candidate fuzzy hit result based on the comparison result between the context matching peak depth and the dynamic judgment threshold. The result output module is used to output the corresponding keyword recognition result when the exact keyword hit result passes the length check or the candidate fuzzy hit result passes the check. The state reset module is used to reset the speech activity detection state and related cache after the keyword recognition results are output.
18. The keyword recognition system according to claim 17, characterized in that, The system also includes a dynamic keyword update module, which receives the wake word text or command word text configured at runtime, converts it into a keyword representation acceptable to the keyword recognition stream, and updates the keyword set and keyword recognition stream based on the conversion result.
19. An electronic device, characterized in that, The device includes a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, causes the electronic device to perform the keyword recognition method according to any one of claims 1 to 16.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the keyword recognition method according to any one of claims 1 to 16.