Methods, systems, devices, and storage media for post-processing of voice command word recognition
By acquiring the phoneme probability matrix and path score for multi-dimensional judgment, the problems of misidentification, mixed identification, and out-of-set word recognition in offline voice command word recognition on the edge are solved, achieving higher recognition accuracy and anti-interference ability, and is suitable for scenarios with limited offline resources on the edge.
Patent Information
- Application Number
- CN202510933856.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing technologies for offline voice command word recognition on the edge rely too heavily on acoustic models, resulting in insufficient anti-interference capabilities, frequent misidentification, mixed identification, and identification of words outside the set. They also cannot effectively handle the structural correlation of command words and the range of words outside the set, leading to poor recognition accuracy.
By obtaining the phoneme probability matrix and command word path score output by the acoustic model, multi-dimensional identification problem judgment is performed, including secondary confirmation of misidentification, mixed identification and out-of-set word identification. The final identification result is generated by analyzing the phoneme probability matrix and command word structure relationship.
It significantly improves the accuracy and anti-interference ability of voice command word recognition, reduces the risk of misidentification, mixed recognition and recognition of words outside the set, and enhances the robustness and efficiency of recognition.
Smart Images

Figure CN120431904B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech decoding, and in particular to a method, system, device and storage medium for post-processing of speech command word recognition. Background Technology
[0002] In edge-side offline voice command word recognition scenarios, existing technologies, when decoding and recognizing command words by constructing acoustic models, suffer from several problems. These include over-reliance on the acoustic model's performance, the inability of multi-condition recognition mechanisms to meet all conditions under interference, and the significant increase in false recognition rates when conditions are relaxed. When the command word list contains strict prefix-suffix inclusion, approximation, or symmetry structures, the model is prone to mixed recognition due to edge-side resource limitations. Furthermore, in number-related command words, the model often misidentifies non-target words (such as fan speed four, which are not target commands) as target words. The essence of these problems stems from the fact that, under edge-side resource constraints, existing recognition schemes struggle to balance anti-interference capabilities and recognition accuracy. They also lack effective mechanisms for handling command word structural correlations and the range of non-target words, leading to frequent issues of false recognition, mixed recognition, and non-target word recognition. These problems urgently require optimization and solutions through post-processing techniques.
[0003] Therefore, existing technologies, due to their over-reliance on acoustic models and limited end-side resources, are unable to accurately identify command words, leading to frequent recognition errors, which is an urgent problem to be solved. Summary of the Invention
[0004] The main purpose of this application is to provide a method, system, device and storage medium for post-processing of voice command word recognition, which aims to solve the technical problem that the existing technology cannot accurately recognize command words due to over-reliance on acoustic models and limited end-side resources, resulting in frequent recognition errors.
[0005] To achieve the above-mentioned objectives, this application proposes a post-processing method for speech command word recognition, the method comprising: obtaining the phoneme probability matrix and command word path score output by the acoustic model;
[0006] Based on the phoneme probability matrix and command word path score, a preliminary judgment is made as to whether there is misidentification, mixed identification, or identification of words outside the set.
[0007] If a preliminary judgment indicates that there is a misidentification, a score vector is extracted based on the phoneme probability matrix and the number of non-target phonemes is counted to make a judgment and generate the first confirmation result.
[0008] If it is initially determined that there is mixed recognition, the phoneme sequence of subsequent frames or historical frames is obtained from the phoneme probability matrix according to the anomaly type corresponding to the mixed recognition, and the analysis is performed based on the command word structure relationship to obtain the second confirmation result;
[0009] If it is initially determined that there is out-of-set word recognition, the key phoneme with the highest score is extracted from the phoneme probability matrix, the candidate phoneme library is queried, and a third confirmation result is generated.
[0010] Based on the first confirmation result, the second confirmation result, or the third confirmation result, the final identification result is output.
[0011] Furthermore, the step of obtaining the phoneme probability matrix and command word path score output by the acoustic model includes:
[0012] The phoneme probability distributions output by the acoustic model at each time frame are received to form a phoneme probability matrix;
[0013] Based on the phoneme probability matrix, the optimal phoneme path matching the command word list is obtained through a decoding algorithm, and the comprehensive score of the path is calculated.
[0014] Cache the phoneme probability matrix of the current recognition frame and the historical N frames, where N is a positive integer preset according to the maximum phoneme length of the command word, for subsequent delayed recognition or historical frame analysis;
[0015] The command word path score is recorded synchronously for preliminary judgment of the result type.
[0016] Furthermore, the step of initially determining whether there is misidentification, mixed identification, or out-of-set word identification based on the phoneme probability matrix and command word path score includes:
[0017] If the command word path score is lower than the preset misidentification threshold, it is preliminarily determined that there is a misidentification;
[0018] Determine the structural relationship between the current command and each entry in the command list;
[0019] If the current command has a pre-defined structural relationship with any word in the command word list, it is initially determined that there is a misidentification.
[0020] If the current command word contains a number and the path score of the current command word is lower than the threshold for triggering out-of-collection words, it is initially determined that out-of-collection word recognition exists.
[0021] Furthermore, the preset structural relationships include strict prefix inclusion, strict suffix inclusion, non-strict inclusion, approximate inclusion, and symmetric relationships; the step of initially determining that there is a misidentification if the current command and any word in the command word list have a preset structural relationship includes:
[0022] Extract the phoneme sequence path corresponding to the current recognition result;
[0023] The current recognition path is compared with other words in the command word list by prefix. If the phoneme sequence of the short path word completely matches the beginning part of the long path word, then there is strict prefix inclusion, and it is initially determined to be a prefix inclusion type mixed recognition.
[0024] If the final phoneme sequence of the long path word completely matches the short path word, then there is strict suffix inclusion, and it is initially determined to be a suffix inclusion type mixed recognition.
[0025] Analyze whether the current identification path is a non-contiguous subset of other long path terms in the command word list. If so, there is non-strict inclusion, and it is initially determined to be a non-strict inclusion mixed identification.
[0026] Calculate the phoneme sequence similarity between the current recognition path and other entries in the command word list. If the similarity exceeds the preset threshold and there are shared phoneme segments, then there is approximate inclusion, and it is initially determined to be an approximate inclusion type mixed recognition.
[0027] If the difference between the current command word and other words in the command word list is symmetrical, then a symmetrical relationship exists, and it is initially determined to be a symmetrical mixed recognition.
[0028] Furthermore, the step of extracting a score vector from the phoneme probability matrix and counting the number of non-target phonemes to generate a first confirmation result if a preliminary judgment indicates misidentification includes:
[0029] If a preliminary judgment indicates that there is misidentification, extract the phoneme probability values corresponding to all non-target command words from the phoneme probability matrix of the current frame to form a score vector;
[0030] The number of non-target phonemes in the statistical score vector whose probability value is higher than the target phoneme or whose difference from the target phoneme is within a first preset difference;
[0031] If the number exceeds a preset threshold, the current identification result is determined to be a misidentification, and a first confirmation result is generated.
[0032] Furthermore, the step of obtaining the phoneme sequence of subsequent or historical frames from the phoneme probability matrix based on the anomaly type corresponding to the mixed recognition, and analyzing it based on the command word structure relationship to obtain the second confirmation result if a preliminary judgment is made, includes:
[0033] If it is a strict prefix inclusion, when the included word is identified, the subsequent first preset number of frame phoneme sequences are delayed to determine whether there is phoneme information related to the included word. If there is, it is determined to be a mixed recognition and a second confirmation result is generated.
[0034] If it is a strict suffix inclusion, when the included word is identified, the second preset number of frame phoneme sequences in the cached data are obtained, and it is determined whether there is phoneme information related to the included word. If there is, it is determined to be a mixed recognition, and a second confirmation result is generated.
[0035] If it is not strictly included, within the time range of the short path phoneme sequence, identify whether there is historical or subsequent frame information of a specific phoneme of the long path. If it exists, it is determined to be mixed recognition, and a second confirmation result is generated.
[0036] If it is an approximate inclusion, the phoneme sequence of the first preset number of frames is obtained or the phoneme sequence of the second preset number of frames in the cache data is obtained according to the inclusion form, and it is determined whether there is phoneme information related to the included word. If it exists, it is determined to be a mixed recognition and a second confirmation result is generated.
[0037] If the relationship is symmetrical, when the phoneme score of the symmetrical position of the symmetrical word is lower than the first score threshold, the scores of the two phonemes symmetrical in the symmetrical phoneme position are identified. If the score difference is greater than the first score difference threshold, the current recognition result is determined to be a mixed recognition, and a second confirmation result is generated.
[0038] Furthermore, the step of extracting the highest-scoring key phoneme from the phoneme probability matrix, querying the candidate phoneme library, and generating a third confirmation result if it is initially determined that there is out-of-set word recognition, includes:
[0039] Extract the key phonemes associated with the current command word from the phoneme probability matrix and identify the phoneme with the highest score.
[0040] Based on the highest-scoring key phoneme, query the pre-established candidate phoneme library to obtain the scores of the relevant candidate phonemes;
[0041] If the score of the relevant candidate phoneme exceeds the preset candidate threshold or the difference between the score of the candidate phoneme and the key phoneme is less than the set value, the current recognition is determined to be an out-of-collection word recognition, and a third confirmation result is generated.
[0042] The second aspect of this application also includes a voice command word recognition post-processing system, comprising:
[0043] The acquisition module is used to acquire the phoneme probability matrix and command word path score output by the acoustic model;
[0044] The preliminary judgment module is used to make a preliminary judgment on whether there is misidentification, mixed identification, or out-of-set word identification based on the phoneme probability matrix and command word path score;
[0045] The misidentification secondary confirmation module is used to extract the score vector based on the phoneme probability matrix and count the number of non-target phonemes to make a judgment and generate the first confirmation result if the initial judgment is that misidentification exists.
[0046] The mixed recognition secondary confirmation module is used to obtain the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix according to the abnormal type corresponding to the mixed recognition if it is initially determined that mixed recognition exists. It then analyzes the result based on the command word structure relationship to obtain a second confirmation result.
[0047] The secondary confirmation module for out-of-collection words is used to extract the key phoneme with the highest score from the phoneme probability matrix, query the candidate phoneme library, and generate a third confirmation result if it is initially determined that out-of-collection words are identified.
[0048] The output module is used to output the final identification result based on the first confirmation result, the second confirmation result, or the third confirmation result.
[0049] A third aspect of this application also includes a computer device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of any of the methods described above.
[0050] The fourth aspect of this application also includes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.
[0051] This application effectively improves the accuracy and anti-interference capability of voice command word recognition by acquiring phoneme probability matrices and path scores and performing multi-dimensional recognition problem judgment and processing. For misidentification, a secondary confirmation mechanism is initiated, eliminating unreliable recognitions by statistically counting the number of non-target phonemes; for mixed recognition, subsequent or historical frame phoneme sequences are extracted based on the structural relationship of command words to solve the confusion problem caused by different types of structures; for out-of-collection words, key phonemes are compared with a candidate phoneme library to eliminate out-of-bounds recognition related to numbers. This solution does not require model retraining and can efficiently handle the three types of recognition problems using post-processing mechanisms. In scenarios with limited offline resources on the edge, it can balance recognition sensitivity and misidentification rate, and improve the robustness of complex command word lists through structured analysis, significantly reducing the risks of misidentification, mixed recognition, and out-of-collection word recognition, achieving rapid and efficient recognition optimization. Attached Figure Description
[0052] Figure 1 This is a schematic flowchart of a voice command word recognition post-processing method according to an embodiment of this application;
[0053] Figure 2 This is a schematic block diagram of the structure of a voice command word recognition post-processing system according to an embodiment of this application;
[0054] Figure 3 This is a schematic block diagram of the structure of a computer device according to an embodiment of this application.
[0055] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0057] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and or” as used herein includes all or any modules and all combinations of one or more associated listed items.
[0058] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0059] Reference Figure 1 This invention provides a post-processing method for voice command word recognition, including steps S1-S6, specifically:
[0060] S1. Obtain the phoneme probability matrix and command word path score output by the acoustic model;
[0061] S2. Based on the phoneme probability matrix and command word path score, make a preliminary judgment on whether there is misidentification, mixed identification, or identification of words outside the set;
[0062] S3. If a preliminary judgment is made that there is a misidentification, the score vector is extracted based on the phoneme probability matrix and the number of non-target phonemes is counted to make a judgment and generate the first confirmation result.
[0063] S4. If it is initially determined that there is mixed recognition, the phoneme sequence of subsequent frames or historical frames is obtained from the phoneme probability matrix according to the abnormal type corresponding to the mixed recognition, and the analysis is performed based on the command word structure relationship to obtain the second confirmation result.
[0064] S5. If it is initially determined that there is out-of-set word recognition, extract the key phoneme with the highest score from the phoneme probability matrix, query the candidate phoneme library, and generate a third confirmation result.
[0065] S6. Based on the first confirmation result, the second confirmation result, or the third confirmation result, output the final identification result.
[0066] As described in step S1, the system first receives the phoneme probability distribution output by the acoustic model in real time across each time frame via an interface. For example, each frame generates a vector containing the probability values of all phonemes. The vectors from consecutive frames are arranged in chronological order to form a two-dimensional phoneme probability matrix (e.g., with dimensions T×N, where T is the number of frames and N is the number of phoneme categories). Simultaneously, the probability matrix is processed using decoding algorithms such as Viterbi. The optimal phoneme path is matched by traversing the command word list. The path's comprehensive score is calculated by accumulating or multiplying the probability values of each phoneme. To support subsequent delayed recognition and historical frame analysis, the system caches the phoneme probability matrix of the current recognition frame and N historical frames according to preset rules (N is usually set based on the maximum phoneme length of the command word; for example, if the longest command word contains 8 phonemes, then N is set to 10), and synchronously records the path scores of each command word to the cached data structure.
[0067] At the data processing level, the phoneme probability matrix is normalized to eliminate the impact of numerical fluctuations, and algorithms such as Dynamic Time Warping (DTW) can be introduced to optimize matching accuracy when calculating path scores. Among different implementation methods, caching strategies can use a circular buffer to save memory, or dynamically adjust the value of N according to the scenario (e.g., increase N in noisy environments to retain more historical information).
[0068] This step uses a caching mechanism to provide a data foundation for subsequent delayed review analysis of mixed recognition and secondary confirmation of misidentification, avoiding resource consumption caused by repeatedly calling the acoustic model. For example, when processing the "turn on the air conditioner" command, the acoustic model outputs a 10-frame phoneme probability matrix, the decoding algorithm finds the corresponding phoneme path and calculates the score, and caches these 10 frames of data. If subsequent judgments indicate mixed recognition (such as possible confusion with "turn on the air conditioner to dehumidify"), the cached historical frame data can be directly called for subsequent analysis without regenerating the probability matrix, thus improving processing efficiency.
[0069] As described in step S2, the system first establishes a score threshold discrimination model based on statistical regularities: by collecting score distribution data of correct and incorrect recognition of command words in real-world scenarios (such as collecting 1000 sets of recognition samples in noisy environments), a critical threshold for incorrect recognition is fitted (usually set to 0.7, which can be dynamically adjusted through cross-validation). When the comprehensive score of the command word path decoded by the acoustic model is lower than this threshold, a preliminary labeling of incorrect recognition is immediately triggered. For example, when a user issues the command "turn on the lights," if the decoding score is only 0.65 (lower than the threshold of 0.7), the system initially determines that there is a risk of incorrect recognition.
[0070] At the data processing level, the score calculation incorporates smoothing processing in the time dimension: the path scores of 5 consecutive frames are weighted and averaged (the weight of the recent frame is 0.3, and the weights of historical frames decrease sequentially), avoiding misjudgment caused by single-frame noise. In different implementation methods, the threshold can be dynamically adjusted according to the scenario - the threshold in a quiet environment is increased to 0.8 to reduce false recognition, and decreased to 0.6 in a noisy environment to ensure the recognition rate. The false judgment rate in different scenarios is controlled within 5% through an adaptive threshold strategy.
[0071] At the same time, the system pre-stores the structural relationship database of the command word list, and extracts the structural associations between the currently recognized command and other entries in the list through a string matching algorithm, such as calculating the prefix matching degree, suffix overlap rate, and subsequence inclusion relationship of the phoneme sequence.
[0072] Specifically, the defined structural associations include:
[0073] Strict prefix inclusion: that is, one command word is the prefix of another command word, such as "turn on the air conditioner" and "turn on the air conditioner and dehumidify". For example, when "turn on the air conditioner" is recognized, the system extracts its phoneme sequence "dǎ", "kāi", "kōng", "tiáo", and compares it with the first half of the phoneme sequence of "turn on the air conditioner and dehumidify". If it matches exactly, it is determined that there is a risk of mixed recognition of the prefix inclusion type. This process is implemented through a string prefix matching algorithm, and the matching accuracy can be refined to the pronunciation feature parameters of phonemes (such as pitch and length).
[0074] Strict suffix inclusion: that is, one command word is the suffix of another command word, such as "turn off the light" and "turn off the light after ten minutes". For example, the end phoneme sequences of short path words and long path words are compared. Here, short and long are comparative descriptions obtained by comparing phoneme sequences. Taking "turn off the light" as an example, if its phoneme sequence "guān", "dēng" is exactly the same as the end phoneme sequence of "turn off the light after ten minutes", a mixed recognition mark of the suffix inclusion type is triggered. The system achieves efficient matching by traversing the phoneme sequence in reverse, and a minimum matching length threshold (such as 80% of the number of phonemes in the short path word) can be set to reduce misjudgment.
[0075] Non-strict inclusion: that is, the recognition path of one command word is a subset of another command word, such as "lowest brightness" and "brightest". For example: a subsequence matching algorithm is used to analyze whether the currently recognized path is a non-continuous subset of the long path word. For example, if the phoneme sequence "zuì", "liàng" of "brightest" exists in the phoneme sequence "zuì", "dī", "liàng", "dù" of "lowest brightness" in a non-continuous manner, and the path score is lower than the preset threshold (such as 0.7), it is determined as a non-strict inclusion type of mixed recognition. Here, the longest common subsequence (LCS) is calculated through a dynamic programming algorithm. For example, when the proportion of the subsequence length to the number of phonemes in the long path word exceeds 50%, a mark is triggered.
[0076] Approximate inclusion means that two command words are mostly similar, such as "white light" and "turn off the light". The phoneme sequence similarity is calculated using an edit distance algorithm (such as the Levenshtein distance). Taking "white light" and "turn off the light" as an example, the system compares the phoneme sequences "bái", "guāng", "dēng" with "guān", "dēng". If the similarity exceeds 0.8 (determined by calculating the minimum cost of substitution, insertion, and deletion operations) and the length of the shared phoneme segment is ≥ 2, it is marked as an approximate inclusion type misrecognition.
[0077] Symmetry means that the results of two command words show a symmetric relationship in the different parts, such as "play previous song" and "play next song". For command words with a symmetric structure, such as "play previous song" and "play next song", the system first locates the phonemes at the symmetric positions (such as "previous" and "next"), and calculates the score difference between them. If the score difference is less than 0.1 and the path structure shows mirror symmetry (such as the rest of the parts are exactly the same except for the symmetric phonemes), it is determined as a symmetric type misrecognition risk. This process is achieved through a predefined symmetric word list and position index for fast positioning.
[0078] For command words related to numbers (such as fan speed), the system establishes a digital range constraint library: First, identify the digital components in the command word (locate through pronunciations such as "yīèrsān" in the phoneme sequence). If the current command word is "fourth gear" and the fan only supports "first gear to third gear", and the path score ≤ 0.75, a preliminary judgment of out-of-vocabulary words is triggered. For example, when the user misspells "fifth gear", the system detects that the number "five" exceeds the preset range (1 - 3), and at the same time the score of 0.68 is lower than the threshold, and it is preliminarily determined as an out-of-vocabulary word recognition.
[0079] In step S3, accurate secondary confirmation of misrecognition is achieved by quantifying the interference degree of non-target phonemes. When there is a preliminary judgment of misrecognition risk in step S2 (such as the command word path score is lower than 0.7), the system extracts the phoneme probability values corresponding to all non-target command words from the current frame phoneme probability matrix to form a score vector with a dimension of N×1 (N is the total number of phoneme categories). For example, if the phonemes of the target command word "turn on the light" are "kāidēng", the score vector needs to exclude the probability values of these two phonemes and retain the probability data of other phonemes such as "shānguān". During data processing, first perform softmax normalization on the probability matrix to eliminate the numerical fluctuations between different frames, and then quickly filter the target phonemes through a mask matrix to improve the extraction efficiency. Subsequently, the system counts the number of non-target phonemes in the score vector that meet the following conditions:
[0080] The probability value of a single phoneme is higher than the maximum probability of the target phoneme; or the difference between the probability of the single phoneme and the probability of the target phoneme is less than a first preset difference (e.g., 0.15, which can be obtained by fitting historical misidentification data). Taking the identification of "turn on the lights" in a noisy environment as an example, if the acoustic model outputs a phoneme probability of "shān" (mountain) as 0.6 due to noise, while the highest probability of the target phoneme "kāidēng" is 0.55, then "shān" will be counted as an interference phoneme; if the number of similar interference phonemes exceeds a preset threshold (e.g., 2), then the current identification is determined to be a misidentification, and a first confirmation result of rejection (confirming that the current identification is a misidentification) is generated.
[0081] The advantage of this step lies in replacing the traditional full re-identification with "non-target phoneme interference quantification," eliminating noise-induced misidentifications without additional model calculations. For example, when a user says "open the curtains," if the system misidentifies it as "turn on the air conditioner," the probabilities of the non-target phoneme "chuānglián" are 0.6 and 0.5, respectively, while the highest probability of the target phoneme "kōngtiáo" is 0.45. With two interfering phonemes (exceeding the threshold of 1), the system will immediately determine and correct the misidentification, reducing the misidentification rate by 45% in noisy environments. Furthermore, this mechanism does not rely on model iteration; it can adapt to different scenarios using only post-processing rules, achieving lightweight and efficient optimization on edge devices.
[0082] As described in step S4 above, precise disambiguation of mixed recognition is achieved through temporal phoneme analysis based on command word structure relationships, specifically divided into five differentiating processing procedures for anomaly types:
[0083] When a mixed recognition is initially determined in step S2, the phoneme sequence of subsequent frames or historical frames needs to be obtained from the phoneme probability matrix according to its corresponding anomaly type (strict prefix inclusion, strict suffix inclusion, non-strict inclusion, approximate inclusion, and symmetric relationship). Analysis is then performed based on the command word structure relationship to obtain a second confirmation result. Specifically, for strict prefix inclusion (such as "turn on the air conditioner" and "turn on the air conditioner to dehumidify"), when the included word is identified, the phoneme sequence of the next 5 frames is obtained after a delay. A sliding window is used to match the subsequent key phonemes of the included word (such as "chúshī") frame by frame. For example, if there is a match in more than 3 consecutive frames and the single-frame probability is ≥0.6 or the matching degree exceeds 80%, the current recognition result is determined to be mixed recognition. This avoids premature confirmation of short-path words and improves the recognition rate of long-path words through temporal buffering. Compared with the no-delay strategy, the mixed recognition rate of prefix inclusion is reduced.
[0084] For strict suffix inclusion (such as "turn off the lights" and "turn off the lights in ten minutes"), when the included word is identified, the phoneme sequence of the first 5 frames is extracted from the cached historical 10 frames. The reverse KMP algorithm is used to match the prefix phonemes of long path words (such as "shífēnzhōng"). If there are more than 2 consecutive starting phonemes in the historical frame and the probability of a single frame is ≥0.5 or the prefix matching length accounts for more than 60% of the long path word, a correction is triggered. By backtracking through historical frames, the omissions caused by temporal dependence are made up for, thereby improving the recognition accuracy of suffix inclusion scenarios.
[0085] For non-strict inclusion (such as "brightest" and "lowest brightness"), the longest common subsequence of short-path words and long-path words is calculated using dynamic programming. If the subsequence length accounts for more than 50% of the number of phonemes in the long-path word and the sum of the scores of the missing phonemes of the long-path word in subsequent frames exceeds 0.7, it is judged as a mixed recognition. For example, the LCS of "brightest" (phoneme "zuìliàng") and "lowest brightness" ("zuìdīliàngdù") is 2, accounting for 50% of the long-path words. If the scores of "dīdù" in subsequent frames are 0.6 and 0.58 respectively, and the total score exceeds 0.7, it is judged as a mixed recognition. It can be corrected to "lowest brightness" in the later stage. This method can solve the mixed recognition problem caused by non-contiguous subsets. Compared with the scheme that only relies on path scores, the mixed recognition rate of non-strict inclusion is reduced.
[0086] For near-inclusion pairs (such as "white light" and "turn off the light"), after calculating the edit distance, the phoneme sequence is extracted by delaying or backtracking according to the inclusion form. For example, the edit distance between the current recognition path and the near-inclusion pair is calculated (e.g., the Levenshtein distance between "white light" and "turn off the light" is 2). If it is a near-prefix inclusion, the delay is 3 frames; if it is a near-suffix inclusion, the backtracking is 2 frames, and the phoneme sequence is extracted. If a different phoneme appears with a probability ≥ 0.6 or the overall similarity exceeds 0.8, it is corrected. For example, when recognizing "turn off the light", if the edit distance with "white light" is 2, and the phoneme "báiguāng" appears in the delayed frame (probabilities 0.7 and 0.6), it is determined to be a near-inclusion mixed recognition and corrected to "white light". By adapting different near-inclusion structures through dynamic strategies, the recognition accuracy of near-inclusion words is improved.
[0087] For symmetrical mixed recognition (such as "up" and "down" in "previous song" and "next song"), the phoneme probabilities of the two frames before and after the symmetrical position are obtained, and the score difference is calculated. If the symmetrical phoneme score difference is ≤0.2, and the context phoneme paths are completely consistent, bidirectional confirmation is initiated. If the symmetrical direction phoneme score is higher in subsequent frames (e.g., "down" score of 0.7 is higher than "up" score of 0.5), it is corrected to the corresponding direction word. Example: When recognizing "play previous song", the "up" phoneme score is 0.55, and the "down" score is 0.58, with a score difference of 0.03 ≤ 0.2. Furthermore, the context phoneme of "down" matches better in subsequent frames, so it is corrected to "play next song". This can solve the directional misjudgment caused by symmetrical structures and improve the recognition accuracy of symmetrical words.
[0088] This step effectively solves the problem of confusion in model recognition of command words caused by the structural relationship of command words by caching historical frames and delaying the acquisition mechanism, thereby improving the robustness of recognition.
[0089] As described in S5 above, when it is initially determined that there is out-of-collection word recognition (the command word is identified to contain numbers and the path score is lower than the out-of-collection word trigger threshold), the first priority is to eliminate out-of-collection recognition by comparing the key phoneme with the candidate phoneme library. Specifically, the system first extracts the key phoneme associated with the current command word from the phoneme probability matrix (such as the phoneme "sì" corresponding to the number "four"), and determines the phoneme with the highest score. Then, based on the key phoneme, it queries the pre-established candidate phoneme library (such as the pre-stored target phonemes such as "yīèrsān" in the fan speed scenario) to obtain the score of the relevant candidate phonemes. If there is a candidate phoneme with a score exceeding the preset candidate threshold (such as 0.6) or a score difference with the key phoneme that is less than a set value (such as 0.15), it is determined that the current recognition is out-of-collection word recognition (such as recognizing "fourth speed" when the fan only supports three speeds), a third confirmation result confirming out-of-collection word recognition is generated, and the output result of the current model is rejected. For example, when a user mistakenly says "fan speed five", the system extracts the phoneme "wǔ" with a score of 0.5 and the candidate phoneme "sān" with a score of 0.65. Since the difference is 0.15 ≤ the set value, it is judged as an out-of-collection word, avoiding the execution of invalid commands and further reducing the probability of misidentification caused by out-of-collection word recognition.
[0090] As described in step S6, based on the confirmation results of misidentification, mixed identification, and out-of-collection word identification, the final identification result is output. If the final identification is confirmed as misidentification or out-of-collection word identification, the current identification result is rejected; if it is mixed identification, it is corrected to the correct command word to ensure reliable command output.
[0091] In one embodiment, the step of obtaining the phoneme probability matrix and command word path score output by the acoustic model includes:
[0092] S10. Receive the phoneme probability distribution output by the acoustic model in each time frame to form a phoneme probability matrix;
[0093] S11. Based on the phoneme probability matrix, obtain the optimal phoneme path that matches the command word list through a decoding algorithm, and calculate the comprehensive score of the path.
[0094] S12. Cache the phoneme probability matrix of the current recognition frame and the historical N frames, where N is a positive integer preset according to the maximum phoneme length of the command word, which is used for subsequent delayed recognition or historical frame analysis.
[0095] S13. Synchronously record the command word path score for preliminary judgment of result type identification.
[0096] In this embodiment, the phoneme probability distribution output by the acoustic model at each time frame is received in real time via an interface. The probability vector of each frame is arranged in chronological order to form a two-dimensional phoneme probability matrix (e.g., T×N dimension, where T is the number of frames and N is the number of phoneme categories). The matrix is processed using decoding algorithms such as Viterbi, traversing the command word list to match the optimal phoneme path. A comprehensive score is calculated by accumulating or multiplying the probabilities of each phoneme. A preset value of N is used based on the maximum phoneme length of the command word (e.g., N is set to 10 if the longest command word contains 8 phonemes). A circular buffer is used to cache the current and historical N-frame matrices, and path scores are recorded synchronously in the cache structure. During data processing, the matrix is normalized to eliminate fluctuations, and the DTW algorithm is introduced into the score calculation to optimize matching accuracy. The caching strategy can dynamically adjust the value of N (e.g., increasing N in noisy environments) to save memory while retaining sufficient historical information. This step provides a data foundation for subsequent mixed recognition delay analysis and secondary confirmation of misidentification, avoiding repeated model calls. For example, when processing "turn on the air conditioner", 10 frames of matrix are cached. After decoding to obtain the path score, if subsequent judgments are mixed with "turn on the air conditioner to dehumidify", the cached data can be directly called for analysis to improve efficiency.
[0097] In one embodiment, the step of initially determining whether there is misidentification, mixed identification, or out-of-set word identification based on the phoneme probability matrix and command word path score includes:
[0098] S20. If the command word path score is lower than the preset misidentification threshold, it is preliminarily determined that there is misidentification.
[0099] S21. Determine the structural relationship between the current command and each entry in the command word list;
[0100] S22. If the current command has a preset structural relationship with any word in the command word list, it is initially judged that there is a misidentification.
[0101] S23. If the current command word contains a number and the current command word path score is lower than the out-of-set word trigger threshold, it is initially determined that out-of-set word recognition exists.
[0102] In this embodiment, the system first compares the command word path score with a preset misidentification threshold (e.g., 0.7, based on fitting the distribution of correct and misidentified scores in historical data). If the score is lower than the threshold, a preliminary judgment is made that there is a risk of misidentification. This process uses a weighted average of the scores of 5 consecutive frames to smooth out the impact of noise. At the same time, the system extracts the structural relationship between the current command and the command word list through a string matching algorithm, and predefines five modes: strict prefix and suffix inclusion, non-strict inclusion, approximate inclusion, and symmetric relationship. For example, prefix inclusion, such as "turn on the air conditioner" and "turn on the air conditioner to dehumidify", detects whether the short path word completely matches the starting phoneme of the long path word through forward longest matching; suffix inclusion, such as "turn off the lights" and "turn off the lights in ten minutes", compares the ending phoneme sequence in reverse; non-strict inclusion calculates the subsequence matching degree through dynamic programming; approximate inclusion uses edit distance to judge the similarity of phoneme sequences; and symmetric relationship locates the score difference of symmetric phonemes (e.g., "up and down"). If a command word containing numbers (e.g., "four gears") is identified and the score is lower than the out-of-collection word threshold (e.g., 0.75), it is preliminarily determined to be an out-of-collection word in combination with a preset number range (e.g., fan speed 1-3). This mechanism quickly screens for anomalies using multi-dimensional rules. For example, when a user says "fan speed 5", the score of 0.68 is below the threshold and the number is out of bounds. A preliminary judgment is completed within 0.05 seconds, reducing computing power consumption by 40% compared to full model analysis.
[0103] In one embodiment, the preset structural relationship includes strict prefix inclusion, strict suffix inclusion, non-strict inclusion, approximate inclusion, and symmetric relationship; the step of initially determining that there is a misidentification if the current command and any word in the command word list have a preset structural relationship includes:
[0104] S30. Extract the phoneme sequence path corresponding to the current recognition result;
[0105] S31. Compare the current recognition path with other words in the command word list using prefixes. If the phoneme sequence of the short path word completely matches the beginning part of the long path word, then there is strict prefix inclusion, and it is initially determined to be a prefix inclusion type mixed recognition.
[0106] S32. If the final phoneme sequence of the long path word completely matches the short path word, then there is strict suffix inclusion, and it is initially determined to be a suffix inclusion type mixed recognition.
[0107] S33. Analyze whether the current identification path is a non-contiguous subset of other long path terms in the command word list. If so, there is non-strict inclusion, and it is initially determined to be a non-strict inclusion mixed identification.
[0108] S34. Calculate the phoneme sequence similarity between the current recognition path and other words in the command word list. If the similarity exceeds the preset threshold and there are shared phoneme segments, then there is approximate inclusion, and it is initially determined to be an approximate inclusion type mixed recognition.
[0109] S35. If the difference between the current command word and other words in the command word list is symmetrical, then there is a symmetrical relationship, and it is initially determined to be a symmetrical mixed recognition.
[0110] In this embodiment, the system extracts five structural relationships from the command word list: strict prefix inclusion, strict suffix inclusion, non-strict inclusion, approximate inclusion, and symmetric relationship. After extracting the phoneme sequence path corresponding to the current recognition result, a multi-dimensional algorithm is used for preliminary judgment of mixed recognition. During prefix comparison, if the phoneme sequence of the short path word completely matches the beginning part of the long path word, such as the phoneme sequence of "turn on the air conditioner" completely matching the first half of "turn on the air conditioner to dehumidify," then strict prefix inclusion exists, and it is initially judged as prefix inclusion mixed recognition. If the end phoneme sequence of the long path word completely matches the short path word, such as the phoneme sequence of "turn off the lights" being consistent with the end part of "turn off the lights after ten minutes," then strict suffix inclusion exists, and it is initially judged as suffix inclusion mixed recognition. The system analyzes whether the current recognition path is a non-contiguous subset of other long path words. For example, if the phoneme sequence of "brightest" is a non-contiguous subset of the phoneme sequence of "lowest brightness," it is initially judged as non-strict inclusion mixed recognition. By calculating the similarity of phoneme sequences, if the similarity of the phoneme sequences of "white light" and "turn off the light" exceeds a preset threshold and there are shared phoneme segments, then there is approximate inclusion, and it is initially judged as an approximate inclusion-type mixed recognition. If the difference between the current command word and other entries has a symmetrical structure, such as the difference between "play previous song" and "play next song" where the "up" and "down" parts have a symmetrical structure, then there is a symmetrical relationship, and it is initially judged as a symmetrical mixed recognition. This process uses forward longest matching, backward matching, dynamic programming to calculate the longest common subsequence, and edit distance algorithms, and can complete the structural analysis within 50ms on the device side, which can mark the risk of mixed recognition in advance, improving the detection efficiency compared to traditional models.
[0111] In one embodiment, the step of generating a first confirmation result by extracting a score vector from the phoneme probability matrix and counting the number of non-target phonemes if a preliminary judgment is made includes:
[0112] S40. If it is initially determined that there is misidentification, extract the phoneme probability values corresponding to all non-target command words from the phoneme probability matrix of the current frame to form a score vector.
[0113] S41. Count the number of non-target phonemes in the score vector whose probability value is higher than the target phoneme or whose difference with the target phoneme is within a first preset difference.
[0114] S42. If the quantity exceeds a preset quantity threshold, the current identification result is determined to be a misidentification, and a first confirmation result is generated.
[0115] In this embodiment, when a misidentification is initially determined, the system extracts the phoneme probability values corresponding to all non-target command words from the current frame's phoneme probability matrix. After filtering the target phonemes using a mask matrix, a score vector is formed. For example, if the phoneme of the target command word "turn on the light" is "kāidēng", the score vector retains the probability data of other phonemes such as "shānguān", and the probability matrix is subjected to softmax normalization to eliminate numerical fluctuations. Subsequently, the number of non-target phonemes in the score vector that meet the conditions is counted: if the probability of a single phoneme is higher than the maximum probability of the target phoneme, or the difference between the probability of the single phoneme and the target phoneme is less than a first preset difference (e.g., 0.15, based on fitting historical misidentification data), then it is counted as an interference phoneme. When this number exceeds a preset threshold (e.g., 2, which can be dynamically adjusted according to the scenario), it is determined to be a misidentification and a first confirmation result of refusal to execute is generated.
[0116] In terms of implementation, the threshold strategy supports dynamic adaptation: in quiet scenes, the quantity threshold is set to 1 for strict filtering, while in noisy scenes it is relaxed to 3 to avoid missed recognition; the difference threshold can be adjusted in real time in conjunction with the signal-to-noise ratio (e.g., the difference threshold increases by 0.05 for every 5dB decrease in the signal-to-noise ratio). The statistical algorithm uses parallel vector operations for acceleration, and the processing time per frame is controlled within 10ms. This mechanism replaces the traditional full re-recognition by quantifying the interference level of non-target phonemes, eliminating noise misjudgments without additional model calculations. For example, when a user says "open the curtains," the system misidentifies it as "turn on the air conditioner." If the probabilities of the non-target phoneme "chuānglián" are 0.6 and 0.5, respectively, while the highest probability of the target phoneme is 0.45, and the number of interfering phonemes reaches 2, exceeding the threshold, the system immediately determines misrecognition, reducing the misrecognition rate in noisy environments by 45%, and can adapt to different scenarios without relying on model iteration, achieving lightweight and efficient optimization on the device side.
[0117] In one embodiment, the step of obtaining the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix based on the anomaly type corresponding to the mixed recognition, and analyzing it based on the command word structure relationship to obtain a second confirmation result, if it is initially determined that there is mixed recognition, and then performing analysis to obtain a second confirmation result, includes:
[0118] S50. If it is a strict prefix inclusion, when the included word is identified, the subsequent first preset number of frame phoneme sequences are obtained after a delay. It is determined whether there is phoneme information related to the included word. If there is, it is determined to be a mixed recognition and a second confirmation result is generated.
[0119] S51. If it is a strict suffix inclusion, when the included word is identified, the second preset number of frame phoneme sequences in the cached data are obtained, and it is determined whether there is phoneme information related to the included word. If there is, it is determined to be a mixed recognition, and a second confirmation result is generated.
[0120] S52. If it is not strictly included, within the time range of the short path phoneme sequence, identify whether there is historical or subsequent frame information of a specific phoneme of the long path. If it exists, it is determined to be mixed recognition, and a second confirmation result is generated.
[0121] S53. If it is an approximate inclusion, according to the inclusion form, delay the acquisition of the subsequent first preset number of frame phoneme sequences or the second preset number of frame phoneme sequences in the cached data, determine whether there is included word-related phoneme information, if there is, determine it as mixed recognition, and generate a second confirmation result.
[0122] S54. If it is a symmetrical relationship, when the phoneme score of the symmetrical position of the symmetrical word is lower than the first score threshold, the scores of the two phonemes with symmetrical phoneme positions are identified. If the score difference is greater than the first score difference threshold, the current recognition result is determined to be a mixed recognition, and a second confirmation result is generated.
[0123] In this embodiment, when a preliminary judgment indicates the presence of mixed recognition, the system retrieves temporal phoneme sequences from the phoneme probability matrix according to the anomaly type for analysis to generate a second confirmation result. For strict prefix inclusion, when the included word is identified, the subsequent frame phoneme sequence is retrieved with a delay, and the key phonemes of the included word are matched; for strict suffix inclusion, the historical frame phoneme sequence is retrieved, and the prefix phonemes of long-path words are traced back; for non-strict inclusion within the short-path time range, the historical or subsequent frame information of specific phonemes in the long path is detected; for approximate inclusion, the phoneme sequence is delayed or traced back according to the inclusion form, and the difference phonemes are matched; in symmetrical relationships, when the symmetrical phoneme score is lower than the threshold, the difference in symmetrical phoneme scores is compared. This mechanism utilizes cached historical frame data and uses lightweight algorithms such as DTW temporal alignment and reverse matching to complete the analysis within 50ms on the device side. For example, in smart homes, the delayed recognition of "turn on the air conditioner" avoids misjudging as "turn on the air conditioner to dehumidify," significantly reducing the mixed recognition rate and adapting to complex command word structures without model iteration.
[0124] In one embodiment, the step of extracting the highest-scoring key phoneme from the phoneme probability matrix, querying the candidate phoneme library, and generating a third confirmation result if it is initially determined that there is out-of-set word recognition includes:
[0125] S60. Extract the key phonemes associated with the current command word from the phoneme probability matrix and determine the phoneme with the highest score.
[0126] S61. Based on the key phoneme with the highest score, query the pre-established candidate phoneme library to obtain the score of the relevant candidate phoneme.
[0127] S62. If the score of the relevant candidate phoneme exceeds the preset candidate threshold or the difference between the score of the candidate phoneme and the key phoneme is less than the set value, then the current recognition is determined to be an out-of-collection word recognition, and a third confirmation result is generated.
[0128] In this embodiment, when it is initially determined that there is out-of-set word recognition, the system extracts the key phoneme associated with the current command word from the phoneme probability matrix (such as the phoneme "sì" corresponding to the number "four"), and determines the phoneme with the highest score. Based on this phoneme, it queries a pre-established candidate phoneme library (such as the target phonemes "yīèrsān" pre-stored in the fan speed scenario) to obtain the score of the candidate phoneme. If the score of the candidate phoneme exceeds a preset threshold (such as 0.6) or the difference between the candidate phoneme score and the key phoneme score is less than a set value (such as 0.15), it is determined to be out-of-set word recognition (such as recognizing "fourth speed" when the fan only supports three speeds). The threshold can be dynamically adjusted according to the environment (such as relaxing the difference to 0.2 in noisy environments). Lightweight rule matching replaces model iteration, reducing the computational power consumption on the edge. For example, when a user mistakenly says "fan five speed", the extracted phoneme "wǔ" scores 0.5, and the candidate phoneme "sān" scores 0.65. Since the difference is 0.15≤set value, it is judged as an out-of-collection word, avoiding the execution of invalid instructions, thus reducing the out-of-collection word recognition rate by more than 60%.
[0129] Reference Figure 2 This is a block diagram of a voice command word recognition post-processing system according to an embodiment of this application. The device includes:
[0130] The acquisition module 100 is used to acquire the phoneme probability matrix and command word path score output by the acoustic model;
[0131] The preliminary judgment module 200 is used to make a preliminary judgment on whether there is misidentification, mixed identification, or out-of-set word identification based on the phoneme probability matrix and command word path score.
[0132] The misidentification secondary confirmation module 300 is used to extract the score vector based on the phoneme probability matrix and count the number of non-target phonemes to make a judgment and generate a first confirmation result if a misidentification is initially judged.
[0133] The mixed recognition secondary confirmation module 400 is used to obtain the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix according to the abnormal type corresponding to the mixed recognition if it is initially determined that mixed recognition exists, and to analyze it based on the command word structure relationship to obtain the second confirmation result.
[0134] The second confirmation module for out-of-collection words 500 is used to extract the key phoneme with the highest score from the phoneme probability matrix, query the candidate phoneme library, and generate a third confirmation result if it is initially determined that out-of-collection words exist.
[0135] The output module 600 is used to output the final recognition result based on the first confirmation result, the second confirmation result, or the third confirmation result.
[0136] In one embodiment, the acquisition module 100 includes:
[0137] The probability matrix generation unit is used to receive the phoneme probability distribution output by the acoustic model at each time frame and form a phoneme probability matrix.
[0138] The optimal path acquisition unit is used to obtain the optimal phoneme path matching the command word list based on the phoneme probability matrix and through a decoding algorithm, and to calculate the comprehensive score of the path.
[0139] The historical frame buffer unit is used to buffer the phoneme probability matrix of the current recognition frame and the historical N frames, where N is a positive integer preset according to the maximum phoneme length of the command word, for subsequent delayed recognition or historical frame analysis.
[0140] The scoring recording unit is used to synchronously record the command word path score, which is used for preliminary judgment of the result type.
[0141] In one embodiment, the preliminary judgment module 200 includes:
[0142] The misidentification judgment unit is used to preliminarily determine that there is a misidentification if the command word path score is lower than the preset misidentification threshold;
[0143] The structural relationship judgment unit is used to determine the structural relationship between the current command and each entry in the command word list;
[0144] The mixed recognition preliminary judgment unit is used to preliminarily judge that there is mixed recognition if the current command has a preset structural relationship with any word in the command word list;
[0145] The out-of-collection word judgment unit is used to initially judge that out-of-collection word recognition exists if the current command word contains numbers and the current command word path score is lower than the out-of-collection word trigger threshold.
[0146] In one embodiment, the mixed identification preliminary judgment unit includes:
[0147] The phoneme sequence extraction subunit is used to extract the phoneme sequence path corresponding to the current recognition result;
[0148] The prefix matching subunit is used to perform prefix matching between the current recognition path and other entries in the command word list. If the phoneme sequence of the short path word completely matches the beginning part of the long path word, then there is strict prefix inclusion, and it is initially determined to be a prefix inclusion type mixed recognition.
[0149] The suffix matching subunit is used to identify suffix inclusion if the last phoneme sequence of a long path word completely matches that of a short path word.
[0150] The non-continuous subset analysis subunit is used to analyze whether the current recognition path is a non-continuous subset of other long path words in the command word list. If so, there is non-strict inclusion, and it is initially determined to be a non-strict inclusion mixed recognition.
[0151] The similarity calculation subunit is used to calculate the phoneme sequence similarity between the current recognition path and other entries in the command word list. If the similarity exceeds the preset threshold and there are shared phoneme segments, then there is approximate inclusion, and it is initially determined to be an approximate inclusion type mixed recognition.
[0152] The symmetry structure recognition subunit is used to identify a symmetry relationship between the current command word and other words in the command word list if the difference between them is symmetrical. This is initially determined to be a symmetry-type mixed recognition.
[0153] In one embodiment, the misidentification secondary confirmation module 300 includes:
[0154] The non-target phoneme extraction unit is used to extract the phoneme probability values corresponding to all non-target command words from the phoneme probability matrix of the current frame if a preliminary judgment indicates that there is misidentification, and form a score vector.
[0155] The quantity statistics unit is used to count the number of non-target phonemes in the score vector whose probability value is higher than the target phoneme or whose difference from the target phoneme is within a first preset difference.
[0156] The misidentification determination unit is used to determine that the current identification result is a misidentification if the number exceeds a preset number threshold, and generate a first confirmation result.
[0157] In one embodiment, the mixed identification secondary confirmation module 400 includes:
[0158] The prefix inclusion processing unit is used to delay acquiring the subsequent first preset number of frame phoneme sequences when the included word is identified if it is a strict prefix inclusion, to determine whether there is phoneme information related to the included word, and if so, to determine that it is a mixed recognition and to generate a second confirmation result.
[0159] The suffix inclusion processing unit is used to, if it is a strict suffix inclusion, when the included word is identified, obtain the second preset number of frame phoneme sequences in the cached data, determine whether there is phoneme information related to the included word, and if so, determine that it is a mixed recognition and generate a second confirmation result.
[0160] The non-strict inclusion processing unit is used to identify whether there is historical or subsequent frame information of a specific phoneme of a long path within the time range of the short path phoneme sequence if it is non-strict inclusion. If it exists, it is determined to be mixed recognition and a second confirmation result is generated.
[0161] The approximate inclusion processing unit is used to, if it is an approximate inclusion, delay the acquisition of a subsequent first preset number of frame phoneme sequences or the acquisition of a second preset number of frame phoneme sequences in the cached data according to the inclusion form, determine whether there is included word-related phoneme information, and if so, determine that it is a mixed recognition and generate a second confirmation result.
[0162] The symmetry relationship processing unit is used to identify the scores of two phonemes whose positions are symmetrical if the phoneme score of the symmetrical position of the symmetrical word is lower than the first score threshold. If the score difference is greater than the first score difference threshold, the current recognition result is determined to be a mixed recognition and a second confirmation result is generated.
[0163] In one embodiment, the out-of-set word recognition secondary confirmation module 500 includes:
[0164] The key phoneme extraction unit is used to extract key phonemes associated with the current command word from the phoneme probability matrix and identify the phoneme with the highest score.
[0165] The candidate phoneme query unit is used to query a pre-established candidate phoneme library based on the key phoneme with the highest score, and obtain the score of the relevant candidate phoneme.
[0166] The out-of-collection word determination unit is used to determine that the current recognition is an out-of-collection word recognition if the score of the relevant candidate phoneme exceeds the preset candidate threshold or the difference between the score of the candidate phoneme and the key phoneme is less than a set value, and to generate a third confirmation result.
[0167] Reference Figure 3 This application also provides a computer device, which may be a server, and its internal structure may be as follows: Figure 3As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data used in the voice command word recognition post-processing method. The network interface allows communication with external terminals via a network connection. Furthermore, the computer device may also include input devices and a display screen. When the aforementioned computer program is executed by a processor, it implements a post-processing method for speech command word recognition, including the following steps: obtaining the phoneme probability matrix and command word path score output by the acoustic model; preliminarily determining whether there is misidentification, mixed recognition, or out-of-collection word recognition based on the phoneme probability matrix and command word path score; if misidentification is preliminarily determined, extracting the score vector based on the phoneme probability matrix and counting the number of non-target phonemes to generate a first confirmation result; if mixed recognition is preliminarily determined, obtaining the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix according to the anomaly type corresponding to the mixed recognition, analyzing it based on the command word structure relationship to obtain a second confirmation result; if out-of-collection word recognition is preliminarily determined, extracting the key phoneme with the highest score from the phoneme probability matrix, querying the candidate phoneme library, and generating a third confirmation result; and outputting the final recognition result based on the first confirmation result, the second confirmation result, or the third confirmation result. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment on which the present application is applied.
[0168] One embodiment of this application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements a post-processing method for speech command word recognition, including the following steps: obtaining a phoneme probability matrix and command word path score output by an acoustic model; preliminarily determining whether there is misidentification, mixed recognition, or out-of-collection word recognition based on the phoneme probability matrix and command word path score; if misidentification is preliminarily determined, extracting a score vector from the phoneme probability matrix and counting the number of non-target phonemes to generate a first confirmation result; if mixed recognition is preliminarily determined, obtaining the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix according to the anomaly type corresponding to the mixed recognition, analyzing it based on the command word structure relationship, and obtaining a second confirmation result; if out-of-collection word recognition is preliminarily determined, extracting the key phoneme with the highest score from the phoneme probability matrix, querying a candidate phoneme library, and generating a third confirmation result; and outputting the final recognition result based on the first confirmation result, the second confirmation result, or the third confirmation result. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media provided in this application and in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0170] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0171] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for post-processing speech command word recognition, characterized in that, include: Obtain the phoneme probability matrix and command word path score output by the acoustic model; Based on the phoneme probability matrix and command word path score, a preliminary judgment is made as to whether there is misidentification, mixed identification, or identification of words outside the set. If a preliminary judgment indicates that there is a misidentification, a score vector is extracted based on the phoneme probability matrix and the number of non-target phonemes is counted to make a judgment and generate the first confirmation result. If it is initially determined that there is mixed recognition, the phoneme sequence of subsequent frames or historical frames is obtained from the phoneme probability matrix according to the anomaly type corresponding to the mixed recognition, and the analysis is performed based on the command word structure relationship to obtain the second confirmation result; If it is initially determined that there is out-of-set word recognition, the key phoneme with the highest score is extracted from the phoneme probability matrix, the candidate phoneme library is queried, and a third confirmation result is generated. Based on the first confirmation result, the second confirmation result, or the third confirmation result, output the final identification result; The step of initially determining whether there is misidentification, mixed identification, or out-of-set word identification based on the phoneme probability matrix and command word path score includes: If the command word path score is lower than the preset misidentification threshold, it is preliminarily determined that there is a misidentification; Determine the structural relationship between the current command and each entry in the command list; If the current command has a pre-defined structural relationship with any word in the command word list, it is initially determined that there is a misidentification. If the current command word contains a number and the path score of the current command word is lower than the threshold for triggering out-of-collection words, it is initially determined that out-of-collection word recognition exists.
2. The post-processing method for voice command word recognition according to claim 1, characterized in that, The steps of obtaining the phoneme probability matrix and command word path score output by the acoustic model include: The phoneme probability distributions output by the acoustic model at each time frame are received to form a phoneme probability matrix; Based on the phoneme probability matrix, the optimal phoneme path matching the command word list is obtained through a decoding algorithm, and the comprehensive score of the optimal phoneme path is calculated. Cache the phoneme probability matrix of the current recognition frame and the historical N frames, where N is a positive integer preset according to the maximum phoneme length of the command word, for subsequent delayed recognition or historical frame analysis; The command word path score is recorded synchronously for preliminary judgment of the result type.
3. The post-processing method for voice command word recognition according to claim 1, characterized in that, The preset structural relationships include strict prefix inclusion, strict suffix inclusion, non-strict inclusion, approximate inclusion, and symmetric relationships; the step of initially determining that there is a misidentification if the current command and any word in the command word list have a preset structural relationship includes: Extract the phoneme sequence path corresponding to the current recognition result; The current recognition path is compared with other words in the command word list by prefix. If the phoneme sequence of the short path word completely matches the beginning part of the long path word, then there is strict prefix inclusion, and it is initially determined to be a prefix inclusion type mixed recognition. If the final phoneme sequence of the long path word completely matches the short path word, then there is strict suffix inclusion, and it is initially determined to be a suffix inclusion type mixed recognition. Analyze whether the current identification path is a non-contiguous subset of other long path terms in the command word list. If so, there is non-strict inclusion, and it is initially determined to be a non-strict inclusion mixed identification. Calculate the phoneme sequence similarity between the current recognition path and other entries in the command word list. If the similarity exceeds the preset threshold and there are shared phoneme segments, then there is approximate inclusion, and it is initially determined to be an approximate inclusion type mixed recognition. If the difference between the current command word and other words in the command word list is symmetrical, then a symmetrical relationship exists, and it is initially determined to be a symmetrical mixed recognition.
4. The post-processing method for voice command word recognition according to claim 1, characterized in that, The step of generating a first confirmation result by extracting a score vector from the phoneme probability matrix and counting the number of non-target phonemes if a preliminary judgment is made includes: If a preliminary judgment indicates that there is misidentification, extract the phoneme probability values corresponding to all non-target command words from the phoneme probability matrix of the current frame to form a score vector; The number of non-target phonemes in the statistical score vector whose probability value is higher than the target phoneme or whose difference from the target phoneme is within a first preset difference; If the number of non-target phonemes exceeds a preset threshold, the current recognition result is determined to be a misrecognition, and a first confirmation result is generated.
5. The post-processing method for voice command word recognition according to claim 1, characterized in that, The step of obtaining the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix based on the anomaly type corresponding to the mixed recognition, and analyzing it based on the command word structure relationship to obtain the second confirmation result, includes: If it is a strict prefix inclusion, when the included word is identified, the subsequent first preset number of frame phoneme sequences are delayed to determine whether there is phoneme information related to the included word. If there is, it is determined to be a mixed recognition and a second confirmation result is generated. If it is a strict suffix inclusion, when the included word is identified, the second preset number of frame phoneme sequences in the cached data are obtained, and it is determined whether there is phoneme information related to the included word. If there is, it is determined to be a mixed recognition, and a second confirmation result is generated. If it is not strictly included, within the time range of the short path phoneme sequence, identify whether there is historical or subsequent frame information of a specific phoneme of the long path. If it exists, it is determined to be mixed recognition, and a second confirmation result is generated. If it is an approximate inclusion, the phoneme sequence of the first preset number of frames is obtained or the phoneme sequence of the second preset number of frames in the cache data is obtained according to the inclusion form, and it is determined whether there is phoneme information related to the included word. If it exists, it is determined to be a mixed recognition and a second confirmation result is generated. If the relationship is symmetrical, when the phoneme score of the symmetrical position of the symmetrical word is lower than the first score threshold, the scores of the two phonemes symmetrical in the symmetrical phoneme position are identified. If the score difference is greater than the first score difference threshold, the current recognition result is determined to be a mixed recognition, and a second confirmation result is generated.
6. The post-processing method for voice command word recognition according to claim 1, characterized in that, The steps of extracting the highest-scoring key phoneme from the phoneme probability matrix, querying the candidate phoneme library, and generating a third confirmation result if it is initially determined that there is out-of-set word recognition include: Extract the key phonemes associated with the current command word from the phoneme probability matrix and identify the phoneme with the highest score. Based on the highest-scoring key phoneme, query the pre-established candidate phoneme library to obtain the scores of the relevant candidate phonemes; If the score of the relevant candidate phoneme exceeds the preset candidate threshold or the difference between the score of the candidate phoneme and the key phoneme is less than the set value, the current recognition is determined to be an out-of-collection word recognition, and a third confirmation result is generated.
7. A voice command word recognition post-processing system, characterized in that, include: The acquisition module is used to acquire the phoneme probability matrix and command word path score output by the acoustic model; The preliminary judgment module is used to make a preliminary judgment on whether there is misidentification, mixed identification, or out-of-set word identification based on the phoneme probability matrix and command word path score; The misidentification secondary confirmation module is used to extract the score vector based on the phoneme probability matrix and count the number of non-target phonemes to make a judgment and generate the first confirmation result if the initial judgment is that misidentification exists. The mixed recognition secondary confirmation module is used to obtain the phoneme sequence of subsequent frames or historical frames from the phoneme probability matrix according to the abnormal type corresponding to the mixed recognition if it is initially determined that mixed recognition exists. It then analyzes the result based on the command word structure relationship to obtain a second confirmation result. The secondary confirmation module for out-of-collection words is used to extract the key phoneme with the highest score from the phoneme probability matrix, query the candidate phoneme library, and generate a third confirmation result if it is initially determined that out-of-collection words are identified. The output module is used to output the final recognition result based on the first confirmation result, the second confirmation result, or the third confirmation result; The preliminary judgment module includes: The misidentification judgment unit is used to preliminarily determine that there is a misidentification if the command word path score is lower than the preset misidentification threshold; The structural relationship judgment unit is used to determine the structural relationship between the current command and each entry in the command word list; The mixed recognition preliminary judgment unit is used to preliminarily judge that there is mixed recognition if the current command has a preset structural relationship with any word in the command word list; The out-of-collection word judgment unit is used to initially judge that out-of-collection word recognition exists if the current command word contains numbers and the current command word path score is lower than the out-of-collection word trigger threshold.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
User-defined command word recognition method and device, and computer equipment
CN113506574A
Training method of command word recognition model and command word recognition method and device
CN116778914A