Misidentification processing method, device, equipment and storage medium
By scoring and verifying the command words in speech recognition technology, the problem of high misrecognition rate in the prior art is solved, and higher recognition accuracy and system performance are achieved.
Patent Information
- Application Number
- CN202510231151.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-28
AI Technical Summary
When faced with colloquial command words, existing speech recognition technology lacks a consistency verification mechanism, resulting in a high rate of misrecognition.
A misidentification processing method is proposed. By scoring each entry in the command word list, determining the command word to be identified, and independently scoring its prefix and suffix parts, it is determined to determine whether the score difference exceeds the preset threshold. If not exceeded, the alignment process is performed to calculate the score difference of each phoneme. If it is lower than the second preset threshold, the command word to be identified is executed.
It significantly improves the accuracy of speech recognition, reduces the phenomenon of misrecognition, enhances the recognition ability of short command words, and improves the overall performance of the system.
Smart Images

Figure CN119724191B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a method, device, equipment and storage medium for handling misrecognition. Background Art
[0002] At present, speech recognition in smart homes and smart terminal devices mainly relies on the CTC (Connectionist Temporal Classification) algorithm for speech recognition. However, since it only relies on the output results to judge command words, it leads to a high misrecognition rate. In addition, command words are usually short and tend to be colloquial, which further increases the possibility of misrecognition. For example, when a user says "open the box", "open the box" does not belong to the vocabulary in the preset command word list, but it may be mistakenly recognized as "open settings" in the command word list. Therefore, the existing methods lack a verification mechanism for the internal consistency of command words, which is prone to misjudgment.
[0003] Therefore, the existing speech recognition only relies on CTC decoding results to judge command words and lacks a consistency verification mechanism. Faced with the recognition environment of spoken command words, it is easy to lead to a high misrecognition rate, which is a technical problem that needs to be solved. Summary of the invention
[0004] The main purpose of this application is to provide a method, device, equipment and storage medium for handling misrecognition, which aims to solve the technical problem that current speech recognition only relies on CTC decoding results to judge command words, lacks a consistency verification mechanism, and is prone to a high misrecognition rate in the face of a recognition environment of spoken command words.
[0005] In order to achieve the above-mentioned invention object, the present application proposes a method for processing misidentification, the method comprising:
[0006] Score each entry in the command word list according to the CTC decoding rules to determine the command word to be recognized;
[0007] Scoring the prefix and suffix of the command word to be recognized;
[0008] Determine whether the score difference between the prefix part and the suffix part exceeds a first preset threshold;
[0009] If it does not exceed the preset threshold, the command word to be recognized is aligned to obtain the score difference of each phoneme;
[0010] If the score difference is lower than a second preset threshold, the command word to be recognized is executed.
[0011] Furthermore, the step of scoring each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized includes:
[0012] Extract the feature vector of the current audio signal and input it into the pre-trained CTC model to obtain a probability matrix;
[0013] Calculate the path score of the phoneme sequence corresponding to each command word using a forward algorithm based on the probability matrix;
[0014] Using a dynamic programming method to find the best path for each command word, and obtaining the total score of the phoneme sequence corresponding to the command word;
[0015] All command words are sorted according to the total scores, and a preset number of command words are selected according to the sorting as command words to be recognized.
[0016] Furthermore, the step of scoring the prefix part and the suffix part of the command word to be recognized includes:
[0017] Based on the command word to be recognized, extract the corresponding phoneme sequence;
[0018] According to the structure of the command word to be recognized, determining the phoneme boundaries of the prefix part and the suffix part;
[0019] Using the probability matrix output by the CTC model, the path score of the prefix partial phoneme sequence and the path score of the suffix partial phoneme sequence are calculated.
[0020] Furthermore, after the step of determining whether the score difference between the prefix part and the suffix part exceeds a first preset threshold, the following steps are included:
[0021] If the score difference between the prefix part and the suffix part exceeds a first preset threshold, it is determined that the command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
[0022] Furthermore, if the difference does not exceed the preset threshold, the step of aligning the command words to be recognized to obtain the score difference of each phoneme includes:
[0023] Use the backtracking algorithm to calculate the best matching path between the audio signal and the phoneme sequence based on the probability matrix output by the CTC model;
[0024] According to the best matching path, calculate the score of each phoneme at the corresponding time step, and calculate the average score of each phoneme in the entire path;
[0025] Based on the average score, the score difference of each phoneme in the command word to be recognized is calculated.
[0026] Furthermore, if the score difference is lower than a second preset threshold, the step of executing the command word to be recognized includes:
[0027] If the score difference of each phoneme is lower than the second preset threshold, the current command word to be recognized is confirmed as the final recognition result;
[0028] Based on the confirmed final recognition result, the corresponding execution instruction is generated and transmitted to the target device or application to perform the corresponding operation.
[0029] Furthermore, if the score difference is lower than a second preset threshold, after executing the step of the command word to be recognized, the method further comprises:
[0030] If the score difference is not lower than the second preset threshold, it is determined that the current command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
[0031] The second aspect of the present application provides a misidentification processing device, comprising:
[0032] A determination module is used to score each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized;
[0033] A scoring module, used for scoring the prefix and suffix of the command word to be recognized;
[0034] A judging module, used to judge whether the score difference between the prefix part and the suffix part exceeds a first preset threshold;
[0035] An alignment module, used for performing alignment processing on the command words to be recognized to obtain the score difference of each phoneme if the score difference does not exceed a preset threshold;
[0036] An execution module is used to execute the command word to be recognized if the score difference is lower than a second preset threshold.
[0037] The third aspect of the present application also includes a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0038] The fourth aspect of the present application also includes a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of any of the above methods when executed by a processor.
[0039] Beneficial effects:
[0040] This application introduces an independent scoring mechanism for the prefix and suffix parts of command words, and combines it with a consistency verification step to significantly improve the voice recognition accuracy of smart homes and smart terminal devices. First, by scoring the prefix and suffix parts of each determined command word to be recognized and calculating the score difference, the system can more finely evaluate the consistency of the command word, thereby effectively reducing the phenomenon of misrecognition. In addition, by comparing the score difference of each phoneme and comparing it with the second preset threshold, the possibility of misrecognition is further reduced through strict score difference control, ensuring that only the command words to be recognized with a score difference below the threshold can be confirmed as the final recognition result. The misrecognition rate is significantly reduced, the recognition ability of short command words is enhanced, and the overall performance of the system is improved through strict score difference control. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A schematic diagram of a flow chart of a misidentification processing method according to an embodiment of the present application;
[0042] Figure 2 A schematic block diagram of the structure of a misidentification processing device according to an embodiment of the present application;
[0043] Figure 3 A schematic block diagram of the structure of a computer device according to an embodiment of the present application.
[0044] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0046] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "a", "an", "above", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any module and all combinations of one or more associated listed items.
[0047] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the field to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0048] Reference Figure 1 The embodiment of the present invention provides a method for processing misidentification, including steps S1-S5, specifically:
[0049] S1. Score each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized;
[0050] S2, scoring the prefix and suffix of the command word to be recognized;
[0051] S3, determining whether the score difference between the prefix part and the suffix part exceeds a first preset threshold;
[0052] S4. If the score difference does not exceed the preset threshold, align the command words to be recognized to obtain the score difference of each phoneme;
[0053] S5. If the score difference is lower than a second preset threshold, execute the command word to be recognized.
[0054] As described in step S1 above, after the user's voice is recognized (for example, "Open XX"), the corresponding feature vector is extracted and then input into the CTC model. Then, all command word lists (such as "Open sports", "Open weather", "Open settings", etc.) are traversed, and each entry is scored according to the CTC decoding rules to find the entry with the highest score, which may be "Open weather" or "Open settings". Among them, the CTC model will output a probability matrix representing the probability distribution of each phoneme at each time step. For example, for "Open settings", the model may output a probability matrix similar to [0.9 (K), 0.05 (A), 0.03(I), ...]. Then, the forward algorithm is used to calculate the score of each possible path, thereby obtaining the path score of the phoneme sequence corresponding to each command word. Next, a dynamic programming method (such as the Viterbi algorithm) is used to find the optimal path, that is, the path with the highest score, and calculate its total score. For example, the optimal path and its score for "Open settings" may be 0.87. Finally, all command words are sorted according to the total score, and one or more command words are selected as the command words to be identified, which means that the next step of judgment and screening can be entered.
[0055] As described in the above steps S2-S3, it is first necessary to extract the corresponding phoneme sequence from the command word to be recognized determined in step S1. For example, assuming that the current command word to be recognized determined in step S1 is "open settings", then extract its corresponding phoneme sequence, and determine the phoneme boundaries of the prefix part and the suffix part according to the structure of the command word. For example, the prefix part of "open settings" is "open", and the corresponding phoneme sequence is [K, A, I, N]; the suffix part is "settings", and the corresponding phoneme sequence is [S, E, T, T, I, N, G]. The probability matrix output by the CTC model is used to calculate the path score of the prefix part phoneme sequence. For example, for the phoneme sequence [K, A, I, N] of "open", the score at each time step may be [0.9, 0.85, 0.87, 0.88], and the total score is 0.9 + 0.85 + 0.87 + 0.88 = 3.5. Similarly, using the probability matrix output by the CTC model, calculate the path score of the suffix part phoneme sequence. For example, for the "set" phoneme sequence [S, E, T, T, I, N, G], the score at each time step may be [0.86, 0.85, 0.83, 0.84, 0.85, 0.86, 0.87], and the total score is 0.86 + 0.85 + 0.83 + 0.84 + 0.85 + 0.86 + 0.87 = 5.96.
[0056] In step S2, the scores of the prefix part and the suffix part of the command word to be recognized have been calculated respectively. For example, suppose the score of the prefix part ("open") of "open settings" is 3.5, and the score of the suffix part ("settings") is 5.96. Calculate the score difference between the prefix part and the suffix part. The specific calculation method can be direct subtraction or calculation of the average score and then subtraction. For example, the score difference can be expressed as: ; Compare the calculated score difference with the first preset threshold set by the system. Assuming that the first preset threshold is 0.1, if the score difference is less than or equal to the first preset threshold, it is considered that the consistency of the entry meets the requirement; otherwise, it is considered that the consistency does not meet the requirement.
[0057] By calculating the scores of the prefix and suffix parts separately, the system can detect potential inconsistencies. For example, if the score of the prefix part is significantly lower than that of the suffix part, it indicates that there may be a risk of misidentification. By introducing the prefix and suffix scoring mechanism, the system can find more reliable matches in shorter command words and improve the robustness and stability of recognition. For example, after the short command words such as "open" and "close" are scored with the prefix and suffix, they will only be considered to have passed the first consistency screening process and can enter the next misidentification judgment step if the score difference is within the preset threshold range. Evaluating the score difference of the prefix and suffix parts separately helps to find those entries with high overall scores but poor internal consistency. For example, when processing "open settings", if the score difference of the prefix part is large, the system can refuse to recognize the entry to avoid misidentification as other similar command words.
[0058] As described in steps S4-S5 above, data is extracted from the probability matrix output by the CTC model obtained in step S1. This matrix represents the probability distribution of each phoneme at each time step. For example, for the "open setting", the model may output a probability matrix like [0.9 (K), 0.05 (A), 0.03 (I), ...]. A backtracking algorithm (such as the Viterbi algorithm) is used to find the best matching path between the audio signal and the phoneme sequence based on the probability matrix. This path not only gives the best arrangement order of the phonemes, but also provides the score of each phoneme at the corresponding time step.
[0059] For example, suppose the best path for "Open Settings" and its score is 0.87, and the scores of each phone at the corresponding time step are [0.9, 0.85, 0.87, 0.88, 0.86, 0.85, 0.83, 0.84, 0.85, 0.86,0.87]. Based on the best path, the score of each phone at the corresponding time step is calculated. This step ensures that the system can accurately evaluate the performance of each phone in the entire path.
[0060] For example, for the "open setting" phone sequence [K, A, I, N, S, E, T, T, I, N, G], the time step scores for each phone might be [0.9, 0.85, 0.87, 0.88, 0.86, 0.85, 0.83, 0.84, 0.85,0.86, 0.87]. The average score of each phone over the entire path is calculated. This step helps evaluate the overall performance of each phone and provides a basis for further score gap calculations.
[0061] For example, for the phoneme sequence [K, A, I, N, S, E, T, T, I, N, G] of "open settings", the average score of each phoneme can be [0.9, 0.85, 0.87, 0.88, 0.86, 0.85, 0.83, 0.84, 0.85, 0.86,0.87]. Calculate the score gap of each phoneme, that is, the difference between the score of the current phoneme and the scores of other adjacent phonemes. This step helps to find those phonemes with large score fluctuations, so as to further verify the consistency of the command word.
[0062] For example, for the phoneme sequence [K, A, I, N, S, E, T, T, I, N, G] of “turn on settings”, the score gap can be expressed as: score gap = [|0.9−0.85||,|0.85−0.87||,|0.87−0.88||,|0.88−0.86||,|0.86−0.85||,|0.85−0.83||,|0.83−0.84||,|0.84−0.85||,|0.85−0.86||,|0.86−0.87||]; the score gap results are [0.05, 0.02, 0.01, 0.02, 0.01, 0.02, 0.01, 0.01, 0.01, 0.01].
[0063] In step S5, these score differences are compared with a second preset threshold set by the system. Assuming that the second preset threshold is 0.05, if the score differences of all phonemes are lower than the second preset threshold (eg, 0.05), the current candidate recognition term is confirmed as the final recognition result.
[0064] For example, for “Open Setting”, all score gaps [0.05, 0.02, 0.01, 0.02, 0.01, 0.02,0.01, 0.01, 0.01, 0.01] are lower than 0.05, so “Open Setting” is confirmed as the final recognition result.
[0065] Steps S4-S5 calculate the time step score and score difference of each phoneme, so that the system can evaluate the performance of each phoneme in more detail and find potential inconsistencies. For example, when processing "open settings", the system can judge the consistency between phonemes by the score difference to ensure the accuracy of the final recognition result. For those phonemes with large score differences, the system can detect and reject them in time to avoid misidentifying them as other similar command words.
[0066] Through steps S1-S5, the present application realizes efficient and accurate recognition of command words. First, in step S1, the system extracts feature vectors (such as MFCC) from the audio signal and inputs them into the pre-trained CTC model. The forward algorithm is used to calculate the path score of the phoneme sequence corresponding to each command word, and the dynamic programming method (such as the Viterbi algorithm) is used to find the optimal path and its total score. All command words are sorted according to the total score, and several entries with the highest scores are selected as command words to be recognized, thereby realizing preliminary screening and basic data preparation. Then, in step S2, the system extracts the corresponding phoneme sequence from the command word to be recognized, and determines the phoneme boundaries of the prefix part and the suffix part, and uses the probability matrix output by the CTC model to calculate the path scores of the prefix part and the suffix part respectively, and preliminarily evaluates the internal consistency of the command word. Then, in step S3, the system calculates the score difference between the prefix part and the suffix part, and compares it with the first preset threshold, and further screens out those entries with a small score difference between the prefix and suffix parts to ensure that their internal consistency is high and effectively reduce the phenomenon of misrecognition. On the basis of the preliminary consistency verification, step S4 finds the best matching path between the audio signal and the phoneme sequence by using a backtracking algorithm (such as the Viterbi algorithm), calculates the score of each phoneme at the corresponding time step, further calculates the average score and score gap of each phoneme, verifies the overall consistency of the command word in detail, finds and excludes those phonemes with large score fluctuations, and ensures the reliability of the final recognition result. Finally, in step S5, the system compares the score gap of each phoneme with the second preset threshold to determine whether it is lower than the threshold. If the score gap of all phonemes is lower than the second preset threshold, the current candidate recognition term is confirmed as the final recognition result, and the corresponding execution instruction is generated and transmitted to the target device or application. At the same time, the user is provided with feedback on the recognition result and the recognition log is recorded. Through a strict score gap check mechanism, it is ensured that only those terms with high internal consistency can be confirmed as the final recognition result, thereby improving the user experience. The whole process is screened by two progressive consistency verifications (preliminary consistency and detailed consistency verification), and the screening conditions are gradually refined to ensure the accuracy and reliability of the final recognition result and avoid unnecessary complex calculations.
[0067] At the same time, this embodiment performs feature extraction and path score calculation based on a pre-trained CTC model. The existing CTC model has been trained on a large amount of data and can handle the mapping relationship between speech signals and phoneme sequences well. There is no need to retrain the model. The system can complete the recognition task in a short time without additional model training, saving time and resources. The modular design (such as feature extraction, path score calculation, prefix and suffix scoring, etc.) makes the scoring code highly reusable. Developers can quickly apply the same logic to different scenarios, simplifying development and maintenance work while maintaining low resource usage.
[0068] In one embodiment, the step of scoring each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized includes:
[0069] S10, extracting the feature vector of the current audio signal and inputting it into the pre-trained CTC model to obtain a probability matrix;
[0070] S11, using a forward algorithm to calculate the path score of the phoneme sequence corresponding to each command word based on the probability matrix;
[0071] S12, using a dynamic programming method to find the best path for each command word, and obtain the total score of the phoneme sequence corresponding to the command word;
[0072] S13. Sort all command words according to the total scores, and select a preset number of command words as command words to be recognized according to the sorting.
[0073] In this embodiment, first, the system extracts feature vectors from the current audio signal, for example, using feature extraction methods such as Mel Frequency Cepstral Coefficients (MFCC) to convert the audio signal into a numerical representation that can be processed by the computer. Then, these feature vectors are input into the pre-trained CTC model, which outputs a probability matrix representing the probability distribution of each phoneme at each time step. This probability matrix provides basic data for subsequent path score calculations.
[0074] Next, based on the probability matrix, the system uses a forward algorithm to calculate the path score of the phoneme sequence corresponding to each command word. The forward algorithm calculates the score of each possible path through dynamic programming, thereby obtaining the path score of the phoneme sequence corresponding to each command word. This process ensures that the system can fully evaluate all possible phoneme sequences and find the optimal path. To further improve recognition accuracy, the system uses dynamic programming to find the best path for each command word, that is, the path with the highest score, and calculates its total score. This method not only considers the score of the overall path, but also finds the optimal path through dynamic programming, which significantly improves the accuracy of the recognition results. By using forward algorithms and dynamic programming methods (such as the Viterbi algorithm) to calculate path scores and optimal paths, the system can complete complex computing tasks in a shorter time. These algorithms are efficient in terms of computational complexity and can quickly obtain results while ensuring accuracy.
[0075] After completing the above calculations, the system sorts the total scores of each command word and selects the words with the highest scores as the command words to be recognized. This step eliminates a large number of low-scoring words through preliminary screening, reducing the workload of subsequent processing. At the same time, this sorting and selection mechanism ensures that only those candidate words with higher scores can enter the next consistency verification stage, further improving the robustness and reliability of the system.
[0076] The entire process extracts audio features, inputs the CTC model to obtain a probability matrix, uses a forward algorithm to calculate the path score, uses dynamic programming to find the best path and calculate the total score, and finally selects a preset number of command words as the command words to be recognized based on the total score, ensuring the accuracy and reliability of the recognition results. This series of steps not only improves the recognition accuracy, but also reduces the possibility of misidentification through a strict screening mechanism, making the voice recognition of smart homes and smart terminal devices more accurate and reliable, greatly improving the user experience.
[0077] In one embodiment, the step of scoring the prefix part and the suffix part of the command word to be recognized includes:
[0078] S20, extracting a corresponding phoneme sequence based on the command word to be recognized;
[0079] S21, determining the phoneme boundaries of the prefix part and the suffix part according to the structure of the command word to be recognized;
[0080] S22. Use the probability matrix output by the CTC model to calculate the path score of the prefix part phoneme sequence and the path score of the suffix part phoneme sequence.
[0081] In this embodiment, first, the system extracts the corresponding phoneme sequence based on the command word to be recognized. For example, assuming that the command word to be recognized is "open settings", the system will convert it into a corresponding phoneme sequence, such as [K, A, I, N, S, E, T, T, I, N, G]. This process ensures that the subsequent steps can accurately process the key information in the audio signal.
[0082] Next, the system determines the phoneme boundaries of the prefix and suffix parts based on the structure of the command word to be recognized. For example, when processing "open settings", the system will divide the command word into the prefix part "open" and the suffix part "set". Specifically, the phoneme sequence of "open" is [K, A, I, N], while the phoneme sequence of "set" is [S, E, T, T, I, N,G]. This division helps the system to more carefully evaluate the consistency within the command word and discover potential inconsistencies.
[0083] The system then uses the probability matrix output by the CTC model to calculate the path scores of the phoneme sequences in the prefix and suffix parts respectively. The probability matrix represents the probability distribution of each phoneme at each time step, and the system uses this probability data to calculate the score of each possible path. For example, for the phoneme sequence [K, A, I, N] with the prefix part "open", the score at each time step may be [0.9, 0.85, 0.87, 0.88], and the total score is 0.9 + 0.85 + 0.87 +0.88 = 3.5. Similarly, for the phoneme sequence [S, E, T, T, I, N, G] with the suffix part “set”, the score at each time step might be [0.86, 0.85, 0.83, 0.84, 0.85, 0.86, 0.87], with a total score of 0.86 + 0.85 + 0.83 + 0.84 + 0.85 + 0.86 + 0.87 = 5.96.
[0084] By scoring the prefix and suffix parts separately, the system can evaluate the consistency within the command word in more detail. For example, if the score of the suffix part is significantly lower than the score of the prefix part, it means that there may be a risk of misrecognition. This method not only improves recognition accuracy, but also reduces the possibility of misrecognition by independently evaluating the score difference between the prefix and suffix parts. In addition, this mechanism can find more reliable matches in shorter command words, improving the robustness and stability of recognition.
[0085] Overall, by extracting phoneme sequences, determining the phoneme boundaries of the prefix and suffix parts, and using the probability matrix output by the CTC model to calculate the path scores of the respective parts, the system is able to more carefully evaluate the overall consistency of the command word, significantly reducing the possibility of misrecognition.
[0086] In one embodiment, after the step of determining whether the score difference between the prefix part and the suffix part exceeds a first preset threshold, the following steps are included:
[0087] S30: If the score difference between the prefix part and the suffix part exceeds a first preset threshold, it is determined that the command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
[0088] In this embodiment, specifically, after completing the score calculation of the prefix part and the suffix part, the system will compare the score difference between the two parts and compare it with a pre-set first preset threshold.
[0089] If the score difference between the prefix part and the suffix part exceeds the first preset threshold, the system will determine that the current command word to be recognized is a misrecognition, mark the current command word to be recognized as an invalid recognition, refuse to recognize the current command word to be recognized, and record relevant information for subsequent analysis and optimization.
[0090] Through this mechanism, the system can exclude those entries with poor internal consistency in the initial screening stage, significantly reducing the risk of misidentification. This approach not only improves recognition accuracy, but also enhances the robustness and reliability of the system. For example, when processing "open settings", the system can effectively distinguish it from similar but incorrect command words (such as "open the box"), avoiding misjudgment due to similar phonemes.
[0091] In one embodiment, if the difference does not exceed the preset threshold, the step of performing alignment processing on the command words to be recognized to obtain the score difference of each phoneme includes:
[0092] S40, using a backtracking algorithm to calculate the best matching path between the audio signal and the phoneme sequence based on the probability matrix output by the CTC model;
[0093] S41. Calculate the score of each phoneme at the corresponding time step according to the best matching path, and calculate the average score of each phoneme in the entire path;
[0094] S42: Based on the average score, calculate the score difference of each phoneme in the command word to be recognized.
[0095] In this embodiment, first, the system uses a backtracking algorithm to calculate the best matching path between the audio signal and the phoneme sequence based on the probability matrix output by the CTC model. Specifically, assuming that the voice command issued by the user is "open settings", the system has extracted the corresponding phoneme sequence through the previous steps and determined the phoneme boundaries of the prefix part and the suffix part. Next, the system uses a backtracking algorithm (such as the Viterbi algorithm) to find the best matching path between the audio signal and the phoneme sequence. This best path not only gives the best arrangement order of the phonemes, but also provides the score of each phoneme at the corresponding time step. Next, the system calculates the score of each phoneme at the corresponding time step according to the best matching path, and further calculates the average score of each phoneme in the entire path. Finally, based on the average score calculated above, the system further calculates the score gap of each phoneme in the command word to be recognized. The score gap is used to evaluate the difference between each phoneme and other adjacent phoneme scores and find potential inconsistencies. If the score gap of all phonemes is lower than the second preset threshold (for example, 0.05), the current command word to be recognized is confirmed as the final recognition result.
[0096] Through this mechanism, the system can evaluate the performance of each phoneme in detail during the alignment process, discover and exclude those phonemes with large score fluctuations, and ensure the reliability and consistency of the final recognition results. This method not only improves recognition accuracy, but also enhances the robustness and stability of the system. For example, when processing "open settings", the system can effectively distinguish it from similar but incorrect command words (such as "open the box"), avoiding misjudgment due to similar phonemes.
[0097] In one embodiment, if the score difference is lower than a second preset threshold, the step of executing the command word to be recognized includes:
[0098] S50, if the score difference of each phoneme is lower than the second preset threshold, confirming the current command word to be recognized as the final recognition result;
[0099] S51. Generate corresponding execution instructions according to the confirmed final recognition result, and transmit them to the target device or application program to execute corresponding operations.
[0100] In this embodiment, the system evaluates the score difference of each phoneme. Assuming that the current command word to be recognized is "open settings", the system has extracted the corresponding phoneme sequence through the previous steps and determined the phoneme boundaries of the prefix and suffix parts. Next, the system uses the probability matrix output by the CTC model to calculate the time step score of each phoneme, and further calculates the average score of each phoneme in the entire path and its score difference. For example, assume that the score difference of each phoneme is [0.05, 0.02, 0.01, 0.02, 0.01, 0.02, 0.01, 0.01, 0.01, 0.01].
[0101] Next, the system compares these score differences with a second preset threshold. Assume that the second preset threshold is set to 0.05. If the score differences of all phonemes are lower than the threshold (for example, [0.05, 0.02, 0.01, 0.02, 0.01, 0.02, 0.01, 0.01, 0.01, 0.01] are all lower than 0.05), the system confirms that the current command word to be recognized is the final recognition result. Once the final recognition result is confirmed, the system generates a corresponding execution instruction based on the result and transmits it to the target device or application to perform the corresponding operation. For example, a control signal is generated to perform the operation of "open settings". Feedback on the recognition result is provided to the user to inform the user that the system has successfully recognized and executed the command. For example, "Executed: Open Settings" is displayed on the user interface.
[0102] In one embodiment, if the score difference is lower than a second preset threshold, after executing the step of the command word to be recognized, the following steps are performed:
[0103] S60: If the score difference is not lower than a second preset threshold, it is determined that the command word to be recognized is misrecognized, and the command word to be recognized is rejected.
[0104] In this embodiment, assuming that the second preset threshold is set to 0.05, the system compares the score difference of each phoneme with the threshold. If the score difference of all phonemes is less than 0.05, the current command word to be recognized is confirmed as the final recognition result; if the score difference of one or more phonemes is not less than 0.05, the current command word to be recognized is determined to be misrecognized, and it is rejected for recognition. The system marks the current command word to be recognized as invalid recognition and records relevant information for subsequent analysis and optimization. For example, suppose the voice command issued by the user is "open the box", but the system always verifies it with "open settings" in the command word list. After consistency verification, it is found that the score difference is large, so "open settings" is marked as invalid recognition and no operation response is made to it.
[0105] Reference Figure 2 , is a structural block diagram of a misidentification processing device in an embodiment of the present application, the device includes:
[0106] A determination module 100, configured to score each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized;
[0107] A scoring module 200, used for scoring the prefix part and the suffix part of the command word to be recognized;
[0108] A judging module 300 is used to judge whether the score difference between the prefix part and the suffix part exceeds a first preset threshold;
[0109] An alignment module 400 is used to align the command words to be recognized to obtain the score difference of each phoneme if the score difference does not exceed a preset threshold;
[0110] The execution module 500 is configured to execute the command word to be recognized if the score difference is lower than a second preset threshold.
[0111] In one embodiment, the determination module 100 includes a planning confirmation unit, which is used to:
[0112] Extract the feature vector of the current audio signal and input it into the pre-trained CTC model to obtain a probability matrix;
[0113] Calculate the path score of the phoneme sequence corresponding to each command word using a forward algorithm based on the probability matrix;
[0114] Using a dynamic programming method to find the best path for each command word, and obtaining the total score of the phoneme sequence corresponding to the command word;
[0115] All command words are sorted according to the total scores, and a preset number of command words are selected according to the sorting as command words to be recognized.
[0116] In one embodiment, the scoring module 200 includes a first scoring unit configured to:
[0117] Based on the command word to be recognized, extract the corresponding phoneme sequence;
[0118] According to the structure of the command word to be recognized, determining the phoneme boundaries of the prefix part and the suffix part;
[0119] Using the probability matrix output by the CTC model, the path score of the prefix partial phoneme sequence and the path score of the suffix partial phoneme sequence are calculated.
[0120] In one embodiment, the above device further includes a first rejection recognition module, which is used to:
[0121] If the score difference between the prefix part and the suffix part exceeds a first preset threshold, it is determined that the command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
[0122] In one embodiment, the alignment module 400 includes a second scoring unit, which is used to:
[0123] Use the backtracking algorithm to calculate the best matching path between the audio signal and the phoneme sequence based on the probability matrix output by the CTC model;
[0124] According to the best matching path, calculate the score of each phoneme at the corresponding time step, and calculate the average score of each phoneme in the entire path;
[0125] Based on the average score, the score difference of each phoneme in the command word to be recognized is calculated.
[0126] In one embodiment, the execution module 500 includes a threshold determination unit, which is used to:
[0127] If the score difference of each phoneme is lower than the second preset threshold, the current command word to be recognized is confirmed as the final recognition result;
[0128] Based on the confirmed final recognition result, the corresponding execution instruction is generated and transmitted to the target device or application to perform the corresponding operation.
[0129] In one embodiment, the above device further includes a second rejection identification module, which is used to:
[0130] If the score difference is not lower than the second preset threshold, it is determined that the current command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
[0131] Reference Figure 3 In an embodiment of the present application, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the usage data and the like in the process of the misidentification processing method. The network interface of the computer device is used to communicate with an external terminal through a network connection. Further, the above-mentioned computer device can also be provided with an input device and a display screen and the like. When the above-mentioned computer program is executed by the processor to implement the misidentification processing method, it includes the following steps: scoring each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized; scoring the prefix part and the suffix part of the command word to be recognized; judging whether the score difference between the prefix part and the suffix part exceeds a first preset threshold; if it does not exceed the preset threshold, aligning the command word to be recognized to obtain the score difference of each phoneme; if the score difference is lower than a second preset threshold, executing the command word to be recognized. Those skilled in the art will understand that Figure 3 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.
[0132] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, a misidentification processing method is implemented, including the following steps: scoring each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized; scoring the prefix part and the suffix part of the command word to be recognized; judging whether the score difference between the prefix part and the suffix part exceeds a first preset threshold; if it does not exceed the preset threshold, aligning the command word to be recognized to obtain the score difference of each phoneme; if the score difference is lower than a second preset threshold, executing the command word to be recognized. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0133] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0134] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, device, article or method including the element.
[0135] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for processing misidentification, characterized in that: The method comprises: Score each entry in the command word list according to the CTC decoding rules to determine the command word to be recognized; Scoring the prefix and suffix of the command word to be recognized; Determine whether the score difference between the prefix part and the suffix part exceeds a first preset threshold; If it does not exceed the preset threshold, the command word to be recognized is aligned to obtain the score difference of each phoneme; If the score difference is lower than a second preset threshold, the command word to be recognized is executed.
2. The method for handling misidentification according to claim 1, characterized in that: The step of scoring each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized includes: Extract the feature vector of the current audio signal and input it into the pre-trained CTC model to obtain a probability matrix; Calculate the path score of the phoneme sequence corresponding to each command word using a forward algorithm based on the probability matrix; Using a dynamic programming method to find the best path for each command word, and obtaining the total score of the phoneme sequence corresponding to the command word; All command words are sorted according to the total scores, and a preset number of command words are selected according to the sorting as command words to be recognized.
3. The method for handling misidentification according to claim 1, characterized in that: The step of scoring the prefix part and the suffix part of the command word to be recognized includes: Based on the command word to be recognized, extract the corresponding phoneme sequence; According to the structure of the command word to be recognized, determining the phoneme boundaries of the prefix part and the suffix part; Using the probability matrix output by the CTC model, the path score of the prefix partial phoneme sequence and the path score of the suffix partial phoneme sequence are calculated.
4. The method for handling misidentification according to claim 1, characterized in that: After the step of determining whether the score difference between the prefix part and the suffix part exceeds the first preset threshold, the method further includes: If the score difference between the prefix part and the suffix part exceeds a first preset threshold, it is determined that the command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
5. The method for handling misidentification according to claim 1, characterized in that: If the difference does not exceed the preset threshold, the step of aligning the command words to be recognized to obtain the score difference of each phoneme includes: Use the backtracking algorithm to calculate the best matching path between the audio signal and the phoneme sequence based on the probability matrix output by the CTC model; According to the best matching path, calculate the score of each phoneme at the corresponding time step, and calculate the average score of each phoneme in the entire path; Based on the average score, the score difference of each phoneme in the command word to be recognized is calculated.
6. The method for handling misidentification according to claim 1, characterized in that: If the score difference is lower than a second preset threshold, the step of executing the command word to be recognized comprises: If the score difference of each phoneme is lower than the second preset threshold, the current command word to be recognized is confirmed as the final recognition result; Based on the confirmed final recognition result, the corresponding execution instruction is generated and transmitted to the target device or application to perform the corresponding operation.
7. The method for handling misidentification according to claim 1, characterized in that: If the score difference is lower than the second preset threshold, after executing the step of the command word to be recognized, the method further comprises: If the score difference is not lower than the second preset threshold, it is determined that the current command word to be recognized is misrecognized, and the current command word to be recognized is rejected.
8. A misidentification processing device, characterized in that: include: A determination module is used to score each entry in the command word list according to the CTC decoding rule to determine the command word to be recognized; A scoring module, used for scoring the prefix and suffix of the command word to be recognized; A judging module, used to judge whether the score difference between the prefix part and the suffix part exceeds a first preset threshold; An alignment module, used for performing alignment processing on the command words to be recognized to obtain the score difference of each phoneme if the score difference does not exceed a preset threshold; An execution module is used to execute the command word to be recognized if the score difference is lower than a second preset threshold.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Voice decoding result processing method and device, equipment and storage medium
CN115497484A
Quick decoding method and device based on voice command word recognition, equipment and medium
CN118748011A