Method, apparatus, device and storage medium for confirming easily confused words
The proposed method enhances speech recognition accuracy by training with general and specific command datasets, extracting embeddings, and calculating similarity to address misidentification of similar-sounding commands, improving adaptability and efficiency.
Patent Information
- Application Number
- CN202510267822.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing CTC algorithms are prone to misrecognition of command words with similar pronunciations in embedded speech recognition systems, and lack a secondary verification mechanism, resulting in low recognition accuracy.
By training the initial speech recognition model based on the preset command word corpus, an optimization recognition model is generated, the embedded representation of audio clips is extracted, the phoneme embedding representation dictionary is constructed, and the similarity between the input audio and the confusing command words is calculated for final recognition.
It significantly improves the accuracy of the speech recognition system for easily confusing command words, reduces the possibility of misidentification, and improves the system's adaptability and flexibility.
Smart Images

Figure CN119763549B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a method, device, equipment and storage medium for confirming easily confused words. Background Art
[0002] The current embedded speech recognition system generally adopts the CTC (Connectionist Temporal Classification) algorithm, which is widely used in resource-constrained scenarios such as smart homes and wearable devices because it does not require mandatory audio-text alignment through end-to-end training. However, the existing solution directly relies on the Top-1 result output by CTC as the final recognition result, which has significant defects: first, for command words with similar pronunciations (such as "twenty-seven degrees" and "twenty-one degrees"), it is easy to cause misrecognition due to the slight difference in CTC path probability; second, the static decision mechanism of single forward propagation lacks the ability to perform secondary verification for easily confused scenes; in addition, the phoneme-level embedded features contained in the middle layer of the model are not effectively utilized, missing the opportunity to improve the discrimination accuracy through fine-grained acoustic feature comparison.
[0003] Therefore, the existing CTC scheme is prone to misrecognition of command words with similar pronunciations, and the lack of a secondary verification mechanism leads to low recognition accuracy, which is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The main purpose of this application is to provide a method, device, equipment and storage medium for confirming easily confused words, aiming to solve the technical problems that the existing CTC scheme is prone to misidentification of command words with similar pronunciations and lacks a secondary verification mechanism, resulting in low recognition accuracy.
[0005] In order to achieve the above-mentioned invention object, the present application proposes a method for confirming easily confused words, the method comprising:
[0006] Based on the preset command word data, the initial speech recognition model is trained to obtain an optimized recognition model;
[0007] Input the real command word data into the optimized recognition model and extract the embedding representation of each audio clip;
[0008] Based on the extracted embedding representation, construct an embedding representation dictionary for each phoneme;
[0009] When receiving new speech input, the corresponding audio embedding sequence is generated by optimizing the recognition model;
[0010] Based on the embedding representation dictionary, obtaining the phoneme embedding representation sequences of all easily confused command words in the easily confused command word list corresponding to the embedding sequence;
[0011] Calculate the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the confusing command words to obtain the final recognition result.
[0012] Further, the step of training the initial speech recognition model based on the preset command word corpus to obtain an optimized recognition model includes:
[0013] Train a speech recognition model based on a general corpus until the loss rate or word error rate on the validation set is reduced to a first preset threshold to obtain the initial speech recognition model;
[0014] Train based on the preset command word corpus on the basis of the initial speech recognition model until the loss rate or word error rate of the initial speech recognition model on the validation set containing command words is reduced to a second preset threshold to obtain the optimized recognition model.
[0015] Further, the step of inputting the real command word corpus into the optimized recognition model and extracting the embedding representation of each audio segment includes:
[0016] Obtain the preset real command word corpus, where the real command word corpus is an audio segment containing command words;
[0017] Input the real command word corpus into the optimized recognition model;
[0018] Use the optimized recognition model to infer the input command word corpus to generate a feature vector corresponding to each audio segment, that is, the embedding representation.
[0019] Further, the step of constructing an embedding representation dictionary for each phoneme according to the extracted embedding representation includes:
[0020] For the audio segments included in the real command word corpus, calculate the path score based on the ctc algorithm;
[0021] Based on the path score, use the backtracking algorithm to find the best alignment path between the audio and phonemes by selecting the maximum path score;
[0022] Create an initial phoneme embedding representation dictionary, where each phoneme corresponds to an initial embedding representation;
[0023] When the embedding representation score of the audio segment is greater than the preset threshold, it is recognized as a high-quality audio embedding representation;
[0024] According to the best alignment path, map the high-quality audio embedding representation to the corresponding phoneme, and update the embedding representation of the corresponding phoneme by the method of moving average to obtain the embedding representation dictionary.
[0025] Further, the step of generating an embedding sequence of the corresponding audio by optimizing the recognition model when a new voice input is received includes:
[0026] Preprocess the voice input to obtain an audio segment to be recognized;
[0027] Input the audio segment to be recognized into the optimized recognition model, and generate an embedding representation of the audio segment to be recognized through forward propagation, that is, the embedding sequence of the corresponding audio.
[0028] Further, the step of obtaining the phoneme embedding representation sequences of all confusing command words in the confusing command word list corresponding to the embedding sequence based on the embedding representation dictionary includes:
[0029] Identify the confusing command word list corresponding to the embedding sequence, and load all confusing command words in the confusing command word list;
[0030] Convert the confusing command words into corresponding phoneme sequences;
[0031] According to each phoneme sequence, obtain the corresponding phoneme embedding representation from the pre-constructed phoneme embedding representation dictionary to obtain the corresponding phoneme embedding representation sequence.
[0032] Further, the step of calculating the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequences of the confusing command words to obtain the final recognition result includes:
[0033] Align the embedding sequence of the input audio with the phoneme embedding representation sequences of the confusing command words;
[0034] Use a similarity measurement method to calculate the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequences of each confusing command word;
[0035] Compare all similarity scores, and select the command word with the highest score as the final recognition result.
[0036] The second aspect of this application proposes a confusing word confirmation device, including:
[0037] An optimization module for training on an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model;
[0038] An extraction module for inputting a real command word corpus into the optimized recognition model to extract the embedding representation of each audio segment;
[0039] A construction module for constructing a phoneme embedding representation dictionary for each phoneme according to the extracted embedding representation;
[0040] A generation module, configured to generate an embedding sequence of the corresponding audio by optimizing an identification model when a new voice input is received;
[0041] An acquisition module, configured to acquire a phoneme embedding representation sequence of all easily confused command words in the list of easily confused command words corresponding to the embedding sequence based on an embedding representation dictionary;
[0042] A calculation module, configured to calculate a similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the easily confused command words to obtain a final recognition result.
[0043] The third aspect of this application further includes a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0044] The fourth aspect of this application further includes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0045] Beneficial effects:
[0046] This solution can significantly improve the recognition accuracy of the voice recognition system for easily confused command words. First, based on the CTC algorithm and the backtracking algorithm, the best alignment path between the audio and the phonemes is found, ensuring the generation of high-quality phoneme embedding representations. By combining the command word corpus for training, the optimized model can more accurately recognize specific command words, reducing the possibility of misrecognition. Using the phoneme embedding representation dictionary, the system can efficiently process newly input audio data in actual applications, and through establishing a list of easily confused words that are easily confused with the currently recognized command word, a secondary confirmation mechanism within the easily confused words is carried out. Specifically, the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the easily confused command words is calculated to accurately distinguish the easily confused command words. This series of steps not only improves the recognition accuracy, but also reduces the computational complexity by optimizing the model and the generation process of the embedding representation, enhancing the adaptability and flexibility of the system. Description of the drawings
[0047] Figure 1 It is a schematic flowchart of a method for confirming easily confused words according to an embodiment of this application;
[0048] Figure 2 It is a schematic block diagram of the structure of a device for confirming easily confused words according to an embodiment of this application;
[0049] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of this application.
[0050] The realization, functional features and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the above" and "the" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of features, integers, steps, operations, elements, modules and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any module and all combinations of one or more related listed items.
[0053] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0054] Referring to Figure 1 , an embodiment of the present invention provides a method for confirming homophones, including steps S1-S6, specifically:
[0055] S1. Train on an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model;
[0056] S2. Input the real command word corpus into the optimized recognition model, and extract the embedded representation of each audio segment;
[0057] S3. According to the extracted embedded representation, construct an embedded representation dictionary for each phoneme;
[0058] S4. When a new voice input is received, generate an embedded sequence of the corresponding audio through the optimized recognition model;
[0059] S5. Based on the embedded representation dictionary, obtain the phoneme embedding representation sequences of all the confusing command words in the list of confusing command words corresponding to the embedding sequence.
[0060] S6. Calculate the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequences of the confusing command words to obtain the final recognition result.
[0061] As described in step S1 above, first, in the basic training stage, a speech recognition model is trained using general corpus. General corpus usually contains a large amount of daily speech data, aiming to let the model learn a wide range of language features. For example, a publicly available speech dataset (such as LibriSpeech) containing multiple languages, different accents, and background noise can be used. During the training process, the model generates a prediction output through forward propagation, calculates the loss function (such as CTC loss, Connectionist Temporal Classification Loss), and then adjusts the model parameters through backpropagation to minimize the loss value. This process continues until the loss rate or word error rate (WER) of the model on the validation set drops to a preset first threshold, obtaining an initial speech recognition model that can be directly used or used for subsequent fine-tuning. This step ensures that the model has basic speech recognition capabilities and can process general speech inputs.
[0062] Next is the fine-tuning stage, where the preset command word corpus is introduced into the training process. The command word corpus is a dataset designed specifically for specific application scenarios (such as command words like "turn on the light" and "turn off the light" in smart home control). The target command words in these corpora include confusing words (such as "twenty-seven degrees" and "twenty-one degrees"). During the fine-tuning process, the initial speech recognition model is continued to be used, and the model parameters are further adjusted in combination with the command word corpus. In this way, the model can learn more specific command word features and continue to be trained on the validation set containing command words until the loss rate or WER drops to a second preset threshold. This step significantly improves the recognition accuracy of the model for specific command words and reduces the possibility of misrecognition.
[0063] For example, assume that we want to train a speech recognition model for a smart speaker, which includes two command words "play music" and "pause music". First, we use the general corpus to perform basic training on the model so that it can recognize general speech instructions. Then, we add the preset command word corpus containing "play music" and "pause music" and other command words used in actual applications to the training set for fine-tuning to obtain an optimized recognition model. During this process, the model will learn how to accurately distinguish the corresponding command words, so as to accurately respond to the user's instructions in actual applications.
[0064] The function of step S1 is to ensure that the model has both extensive speech recognition capabilities and can accurately recognize specific command words, especially those that are easily confused, through a two-stage training method. This optimized model improves the recognition accuracy, can operate stably in complex speech environments, and enhances the user experience.
[0065] As described in steps S2 - S3 above, first, real command word corpora need to be prepared. These corpora should cover all target command words and be as diverse as possible to ensure that the model can learn the features in various situations. For example, for the speech recognition task of a smart home control system, the corpora should include common command words such as "turn on the light", "turn off the light", "raise the temperature", "lower the temperature", etc., and cover different pronunciations, accents, and background noise situations. Next, these corpora are passed as inputs to the fine-tuned optimized recognition model. Before input, it is usually necessary to preprocess the audio data, such as noise reduction, normalization, and segmentation, etc., to improve the stability and accuracy of model processing.
[0066] Then, the embedding representation of each audio segment is generated through model inference. Specifically, the forward propagation process inside the model generates a series of high-dimensional feature vectors, which represent the feature representations of the audio segment at different time steps. The embedding representation here is usually extracted from the feature layer before the output layer of the model, such as the feature before the linear layer or the fully connected layer. For the CTC algorithm, the model generates a probability matrix with the time steps matching the audio length. The maximum score path is found through the backtracking algorithm, and the best alignment path between the audio sequence and phonemes is determined. Based on this best alignment path, the audio embedding representation corresponding to each phoneme can be extracted.
[0067] The goal of step S3 is to construct an embedding representation dictionary for each phoneme based on the extracted embedding representations. First, create an initial phoneme embedding representation dictionary. This dictionary can be initialized based on the embedding representations generated by a pre-trained model or randomly initialized. For example, assume we have a list containing common phonemes (such as "sh", "ang", "yi", etc.), and we can assign an initial embedding vector to each phoneme. These initial embedding vectors can be obtained from a pre-trained language model or randomly generated high-dimensional vectors. For each input real command word corpus, use the CTC algorithm to calculate the scores of each possible path and find the best alignment path between the audio and phonemes through the backtracking algorithm. This step ensures an accurate correspondence between the audio segments and phonemes. Set a quality threshold, and only when the embedding representation score of an audio segment is higher than this threshold is it considered a high-quality embedding representation. For example, if the embedding representation score of a certain audio segment is 0.9 and the threshold is set to 0.8, the embedding representation of this segment will be recognized as high-quality and used for subsequent updates. According to the best alignment path, map the high-quality audio embedding representations to the corresponding phonemes and update the phoneme embedding representations by the method of moving average. Specifically, for each phoneme, collect all high-quality audio embedding representations and update the embedding representation of this phoneme by weighted average. The formula example is as follows: ; where α is a weight coefficient between 0 and 1, used to control the update speed and smoothness. For example, assume the current embedding representation of the phoneme "sh" is E current , and the newly obtained embedding representation is E new , then update it through the above formula. After the update is completed, save the final phoneme embedding representation dictionary and evaluate it using the test set to check whether it can effectively reduce the recognition errors of homophones. If it is found that the embedding representations of some phonemes are not accurate enough or there are large errors, the model parameters can be further adjusted or part of the data can be retrained to improve the overall performance.
[0068] Suppose the input "twenty-seven degrees" is fed into the optimized recognition model to update the corresponding content in the dictionary and construct the dictionary. First, in step S2, we preprocess the received speech sample, removing background noise and normalizing the volume, and then input it into the optimized speech recognition model. The model generates a corresponding embedding sequence, such as [E_1, E_2,..., E_n], where E_i represents the embedding vector at the i-th time step. Then, in step S3, based on these high-quality audio embedding representations, we gradually update the phoneme embedding representation dictionary. For example, the phoneme sequence of "twenty-seven degrees" is ["er4", "shi2", "qi1", "du4"], and we update the high-quality embedding representation corresponding to each phoneme into the dictionary by the moving average method, where the audio is mapped from length n to length m, m is the length of the phoneme sequence and n is the length of the audio sequence. Finally, an accurate phoneme embedding representation dictionary can be obtained, which contains the high-quality embedding representations of all phonemes.
[0069] The functions of steps S2 - S3 are to generate high-quality audio embedding representations. These embedding representations not only reflect the features of the audio segments but also align with specific phonemes, thus providing accurate data support for subsequent steps. In this way, the system can effectively construct the phoneme embedding representation dictionary and efficiently process newly input audio data in practical applications. High-quality embedding representations can be extracted from real speech samples and mapped to the corresponding phonemes, thereby constructing an accurate phoneme embedding representation dictionary. In this way, in subsequent steps, the system can distinguish these confusing command words by calculating similarities, significantly improving the accuracy and reliability of recognition. Ultimately, such high-quality embedding representations provide a solid foundation for the performance improvement of the entire speech recognition system.
[0070] As described in step S4 above, this step aims to convert newly input audio data into high-dimensional feature representations (embedding representations) for subsequent processing and recognition of confusing command words.
[0071] First, in the specific implementation process, necessary preprocessing operations need to be performed on the newly received speech input. These preprocessing steps include noise reduction, normalization, and segmentation, etc., to ensure the quality and consistency of the audio data. For example, suppose the user issues a voice command "play music" through a smart speaker. The system will first perform noise reduction on this audio signal, removing background noise and normalizing it to the same volume level. Then, the audio signal is converted into a format suitable for model input, such as a Mel-spectrogram or other common audio feature representation forms.
[0072] Next, load the trained and fine-tuned speech recognition model. This model has been fully trained on general corpora and command word corpora and has performed well on the validation set. After loading the model, pass the preprocessed audio segments as input to the model. The model processes these data step by step according to its internal structure, generating high-dimensional feature vectors at each time step. Specifically, the model generates a series of embedding representations through forward propagation, usually extracting the features before the output layer of the model (such as the features before the linear layer or the fully connected layer). For example, assume that the last layer of the model is a linear layer, then the features before this layer are extracted as the audio embedding representation.
[0073] To ensure that the generated embedding representations can be aligned with the phoneme embedding representation dictionary in the subsequent steps, some additional processing steps may be required. For example, use the backtracking method of the CTC algorithm to align the audio segments and phoneme sequences, or apply a certain smoothing technique to reduce the impact of noise. This step helps to improve the quality of the embedding representations, making the subsequent similarity calculations more accurate.
[0074] The role of step S4 is to convert the newly input audio data into high-quality embedding representations. These embedding representations not only reflect the characteristics of the audio segments but also are aligned with specific phonemes, thus providing accurate data support for the subsequent steps. In this way, the system can efficiently process newly input audio data in practical applications and distinguish between easily confused command words by calculating similarities, significantly improving the accuracy and reliability of recognition.
[0075] As described in step S5 above, first, at the beginning of step S5, the embedding sequence of the input audio (embedding representation) has been generated through step S4. Assume that the embedding sequence of the input audio is [E_1, E_2,..., E_n], where E_i represents the embedding vector at the i-th time step.
[0076] Next, prepare a list of easily confused command words. For example, construct or load a list containing easily confused command words, such as [[twenty-seven degrees, twenty-one degrees], [previous song, next song]]. The command words in each sub-list are words that are easily misrecognized. For a smart home control system, these easily confused command words may include "turn on the light" and "turn off the light", "raise the temperature" and "lower the temperature", etc.
[0077] Then, convert each easily confused command word into its corresponding phoneme sequence. This step usually needs to be completed using a phoneme mapping tool (such as G2P, Grapheme-to-Phoneme converter). For example, "twenty-seven degrees" can be converted into the phoneme sequence ["er4", "shi2", "qi1", "du4"].
[0078] According to each phoneme sequence, obtain the corresponding phoneme embedding representation from a pre-constructed phoneme embedding representation dictionary. The specific steps are as follows: For each phoneme, look up its corresponding embedding vector in the embedding representation dictionary. In the order of the phoneme sequence, concatenate the embedding representation vectors of each phoneme to form the complete phoneme embedding representation sequence of the command word. For example, the phoneme embedding representation sequence of "twenty-seven degrees" is [E_{er4}, E_{shi2}, E_{qi1}, E_{du4}], where E_{x} represents the embedding representation vector of phoneme x.
[0079] For example, assume that the user issues a voice command "twenty-seven degrees". The system first preprocesses the received audio signal and generates an embedding sequence [E_1, E_2,..., E_n] through an optimized speech recognition model. Based on the recognition result of the sequence, a list of confusing command words that may contain "twenty-seven degrees" is recognized. Then, the system prepares a list of confusing command words [[twenty-seven degrees, twenty-one degrees], [previous song, next song]], and uses the list of confusing command words containing "twenty-seven degrees" as the operation list for the next secondary comparison. Then, "twenty-seven degrees" and "twenty-one degrees" in the list of confusing command words are respectively converted into phoneme sequences ["er4", "shi2", "qi1", "du4"] and ["er4", "yi2", "shi2", "du4"].
[0080] Then, the system obtains the phoneme embedding representation sequences of these two command words from the phoneme embedding representation dictionary: the phoneme embedding representation sequence of "twenty-seven degrees" is [E_{er4}, E_{shi2}, E_{qi1}, E_{du4}]; the phoneme embedding representation sequence of "twenty-one degrees" is [E_{er4}, E_{yi2}, E_{shi2}, E_{du4}]; in this way, the system provides accurate phoneme embedding representation sequences for subsequent steps. In step S6, the system will calculate the similarity between the embedding sequence of the input audio and these two phoneme embedding representation sequences, and finally determine whether the command word issued by the user is "twenty-seven degrees" or "twenty-one degrees".
[0081] The function of step S5 is to ensure that the embedding sequence of the newly input audio can be accurately matched and compared with the confusing command words, thereby improving the recognition accuracy. By constructing the phoneme embedding representation sequences of the confusing command words, the system can efficiently process the newly input audio data in actual applications, which is convenient for subsequent distinguishing of the confusing command words by calculating the similarity.
[0082] As described in step S6 above, first, at the start of step S6, the embedding sequence of the input audio has been generated through step S4, and the sequence of phoneme embedding representations for each command word in the list of confusing command words has been obtained through step S5. Suppose the embedding sequence of the input audio is [E_1, E_2, ..., E_n], and the sequence of phoneme embedding representations for "twenty-seven degrees" is [E_{er4}, E_{shi2}, E_{qi1}, E_{du4}], and the sequence of phoneme embedding representations for "twenty-one degrees" is [E_{er4}, E_{yi2}, E_{shi2}, E_{du4}].
[0083] Next, using a suitable similarity metric method, calculate the similarity between the embedding sequence of the input audio and the sequence of phoneme embedding representations for each confusing command word.
[0084] Align the embedding sequence of the input audio with the sequence of phoneme embedding representations for each confusing command word. This can be achieved through Dynamic Time Warping (DTW) or other alignment algorithms. DTW can handle time series data of different lengths, find the optimal alignment path, and minimize the difference between the two sequences. According to the aligned embedding sequences, use the similarity metric method to calculate the similarity scores. For example: Calculate the cosine of the angle between two vectors, which ranges from -1 to 1, and the closer the value is to 1, the more similar they are. Calculate the Euclidean distance between two vectors, and the smaller the distance, the more similar they are. Compare all the similarity scores and select the command word with the highest score as the final recognition result. For example, suppose the similarity score for "twenty-seven degrees" is 0.9 and the similarity score for "twenty-one degrees" is 0.8, then select "twenty-seven degrees" as the final recognition result. Output the command word with the highest score as the final recognition result and perform subsequent processing or feedback to the user as needed. Step S6 ensures the accuracy and reliability of the speech recognition system through high-quality similarity calculations. This precise similarity calculation not only improves the recognition accuracy of the system but also enhances the adaptability and flexibility of the system, enabling the system to operate stably in complex speech environments and improving the user experience. Finally, this method ensures the efficiency and reliability of the speech recognition system in various application scenarios such as smart homes and smart terminals.
[0085] Through steps S1 - S6, the present application implements an efficient and accurate speech recognition system, which is particularly suitable for dealing with the recognition problem of confusing command words. First, in step S1, based on the preset command word corpus, training is carried out on the initial speech recognition model to obtain an optimized recognition model, ensuring that the model can accurately recognize specific command words and reduce the possibility of misrecognition. Then, in step S2, the real command word corpus is input into the optimized recognition model to extract the embedding representation of each audio segment, generating high-quality audio embedding representations, providing a solid foundation for constructing the phoneme embedding representation dictionary in the follow-up. In step S3, according to the extracted embedding representations, the embedding representation dictionary of each phoneme is constructed. The corresponding relationship between the audio and the phoneme is found through the optimal alignment path, and the embedding representations of the same phoneme are averaged to generate a high-quality phoneme embedding representation dictionary. In step S4, when a new speech input is received, the embedding sequence of the corresponding audio is generated through the optimized recognition model, ensuring that the newly input audio can be converted into high-quality embedding representations. In step S5, based on the embedding representation dictionary, the phoneme embedding representation sequences of all confusing command words in the list of confusing command words corresponding to the embedding sequence of the input audio are obtained, preparing the data for similarity calculation. Finally, in step S6, the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequences of the confusing command words is calculated, and the final recognition result is determined through cosine similarity, thereby accurately distinguishing the confusing command words. This series of steps significantly improves the accuracy, adaptability, and flexibility of the speech recognition system, enabling the system to operate stably in a complex speech environment, enhancing the user experience, and being widely applied to various application scenarios such as smart homes and smart terminals.
[0086] In one embodiment, the step of training on the initial speech recognition model based on the preset command word corpus to obtain the optimized recognition model includes:
[0087] S10. Training the speech recognition model based on the general corpus until the loss rate or word error rate on the validation set is reduced to the first preset threshold to obtain the initial speech recognition model;
[0088] S11. Training on the basis of the initial speech recognition model based on the preset command word corpus until the loss rate or word error rate of the initial speech recognition model on the validation set containing command words is reduced to the second preset threshold to obtain the optimized recognition model.
[0089] In this embodiment, in the process of training an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model, first, the speech recognition model is initially trained using a general corpus until the loss rate or word error rate of the model on the validation set drops to a preset first threshold, thereby obtaining the initial speech recognition model. The goal of this stage is to enable the model to learn extensive language features and possess the basic ability to process various speech inputs. The general corpus usually contains a large amount of daily speech data, covering different languages, accents, and background noises, ensuring that the model has strong generalization ability. For example, the publicly available LibriSpeech dataset can be used, which contains a large number of English speech segments to help the model learn basic speech features.
[0090] Next, after obtaining the initial speech recognition model, it is further fine-tuned using specific command word corpora. These command word corpora are designed specifically for the specific command words in the application scenario, such as "turn on the light", "turn off the light", etc. in smart home control. In this way, the model can learn more specific command word features and significantly improve the recognition accuracy of these command words. During the fine-tuning process, the results of the initial model's training on the general corpus are continued to be used as a starting point, and the model parameters are further adjusted in combination with the command word corpus. Specifically, the command word corpus is input into the model, the scores of each path are calculated, and the backtracking algorithm is used to find the best alignment path between the audio and phonemes, generating high-quality embedding representations. This process continues until the loss rate or word error rate of the model on the validation set containing the command words drops to a second preset threshold. This step significantly improves the model's recognition accuracy for specific command words and reduces the possibility of misrecognition.
[0091] Through this two-stage training method, the system not only has extensive speech recognition capabilities but also can accurately recognize specific command words, especially those that are easily confused. This method provides a solid foundation for subsequent steps (such as extracting embedding representations, constructing a phoneme embedding representation dictionary, etc.), and ultimately realizes an efficient and accurate speech recognition system. The system can operate stably in a complex speech environment, improve the user experience, and be widely applied to various application scenarios such as smart homes and smart terminals. In this way, the system can efficiently process newly input audio data in actual applications, significantly improving the recognition accuracy and reliability.
[0092] In one embodiment, the step of inputting the real command word corpus into the optimized recognition model and extracting the embedding representation of each audio segment includes:
[0093] S20. Obtain a preset real command word corpus, where the real command word corpus is an audio segment containing command words;
[0094] S21. Input the real command word corpus into the optimized recognition model;
[0095] S22. Use the optimized recognition model to perform inference on the input command word corpus, generate a feature vector corresponding to each audio segment, that is, the embedding representation.
[0096] In this embodiment, in the process of inputting the real command word corpus into the optimized recognition model to extract the embedding representation of each audio segment, it is first necessary to obtain the preset real command word corpus. These corpora contain command word audio segments in actual application scenarios, including the collected corpora and the corpora synthesized based on actual applications; use the optimized recognition model to perform inference on the input command word corpus, generate a feature vector corresponding to each audio segment, that is, the embedding representation. Specifically, the model generates a series of high-dimensional feature vectors through forward propagation, and these vectors represent the feature representations of the audio segment at different time steps. To ensure high-quality embedding representations, the system calculates the score of each path and finds the best alignment path between the audio and phonemes through the backtracking algorithm. Based on this best alignment path, the audio embedding representation corresponding to each phoneme can be extracted. For each audio segment, the probability matrix generated by the model with the time steps matching the audio length is used to determine the best alignment path, so as to generate high-quality embedding representations.
[0097] In addition, to further improve the quality of the embedding representation, the system also sets a quality threshold. Only when the embedding representation score of the audio segment is higher than this threshold, it is considered a high-quality embedding representation. This step ensures that only those embedding representations with higher confidence are used to update the phoneme embedding representation dictionary subsequently. In this way, the system can efficiently and accurately process the newly input audio data and provide a solid foundation for constructing the phoneme embedding representation dictionary subsequently.
[0098] In one embodiment, the step of constructing the embedding representation dictionary of each phoneme according to the extracted embedding representation includes:
[0099] S30. For the audio segments included in the real command word corpus, calculate the path score based on the ctc algorithm;
[0100] S31. Based on the path score, use the backtracking algorithm to find the best alignment path between the audio and phonemes by selecting the maximum path score;
[0101] S32. Create an initial phoneme embedding representation dictionary, where each phoneme corresponds to an initial embedding representation;
[0102] S33. When the embedding representation score of the audio segment is greater than the preset threshold, it is recognized as a high-quality audio embedding representation;
[0103] S34. Map the high-quality audio embedding representation to the corresponding phonemes according to the optimal alignment path, and update the embedding representation of the corresponding phonemes by the method of moving average to obtain the embedding representation dictionary.
[0104] In this embodiment, in the process of constructing the embedding representation dictionary for each phoneme based on the extracted embedding representation, it is first necessary to process the audio segments in the real command word corpus. Specifically, for these audio segments containing command words, calculate the path scores based on the CTC (Connectionist Temporal Classification) algorithm. The CTC algorithm can handle the alignment problem of variable-length sequences and is particularly suitable for the alignment of audio and phonemes in speech recognition tasks. Next, based on the calculated path scores, use the backtracking algorithm to find the optimal alignment path between the audio and phonemes by selecting the maximum path score. The optimal alignment path refers to the path with the highest score among all possible paths, representing the most likely correspondence between the audio sequence and the phonemes. For example, assume that the optimal alignment path of an audio segment is ["er4", "shi2", "qi1"], then the correspondence between this audio segment and the phonemes is determined.
[0105] Then, create an initial phoneme embedding representation dictionary, where each phoneme corresponds to an initial embedding representation. This initial dictionary can be initialized based on the embedding representations generated by a pre-trained model or randomly initialized. For example, assume that we have a list containing common phonemes (such as "sh", "ang", "yi", etc.), and we can assign an initial embedding vector to each phoneme. These initial embedding vectors can be obtained from a pre-trained language model or randomly generated high-dimensional vectors.
[0106] To ensure high-quality embedding representations, the system sets a quality threshold, and only when the embedding representation score of the audio segment is greater than this threshold, it is recognized as a high-quality audio embedding representation. This step ensures that only those embedding representations with higher confidence are used to update the phoneme embedding representation dictionary. For example, if the embedding representation score of an audio segment is 0.9 and the threshold is set to 0.8, then the embedding representation of this segment will be recognized as high-quality and used for subsequent updates.
[0107] In this way, the system can efficiently and accurately construct a high-quality phoneme embedding representation dictionary, providing a solid foundation for subsequent steps. This method not only improves the recognition accuracy of the system but also enhances the adaptability and flexibility of the system, enabling the system to operate stably in a complex speech environment, improving the user experience, and being widely applied to various application scenarios such as smart homes and smart terminals. Ultimately, this high-quality phoneme embedding representation ensures the efficiency and reliability of the speech recognition system in practical applications.
[0108] In one embodiment, the step of generating an embedding sequence of the corresponding audio by optimizing the recognition model when a new voice input is received includes:
[0109] S40. Preprocess the voice input to obtain an audio segment to be recognized;
[0110] S41. Input the audio segment to be recognized into the optimized recognition model, and generate an embedding representation of the audio segment to be recognized through forward propagation, that is, the embedding sequence of the corresponding audio.
[0111] In this embodiment, first, the received new voice input is preprocessed to obtain an audio segment to be recognized. The preprocessing steps usually include operations such as noise reduction, normalization, and segmentation to ensure the quality and consistency of the audio data. Next, the preprocessed audio segment to be recognized is input into an optimized recognition model that has been trained and fine-tuned. This optimized recognition model has been fully trained on general corpora and command word corpora and has performed well on the validation set. After loading the model, the audio segment to be recognized is passed as input to the model. The model will gradually process this data according to its internal structure and generate a series of high-dimensional feature vectors through forward propagation. These vectors represent the feature representations of the audio segment at different time steps. Specifically, the model generates an embedding representation at each time step, usually extracting the features before the output layer of the model (such as the features before the linear layer or the fully connected layer). To ensure that the generated embedding representation can be aligned with the phoneme embedding representation dictionary in the subsequent steps, some additional processing steps may be required. For example, using the backtracking method of the CTC algorithm to align the audio segment and the phoneme sequence, or applying a certain smoothing technique to reduce the influence of noise. This step helps to improve the quality of the embedding representation and makes the subsequent similarity calculation more accurate. Finally, the embedding sequence generated by the model can be represented as [E_1, E_2, ..., E_n], where E_i represents the embedding vector at the i-th time step. In this way, the system can efficiently and accurately process the newly input audio data and generate high-quality embedding representations. These embedding representations not only reflect the features of the audio segment but also are aligned with specific phonemes, thus providing accurate data support for the subsequent steps. For example, in the subsequent steps, the system can accurately distinguish these confusing command words by calculating the similarity between the embedding sequence of the newly input audio and the phoneme embedding representation sequences of the confusing command words, significantly improving the accuracy and reliability of recognition. In summary, by preprocessing the received new voice input and using the optimized recognition model to generate high-quality embedding sequences, the system can efficiently process various complex voice inputs in practical applications. This method not only improves the recognition accuracy of the system but also enhances the adaptability and flexibility of the system, enabling the system to operate stably in a complex voice environment, improving the user experience, and being widely applied to various application scenarios such as smart homes and smart terminals. Finally, this high-quality embedding representation ensures the efficiency and reliability of the speech recognition system in practical applications.
[0112] In one embodiment, the step of obtaining the phoneme embedding representation sequences of all confusing command words in the confusing command word list corresponding to the embedding sequence based on the embedding representation dictionary includes:
[0113] S50. Identify the confusing command word list corresponding to the embedding sequence and load all confusing command words in the confusing command word list;
[0114] S501 converts the confusing command words into corresponding phoneme sequences;
[0115] S52. According to each phoneme sequence, obtain the corresponding phoneme embedding representation from the pre-constructed phoneme embedding representation dictionary, and obtain the corresponding phoneme embedding representation sequence.
[0116] In this embodiment, first, identify the list of confusing command words corresponding to the embedded sequence and load all the confusing command words in this list. This step involves determining the set of confusing command words that the current input audio may correspond to. For example, in a smart home control system, common confusing command words may include "turn on the light" and "turn off the light", "raise the temperature" and "lower the temperature", etc. By analyzing the context of the input audio or preset rules, the system can identify a set of potential confusing command words. Next, convert the confusing command words into corresponding phoneme sequences. This process usually needs to be completed using a phoneme mapping tool (such as G2P, Grapheme-to-Phoneme converter). For example, "twenty-seven degrees" can be converted into the phoneme sequence ["er4", "shi2", "qi1", "du4"], and "twenty-one degrees" can be converted into the phoneme sequence ["er4", "yi2", "shi2", "du4"]. The phoneme sequence provides a pronunciation feature representation of each command word, enabling subsequent steps to calculate similarity through comparing phoneme embedding representations. Then, according to each phoneme sequence, obtain the corresponding phoneme embedding representation from the pre-constructed phoneme embedding representation dictionary, resulting in a corresponding phoneme embedding representation sequence. The specific steps are as follows: For each phoneme, look up its corresponding embedding vector in the embedding representation dictionary. For example, assume that the phoneme embedding representation dictionary contains the embedding vector E_{er4} for the phoneme "er4", the embedding vector E_{shi2} for the phoneme "shi2", etc. In the order of the phoneme sequence, concatenate the embedding representation vectors of each phoneme to form the complete phoneme embedding representation sequence of this command word. For example, the phoneme embedding representation sequence of "twenty-seven degrees" is [E_{er4}, E_{shi2}, E_{qi1}, E_{du4}], and the phoneme embedding representation sequence of "twenty-one degrees" is [E_{er4}, E_{yi2}, E_{shi2}, E_{du4}]. In this way, the system provides an accurate phoneme embedding representation sequence for the subsequent steps. These phoneme embedding representation sequences will be used in the next step to calculate similarity scores to distinguish confusing command words. For example, assume that the user issues a voice command "twenty-seven degrees". The system first preprocesses the received audio signal and generates an embedded sequence [E_1, E_2, ..., E_n] through an optimized speech recognition model. Then, the system identifies the list of potential confusing command words [[twenty-seven degrees, twenty-one degrees], [previous song, next song]], and converts "twenty-seven degrees" and "twenty-one degrees" into phoneme sequences ["er4", "shi2", "qi1", "du4"] and ["er4", "yi2", "shi2", "du4"] respectively.Then, the system obtains the phoneme embedding representation sequences of these two command words from the phoneme embedding representation dictionary: the phoneme embedding representation sequence of "twenty-seven degrees" is [E_{er4}, E_{shi2}, E_{qi1}, E_{du4}]; the phoneme embedding representation sequence of "twenty-one degrees" is [E_{er4}, E_{yi2}, E_{shi2}, E_{du4}]; these phoneme embedding representation sequences will be used for similarity calculation in the subsequent steps. By calculating the similarity between the embedding sequence of the input audio and these phoneme embedding representation sequences, the system can accurately identify whether the command word issued by the user is "twenty-seven degrees" or "twenty-one degrees".
[0117] In summary, by identifying the list of confusing command words corresponding to the embedding sequence, converting these command words into phoneme sequences, and then obtaining the corresponding phoneme embedding representation sequences from the embedding representation dictionary, the system can efficiently process newly input audio data in practical applications and distinguish confusing command words by calculating similarity. This method significantly improves the accuracy, adaptability, and flexibility of the speech recognition system, enabling the system to operate stably in complex speech environments, enhancing the user experience, and being widely applied to various application scenarios such as smart homes and smart terminals. Ultimately, this high-quality phoneme embedding representation ensures the efficiency and reliability of the speech recognition system in practical applications.
[0118] In one embodiment, the step of calculating the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the confusing command word to obtain the final recognition result includes:
[0119] S61. Align the embedding sequence of the input audio with the phoneme embedding representation sequence of the confusing command word;
[0120] S62. Use a similarity metric method to calculate the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of each confusing command word;
[0121] S63. Compare all similarity scores and select the command word with the highest score as the final recognition result.
[0122] In this embodiment, the process of calculating the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the confusing command words to obtain the final recognition result is as follows: First, use the Dynamic Time Warping (DTW) algorithm to align the embedding sequence of the input audio with the phoneme embedding representation sequence of each confusing command word to ensure that the two are compared at the same time step. Then, use similarity measurement methods such as cosine similarity or Euclidean distance to calculate the similarity score between the two. For example, assume that the embedding sequence of the input audio is [E_1, E_2, ..., E_n], and it is compared with the phoneme embedding representation sequences of "twenty-seven degrees" and "twenty-one degrees" respectively to obtain their respective similarity scores. Finally, the system compares all similarity scores and selects the command word with the highest score as the final recognition result. Through this method, the system can efficiently and accurately process newly input audio data, distinguish confusing command words, and significantly improve the accuracy, adaptability, and user experience of the speech recognition system.
[0123] Referring to Figure 2 , which is the structural block diagram of the confusing word confirmation device in an embodiment of the present application. The device includes:
[0124] An optimization module 100, configured to train an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model;
[0125] An extraction module 200, configured to input the real command word corpus into the optimized recognition model to extract the embedding representation of each audio segment;
[0126] A construction module 300, configured to construct an embedding representation dictionary for each phoneme according to the extracted embedding representation;
[0127] A generation module 400, configured to generate an embedding sequence of the corresponding audio through the optimized recognition model when a new speech input is received;
[0128] An acquisition module 500, configured to obtain the phoneme embedding representation sequences of all confusing command words in the confusing command word list corresponding to the embedding sequence based on the embedding representation dictionary;
[0129] A calculation module 600, configured to calculate the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the confusing command words to obtain the final recognition result.
[0130] In one embodiment, the above optimization module 100 includes a training unit, configured to:
[0131] Train a speech recognition model based on a general corpus until the loss rate or word error rate on the validation set is reduced to a first preset threshold to obtain the initial speech recognition model;
[0132] Train on the basis of a preset command word corpus on the initial speech recognition model until the loss rate or word error rate of the initial speech recognition model on the validation set containing command words is reduced to a second preset threshold to obtain the optimized recognition model.
[0133] In one embodiment, the above extraction module 200 includes an inference unit for:
[0134] Obtain a preset real command word corpus, where the real command word corpus is an audio segment containing command words;
[0135] Input the real command word corpus into the optimized recognition model;
[0136] Use the optimized recognition model to infer the input command word corpus and generate a feature vector corresponding to each audio segment, that is, the embedding representation.
[0137] In one embodiment, the above construction module 300 includes an update unit for:
[0138] For the audio segments included in the real command word corpus, calculate the path score based on the ctc algorithm;
[0139] Based on the path score, use the backtracking algorithm to find the best alignment path between the audio and phonemes by selecting the maximum path score;
[0140] Create an initial phoneme embedding representation dictionary, where each phoneme corresponds to an initial embedding representation;
[0141] When the embedding representation score of the audio segment is greater than the preset threshold, it is determined as a high-quality audio embedding representation;
[0142] According to the best alignment path, map the high-quality audio embedding representation to the corresponding phoneme, and update the embedding representation of the corresponding phoneme by the method of moving average to obtain the embedding representation dictionary.
[0143] In one embodiment, the above generation module 400 includes a sequence generation unit for:
[0144] Preprocess the speech input to obtain an audio segment to be recognized;
[0145] Input the audio segment to be recognized into the optimized recognition model, and generate an embedding representation of the audio segment to be recognized through forward propagation, that is, the embedding sequence of the corresponding audio.
[0146] In one embodiment, the above acquisition module 500 includes a phoneme embedding acquisition unit for:
[0147] Identify the list of confusing command words corresponding to the embedded sequence, and load all the confusing command words in the list of confusing command words;
[0148] Convert the confusing command words into corresponding phoneme sequences;
[0149] According to each phoneme sequence, obtain the corresponding phoneme embedding representation from the pre-constructed phoneme embedding representation dictionary, and obtain the corresponding phoneme embedding representation sequence.
[0150] In one embodiment, the above computing module 600 includes a result output unit for:
[0151] Align the embedded sequence of the input audio with the phoneme embedding representation sequences of the confusing command words;
[0152] Use a similarity metric method to calculate the similarity between the embedded sequence of the input audio and the phoneme embedding representation sequences of each confusing command word;
[0153] Compare all similarity scores, and select the command word with the highest score as the final recognition result.
[0154] Refer to Figure 3 , in the embodiments of the present application, a computer device is further provided. The computer device may be a server, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the usage data during the process of the confusing word confirmation method, etc. The network interface of the computer device is used to communicate with an external terminal through a network connection. Further, the above computer device may also be provided with an input device and a display screen, etc. When the above computer program is executed by the processor, it realizes the confusing word confirmation method, including the following steps: training on an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model; inputting the real command word corpus into the optimized recognition model, and extracting the embedded representation of each audio segment; constructing a phoneme embedding representation dictionary according to the extracted embedded representation; when receiving new speech input, generating an embedded sequence of the corresponding audio through the optimized recognition model; based on the embedded representation dictionary, obtaining the phoneme embedding representation sequences of all the confusing command words in the list of confusing command words corresponding to the embedded sequence; calculating the similarity between the embedded sequence of the input audio and the phoneme embedding representation sequences of the confusing command words to obtain the final recognition result. Those skilled in the art can understand, Figure 3The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied.
[0155] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, an easily confused word confirmation method is implemented, including the following steps: training on an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model; inputting a real command word corpus into the optimized recognition model to extract the embedded representation of each audio segment; constructing an embedded representation dictionary for each phoneme according to the extracted embedded representation; when a new speech input is received, generating an embedded sequence corresponding to the corresponding audio through the optimized recognition model; based on the embedded representation dictionary, obtaining the phoneme embedded representation sequences of all easily confused command words in the easily confused command word list corresponding to the embedded sequence; calculating the similarity between the embedded sequence of the input audio and the phoneme embedded representation sequences of the easily confused command words to obtain a final recognition result. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0156] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0157] It should be noted that in this text, the terms "including", "comprising", or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, apparatus, article, or method that includes such an element.
[0158] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. A method for identifying easily confused words, characterized in that, The method includes: Training an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model; Inputting a real command word corpus into the optimized recognition model to extract the embedding representation of each audio segment; Constructing an embedding representation dictionary for each phoneme according to the extracted embedding representations; When receiving a new speech input, generating an embedding sequence corresponding to the corresponding audio through the optimized recognition model; Based on the embedding representation dictionary, obtaining the phoneme embedding representation sequences of all easily confused command words in the list of easily confused command words corresponding to the embedding sequence; Calculating the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequences of the easily confused command words to obtain the final recognition result; The step of constructing an embedding representation dictionary for each phoneme according to the extracted embedding representations includes: For the audio segments included in the real command word corpus, calculating the path score based on the ctc algorithm; Based on the path score, using the backtracking algorithm to find the best alignment path between the audio and phonemes by selecting the maximum path score; Creating an initial phoneme embedding representation dictionary, where each phoneme corresponds to an initial embedding representation; When the embedding representation score of an audio segment is greater than a preset threshold, it is recognized as a high-quality audio embedding representation; According to the best alignment path, mapping the high-quality audio embedding representation to the corresponding phoneme, and updating the embedding representation of the corresponding phoneme by the method of moving average to obtain the embedding representation dictionary.
2. The method for confirming homophonic and polysemous words according to claim 1, wherein The step of training an initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model includes: Training a speech recognition model based on a general corpus until the loss rate or word error rate on the validation set is reduced to a first preset threshold to obtain the initial speech recognition model; Training based on a preset command word corpus on the basis of the initial speech recognition model until the loss rate or word error rate of the initial speech recognition model on the validation set containing command words is reduced to a second preset threshold to obtain the optimized recognition model.
3. The method for confirming confusing words according to claim 1, wherein The step of inputting a real command word corpus into the optimized recognition model to extract the embedding representation of each audio segment includes: Obtaining a preset real command word corpus, where the real command word corpus is an audio segment containing command words; Inputting the real command word corpus into the optimized recognition model; Using the optimized recognition model to perform inference on the input command word corpus to generate a feature vector corresponding to each audio segment, that is, the embedding representation.
4. The method for identifying easily confused words according to claim 1, wherein The step of generating an embedding sequence corresponding to the corresponding audio through the optimized recognition model when receiving a new speech input includes: Preprocessing the speech input to obtain an audio segment to be recognized; Inputting the audio segment to be recognized into the optimized recognition model, and generating an embedding representation of the audio segment to be recognized through forward propagation, that is, the embedding sequence corresponding to the corresponding audio.
5. The method for identifying easily confused words according to claim 1, characterized in that The step of obtaining the phoneme embedding representation sequences of all easily confused command words in the list of easily confused command words corresponding to the embedding sequence based on the embedding representation dictionary includes: Identifying the list of easily confused command words corresponding to the embedding sequence, and loading all the easily confused command words in the list of easily confused command words; Convert the confusing command words into corresponding phoneme sequences; According to each phoneme sequence, obtain the corresponding phoneme embedding representation from the pre-constructed phoneme embedding representation dictionary, and obtain the corresponding phoneme embedding representation sequence.
6. The method for identifying easily confused words according to claim 1, wherein The step of calculating the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the confusing command words to obtain the final recognition result includes: Align the embedding sequence of the input audio with the phoneme embedding representation sequence of the confusing command words; Use a similarity metric method to calculate the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of each confusing command word; Compare all similarity scores and select the command word with the highest score as the final recognition result.
7. An easily confused word confirmation device for performing the method according to any one of claims 1-6, characterized in that, Including: An optimization module for training on the initial speech recognition model based on a preset command word corpus to obtain an optimized recognition model; An extraction module for inputting the real command word corpus into the optimized recognition model to extract the embedding representation of each audio segment; A construction module for constructing a phoneme embedding representation dictionary according to the extracted embedding representation; A generation module for generating the embedding sequence of the corresponding audio through the optimized recognition model when receiving a new speech input; An acquisition module for obtaining the phoneme embedding representation sequences of all confusing command words in the list of confusing command words corresponding to the embedding sequence based on the embedding representation dictionary; A calculation module for calculating the similarity between the embedding sequence of the input audio and the phoneme embedding representation sequence of the confusing command words to obtain the final recognition result.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Re-recognizing speech with external data sources
CN107045871A
Voice instruction processing method and device, equipment and storage medium
CN112599127A
Speech recognition model training method and device and computer equipment
CN113870844A
Language classification method and device and computer readable storage medium
CN115132170A
Mixing identification processing method and device, equipment and medium
CN119600997A