Method and device for improving speech recognition accuracy, computer equipment and medium
By extracting voice features and using voice classification models to distinguish between human voice and interfering sound, and combining voiceprint matching to verify user identity, the problem of inaccurate recognition of speech recognition system in complex environments is solved, achieving higher recognition accuracy and security.
Patent Information
- Application Number
- CN202510299867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
Existing voice recognition systems are difficult to accurately recognize voice commands in complex environments, and are greatly affected by the differences in interfering sounds and the voices of different speakers.
By obtaining the voice data to be identified, extracting its voice features, and inputting these features into a pre-trained voice classification model to distinguish vocal data from other interfering sounds. Then, the voice data is matched with the voiceprint characteristics of the pre-registered user, and voice commands are recognized and executed only when the match is successful.
It effectively filters out interference factors in complex voice environments, improves the accuracy and stability of speech recognition, and ensures the accuracy and security of voice interactive services.
Smart Images

Figure CN120148490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a method, apparatus, computer device, and storage medium for improving the accuracy of speech recognition. Background Art
[0002] With the development of intelligent technology, the application of speech recognition technology has become increasingly widespread, such as in-vehicle systems, conference systems, etc. Currently, speech recognition systems mainly rely on automatic speech recognition (ASR) models, aiming to convert the recognized speech commands into executable operations, such as navigation settings, music playback control, conference recording, etc., to provide users with a convenient interaction experience.
[0003] However, since the speech in the environment is usually rich and diverse, in addition to the speech commands of the speaker, there are also various interference sounds. The voice differences of different speakers (such as accents, intonations, volumes, etc.) themselves increase the recognition difficulty. At the same time, background sounds such as song sounds, audiobook sounds, navigation broadcast sounds, and engine noise sounds will seriously interfere with the accuracy of speech recognition. These interference sounds are mixed with the speech commands of the speaker, making it difficult for the ASR model to accurately extract effective speech features, resulting in deviations in the recognition results. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, apparatus, computer device, and storage medium for improving the accuracy of speech recognition to solve at least one of the above problems existing in the prior art.
[0005] In a first aspect, a method for improving the accuracy of speech recognition is provided, including:
[0006] Obtain the speech data to be recognized, and extract the speech features corresponding to the speech data to be recognized;
[0007] Input the speech features into a pre-trained speech classification model to obtain human voice data;
[0008] Match the human voice data with the voiceprint features of pre-registered users;
[0009] If the match is successful, recognize the speech command corresponding to the human voice data, and perform corresponding operations based on the speech command.
[0010] In an embodiment, the extracting the speech features corresponding to the speech data to be recognized includes:
[0011] Preprocess the speech data to be recognized;
[0012] Extract speech features from the preprocessed speech data to be recognized, where the speech features include any one or any combination of spectral features, Mel spectral features, rhythm features, and intonation features.
[0013] In one embodiment, the preprocessing of the speech data to be recognized includes:
[0014] Perform pre-emphasis processing on the speech data to be recognized;
[0015] Perform frame segmentation on the pre-emphasized speech data to be recognized to obtain multiple audio frames;
[0016] Perform windowing processing on each audio frame through a preset windowing algorithm.
[0017] In one embodiment, the recognizing the speech command corresponding to the human voice data and performing corresponding operations based on the speech command includes:
[0018] Perform security detection on the speech command;
[0019] If there is a risk in the speech command, obtain the user's biometric feature corresponding to the speech command;
[0020] Match the user's biometric feature with the biometric features of the users whose voiceprint features match successfully;
[0021] If the match is successful, perform corresponding operations based on the speech command.
[0022] In one embodiment, there are multiple pieces of human voice data, and the recognizing the speech command corresponding to the human voice data and performing corresponding operations based on the speech command:
[0023] Recognize the speech command corresponding to each piece of human voice data;
[0024] Determine the permission priority of the user corresponding to each piece of human voice data;
[0025] Execute the speech command based on the permission priority.
[0026] In one embodiment, the matching of the human voice data with the voiceprint features of the pre-registered users includes:
[0027] Extract the voiceprint features corresponding to the human voice data;
[0028] Determine the similarity value between the extracted voiceprint features and the voiceprint features of the pre-registered users;
[0029] When the similarity value is greater than the preset similarity threshold, it indicates that the human voice data matches the voiceprint data of the pre-registered user successfully.
[0030] In one embodiment, before extracting the voice features corresponding to the voice data to be recognized, the following steps are included:
[0031] Convert the voice data to be recognized into text data;
[0032] Perform keyword detection on the text data;
[0033] If there are keywords in the text data that are associated with a preset voice command, then perform the step of extracting the voice features corresponding to the voice data to be recognized.
[0034] In a second aspect, a device for improving the accuracy of voice recognition is provided, including:
[0035] A voice feature extraction unit, configured to obtain voice data to be recognized and extract the voice features corresponding to the voice data to be recognized;
[0036] A human voice data acquisition unit, configured to input the voice features into a pre-trained voice classification model to obtain human voice data;
[0037] A voiceprint feature matching unit, configured to match the human voice data with the voiceprint features of a pre-registered user;
[0038] A voice command recognition and execution unit, configured to, if the matching is successful, recognize the voice command corresponding to the human voice data and perform a corresponding operation based on the voice command.
[0039] In a third aspect, a computer device is provided, including a memory, a processor, and computer-readable instructions stored on the memory and running on the processor. When the processor executes the computer-readable instructions, the method for improving the accuracy of voice recognition as described above is implemented.
[0040] In a fourth aspect, a readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the method for improving the accuracy of voice recognition as described above is implemented.
[0041] The above method, device, computer equipment and storage medium for improving speech recognition accuracy, the implementation of the method includes: obtaining speech data to be recognized, and extracting speech features corresponding to the speech data to be recognized; inputting the speech features into a pre-trained speech classification model to obtain human voice data; matching the human voice data with the voiceprint features of a pre-registered user; if the match is successful, recognizing the speech command corresponding to the human voice data, and performing corresponding operations based on the speech command. In the embodiments of the present application, by obtaining speech data to be recognized and extracting its corresponding speech features, and inputting these features into a pre-trained speech classification model, it is possible to accurately distinguish human voice data from other interfering sounds (such as song sounds, navigation broadcast sounds, engine noise sounds, etc.), so as to obtain pure human voice data. The interference factors in the complex speech environment are effectively filtered, laying a foundation for accurately recognizing speech commands subsequently. Matching the obtained human voice data with the voiceprint features of a pre-registered user, only when the match is successful, will the speech command be further recognized and the operation be performed. Voiceprint matching, as an identity verification mechanism, can effectively exclude the speech interference of non-target users, and by using the unique voiceprint features of the target user and combining with the speech recognition model, it can better adapt to the speech characteristics of the target user, improve the recognition accuracy of diverse commands of the target user, and further increase the hit rate of commands, providing a more stable and accurate speech interaction service for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 is a flowchart of a method for improving speech recognition accuracy in an embodiment of the present invention;
[0044] Figure 2 is a structural schematic diagram of a device for improving speech recognition accuracy in an embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] In one embodiment, as Figure 1 shown, a method for improving the accuracy of speech recognition is provided, including the following steps:
[0048] In step S110, the speech data to be recognized is obtained, and the speech features corresponding to the speech data to be recognized are extracted;
[0049] It should be noted that the speech features may include any one or any combination of spectral features, Mel spectral features, rhythm features, and intonation features.
[0050] Optionally, after obtaining the speech data to be recognized, the speech data to be recognized can be pre-emphasized, framed, and windowed to obtain a preprocessed audio signal. Then, speech features are extracted from the audio signal. Taking the Mel spectral feature as an example, a Mel spectral filter can be constructed. The center frequencies of the Mel filters are evenly distributed on the Mel frequency axis, and then these Mel filters are mapped back to the linear frequency axis. The Fourier transform can be performed on the preprocessed audio signal to obtain a spectrum, and then the spectrum is convolved with the Mel spectral filter to obtain a Mel spectrum. Each Mel filter performs weighted summation on the spectrum to obtain the energy value of the corresponding Mel frequency band. And the discrete cosine transform is performed on the energy value of the Mel frequency band to obtain the Mel spectral feature.
[0051] In step S120, the speech features are input into a pre-trained speech classification model to obtain human voice data;
[0052] It should be noted that the speech classification model can be a convolutional neural network (CNN), a recurrent neural network (RNN), or a long short-term memory network (LSTM) and a Transformer model. Among them, the long short-term memory network (LSTM) and the Transformer model perform excellently in processing sequence data (such as speech).
[0053] Optionally, if the voice classification model is a convolutional neural network, the convolutional neural network includes multiple stacked convolutional layers. Each convolutional layer performs a convolution operation on the voice features through a convolution kernel to extract local voice features. The weights of the convolution kernels can automatically learn different voice features during the training process. Different convolution kernels can capture different voice features. Then, pooling processing is performed on the voice features after convolution processing, such as max pooling or average pooling. Then, the voice features after pooling processing are input into a fully connected layer for processing and then into the output layer. Through an activation function, such as the Softmax activation function, the feature vector output by the fully connected layer is converted into a probability distribution of various types of sounds (human voice, navigation broadcast sound, in-vehicle background sound, app broadcast sound), and the output dimension is the number of sound categories. For example, if there are 4 types of sounds, a 4-dimensional vector is output, and the value of each dimension represents the probability of the corresponding category of sound. According to the output probability distribution, the category with the highest probability is selected as the prediction result. If the prediction result is the human voice category, then this part of the voice data is recognized as human voice data, thereby realizing the extraction of human voice data.
[0054] Optionally, if the voice classification model is a Transformer model. Then the voice features can be constructed into a voice feature sequence and input into the Transformer model. First, the input voice feature sequence is subjected to attention calculation through the multi-head attention mechanism, and then the calculated voice features are input into the feed-forward neural network layer for processing, which usually consists of two fully connected layers, and the ReLU activation function is used in the middle. The feed-forward neural network layer performs further non-linear transformation and feature extraction on the features output by the self-attention mechanism. Finally, through the activation function preset in the output layer, such as Softmax, activation processing is performed to output the probability distribution of various types of sounds (human voice, navigation broadcast sound, in-vehicle background sound, app broadcast sound), and the output dimension is the number of sound categories. For example, if there are 4 types of sounds, a 4-dimensional vector is output, and the value of each dimension represents the probability of the corresponding category of sound. According to the output probability distribution, the category with the highest probability is selected as the prediction result. If the prediction result is the human voice category, then this part of the voice data is recognized as human voice data, thereby realizing the extraction of human voice data.
[0055] It should be noted that the training process of the above voice classification model can be specifically as follows: Collect a large number of voice sample data such as human voices, navigation broadcasts, in-vehicle background sounds, conference background sounds, app broadcast sounds, etc. Then, through manual annotation, voice category annotation can be performed on the voice sample data to distinguish different types of sounds. Then, a training data set, a test data set, and a validation data set are constructed based on the annotated voice sample data. The voice classification model is iteratively trained based on the training data set until the model converges, for example, the number of iterations reaches a preset number, or the loss value is less than a preset loss threshold. The iterative result can obtain a trained voice classification model.
[0056] It can be understood that the speech classification model after training can also be evaluated and optimized based on the test set. For example, the test set can be further subdivided and grouped according to different environmental types (such as different levels of noise environments), speaker attributes (such as different genders, age groups), etc., so as to analyze the performance of the model in different sub-scenarios targeted, such as the model performance test in scenarios of different background noise conditions, speaker changes, etc. And indicators such as accuracy, recall rate, and F1 value can be used to evaluate the accuracy of the classification model. The accuracy measures the proportion of samples correctly classified by the model in the total samples; the recall rate reflects the ability of the model to correctly identify positive samples (such as human voices); the F1 value comprehensively considers the accuracy and recall rate and evaluates the model performance more comprehensively. At the same time, hyperparameter tuning techniques, such as grid search, random search, or Bayesian optimization, can be used to optimize the hyperparameters of the classification model.
[0057] In step S130, the human voice data is matched with the voiceprint features of the pre-registered user.
[0058] Optionally, the target user can perform voiceprint registration in advance, such as drivers, conference participants, speakers, etc., so as to extract the voice belonging to the target user subsequently. Exemplarily, taking a driver as an example, the voiceprint feature information of the driver can be input in advance. For example, the driver can be allowed to read a series of specific sentences in a quiet environment, collect his speech data, and use feature extraction methods such as Mel Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficients (LPCC), etc. to obtain the feature vector that can represent the unique voiceprint of the driver. Then, voiceprint recognition models such as Gaussian Mixture Model-Universal Background Model (GMM-UBM), Deep Neural Network (DNN), etc. can be used to train the voiceprint features of the driver and construct a driver voiceprint template. When the human voice data is determined through the speech classification model, the human voice data can be subjected to voiceprint extraction, and the extracted voiceprint is compared with the voiceprint template. If the comparison is consistent, it means that subsequent operations can be performed based on the human voice data. Otherwise, the human voice data is filtered.
[0059] It should be noted that when registering the voiceprint features of the target user, the collection of voiceprint data in different environments can be increased to make the voiceprint template more robust. For example, collect voiceprints under different noise levels and different microphone devices, so that the model can learn more comprehensive user voiceprint features and reduce the matching failure caused by environmental changes.
[0060] It should be noted that there can be multiple voiceprint features of the pre-registered user. For multi-user collaborative operation scenarios, such as in a scenario where multiple people are riding in a car and multiple people jointly issue voice commands to control in-vehicle devices, the voice commands of different users can be accurately identified and corresponding processing can be performed.
[0061] In step S140, if the matching is successful, the voice command corresponding to the voice data is recognized, and the corresponding operation is performed based on the voice command.
[0062] Optionally, if the matching is successful, it means that the voice data is the voice of a user who has been pre-registered for voice control. At this time, the voice data can be input into the automatic speech recognition (ASR) system, and the voice command is recognized by the ASR system, and the corresponding operation is performed based on the voice command.
[0063] Exemplarily, this method can be applied to various scenarios. For example, in an audio-visual recommendation platform, corresponding audio-visual data can be played according to the recognized control command; in an intelligent voice assistant system, corresponding response information can be output or corresponding devices can be turned on; in a vehicle-mounted system, the conversation sounds of passengers and navigation broadcast sounds can be filtered, and the voice of the driver can be accurately recognized to perform operations such as music playback and air-conditioning temperature adjustment; in a meeting scenario, the voices of different participants can be automatically recognized and distinguished, and corresponding meeting records can be generated for different participants. Moreover, this method is applicable to scenarios such as background noise suppression, security monitoring, smart home control, audio analysis, data cleaning, content review, copyright detection, automatic annotation and classification, sentiment analysis, translation and transcription, personalized playlist generation, advertisement insertion, and recommendation discovery.
[0064] The embodiment of the present application provides a method for improving the accuracy of speech recognition. The implementation of the method includes: obtaining the speech data to be recognized, and extracting the speech features corresponding to the speech data to be recognized; inputting the speech features into a pre-trained speech classification model to obtain voice data; matching the voice data with the voiceprint features of a pre-registered user; if the matching is successful, the voice command corresponding to the voice data is recognized, and the corresponding operation is performed based on the voice command. In the embodiment of the present application, by obtaining the speech data to be recognized and extracting its corresponding speech features, and inputting these features into a pre-trained speech classification model, it is possible to accurately distinguish voice data from other interfering sounds (such as song sounds, navigation broadcast sounds, engine noise sounds, etc.), so as to obtain pure voice data. The interference factors in the complex speech environment are effectively filtered, laying a foundation for accurately recognizing voice commands subsequently. And the obtained voice data is matched with the voiceprint features of a pre-registered user. Only when the matching is successful, the voice command is further recognized and the operation is performed. Voiceprint matching, as an identity verification mechanism, can effectively exclude the voice interference of non-target users, and by using the unique voiceprint features of the target user and combining with the speech recognition model, it can better adapt to the speech characteristics of the target user, improve the recognition accuracy of diverse commands of the target user, and thus increase the hit rate of commands, providing a more stable and accurate voice interaction service for users.
[0065] In an embodiment of the present application, extracting the speech features corresponding to the speech data to be recognized includes:
[0066] Preprocessing the speech data to be recognized;
[0067] Performing speech feature extraction on the preprocessed speech data to be recognized, where the speech features include any one or any combination of spectral features, mel-spectral features, rhythm features, and intonation features.
[0068] Optionally, after obtaining the speech data to be recognized, the speech data to be recognized can be pre-emphasized, framed, and windowed to obtain a preprocessed audio signal. Then, speech feature extraction is performed on the audio signal to extract any one or any combination of spectral features, mel-spectral features, rhythm features, and intonation features respectively.
[0069] Taking the mel-spectral feature as an example, a mel-spectral filter can be constructed. The center frequencies of the mel filters are evenly distributed on the mel frequency axis, and then these mel filters are mapped back to the linear frequency axis. Fourier transform can be performed on the preprocessed audio signal to obtain a spectrum, and then the spectrum is convolved with the mel-spectral filter to obtain a mel spectrum. Each mel filter performs weighted summation on the spectrum to obtain the energy value of the corresponding mel frequency band. And discrete cosine transform is performed on the energy value of the mel frequency band to obtain the mel-spectral feature.
[0070] Taking the spectral feature as an example, short-time Fourier transform can be performed on the preprocessed speech data to be recognized to convert the audio signal in the time domain to the frequency domain, and a complex matrix representing different frequency components of each frame can be obtained. This is used as the spectral feature.
[0071] Taking the rhythm feature as an example, the autocorrelation function of the preprocessed audio signal can be calculated. The autocorrelation function reflects the similarity degree of the signal at different time delays. By analyzing the peak positions of the autocorrelation function, the fundamental frequency period of the speech signal can be estimated, and the change of the fundamental frequency period is closely related to the rhythm of the speech, thereby obtaining rhythm information.
[0072] Taking the intonation feature as an example, the phase information of the preprocessed speech data can be used to estimate the fundamental frequency. First, STFT is performed on the audio signal to obtain a phase spectrum. Then, through a series of processes, such as phase unwrapping and phase difference calculation, the periodicity of the phase change is found to estimate the fundamental frequency. Then, the change of the fundamental frequency over time is plotted as a curve, that is, the fundamental frequency contour. The fundamental frequency contour reflects the intonation change of the speech. For example, an ascending fundamental frequency contour may indicate an interrogative tone, and a descending fundamental frequency contour may indicate a declarative tone.
[0073] In an embodiment of the present application, preprocessing the speech data to be recognized includes:
[0074] Perform pre-emphasis processing on the speech data to be recognized;
[0075] Perform frame segmentation processing on the pre-emphasized speech data to be recognized to obtain multiple frames of audio frames;
[0076] Perform windowing processing on each frame of the audio frames through a preset windowing algorithm.
[0077] Optionally, pre-emphasis processing can be performed on the speech data to be recognized through a high-pass filter to enhance high-frequency information. Then, the pre-emphasized audio data is segmented into short-frame audio data. For example, the frame length is set to be between 10 and 40 milliseconds, and then windowing processing, such as Hamming window, is performed on each frame of the segmented audio data, thereby realizing the preprocessing of the speech data to be recognized.
[0078] In an embodiment of the present application, after recognizing the voice command corresponding to the human voice data and performing corresponding operations based on the voice command, it includes:
[0079] Perform security detection on the voice command;
[0080] If there is a risk in the voice command, obtain the user biometric characteristics of the target user;
[0081] Match the user biometric characteristics with the biometric characteristics of the user whose voiceprint characteristics match successfully;
[0082] If the match is successful, perform corresponding operations based on the voice command.
[0083] Optionally, before executing the voice command, security detection can be performed on the voice command first. For example, in a vehicle system, if the recognized voice command involves key operations for vehicle control, such as vehicle start, stop, acceleration, deceleration, steering, opening and closing doors, etc., which are commands related to vehicle driving safety or important functions. Then, the security check process can be triggered. At this time, a face recognition camera can be turned on to extract the facial features of the driver and match them with the facial features of the target user who is pre-registered and whose voiceprint matches successfully. If the match is successful, the corresponding voice command can be executed.
[0084] It is understandable that the target user can be the initiator of the voice data or the preset executor of the voice command. For example, taking the in-vehicle system as an example, the camera can be turned on to collect the facial features of the driver from multiple angles and with multiple expressions, and the clearest facial features can be selected as the basis for subsequent comparison. Then, it can be compared with the facial features of the successfully matched user. If the comparison is unsuccessful, even if the voiceprint is successfully matched, the voice command will not be executed at this time, which can avoid the vehicle being accidentally controlled by the voice of the passenger and cause driving risks. It can improve driving safety, user experience, and operation efficiency.
[0085] It should be noted that the biometric feature can also be an iris feature, a fingerprint feature, etc.
[0086] In an embodiment of the present application, there are multiple pieces of the voice data, and the corresponding voice command of the voice data is recognized, and the corresponding operation is executed based on the voice command:
[0087] Recognize the voice command corresponding to each piece of voice data;
[0088] Determine the permission priority of the user corresponding to each piece of voice data;
[0089] Execute the voice command based on the permission priority.
[0090] Optionally, in a multi-person scenario, such as a multi-person car-riding scenario, multiple voices can be collected. At this time, voiceprint recognition technology can be used to distinguish the voices of different users. Since the voiceprint features of each user are unique, the voice commands issued by different users can be accurately recognized. It should be noted that the voiceprint features of multiple users can be allowed to be registered, and a corresponding user management system can be established. Each user can have their own personalized settings and permissions. For example, different users have different operation permissions for certain functions. Therefore, when multiple pieces of voice data are obtained and multiple voiceprints are successfully matched, multiple voice commands can be obtained, and these voice commands can be arranged in the order from high to low according to the priority of the user permissions. If the high-priority voice command is to increase the temperature by 2 degrees and the low-priority voice command is to decrease the temperature by 2 degrees, the high-priority voice command will be executed.
[0091] It should be noted that the priority rules can be set based on factors such as user identity (such as the driver having a higher priority than the passenger), command type (such as safety-related commands having a higher priority), etc. When it cannot be solved by the priority, the user can be prompted by voice to negotiate a solution, or multiple solutions can be provided for the user to select the voice command to be executed.
[0092] It should be noted that in a multi-person scenario, during the execution of the voice command, it can be switched based on the voice commands of different users, that is, when the voice of a new user is detected and successfully matched, the voice command of this user can be automatically switched and executed.
[0093] In one embodiment of the present application, the matching of the voice data with the voiceprint features of a pre-registered user includes:
[0094] Extracting the voiceprint features corresponding to the voice data;
[0095] Determining a similarity value between the extracted voiceprint features and the voiceprint features of the pre-registered user;
[0096] When the similarity value is greater than a preset similarity threshold, it indicates that the voice data matches successfully with the voiceprint data of the pre-registered user.
[0097] Optionally, a target user can perform voiceprint registration in advance, such as a driver, a conference participant, a keynote speaker, etc., so as to extract the voice belonging to the target user subsequently. Exemplarily, taking a driver as an example, the voiceprint feature information of the driver can be input in advance. For example, the driver can be allowed to read a series of specific sentences in a quiet environment, and their voice data can be collected. Feature extraction methods such as Mel Frequency Cepstral Coefficients (MFCC) and Linear Predictive Cepstral Coefficients (LPCC) can be used to obtain a feature vector representing the unique voiceprint of the driver. Then, voiceprint recognition models such as Gaussian Mixture Model-Universal Background Model (GMM-UBM) and Deep Neural Network (DNN) can be used to train the voiceprint features of the driver to construct a driver voiceprint template. Then, after extracting the voiceprint features of the voice data, the log-likelihood under the GMM model can be calculated. The log-likelihood represents the likelihood that this feature belongs to this GMM model (i.e., the driver voiceprint). Specifically in the calculation, the feature vector is substituted into the probability density function of the GMM model, the contribution of each Gaussian component is calculated and summed, and then the logarithm is taken to obtain the log-likelihood. The higher the log-likelihood, the more similar the voice data is to the driver voiceprint template.
[0098] If the voiceprint feature template is a voiceprint feature mapping learned through a deep learning model, such as DNN, the matching of the voiceprint features can be calculated through methods such as Euclidean distance and cosine similarity.
[0099] Optionally, before extracting the voice features corresponding to the voice data to be recognized, it includes:
[0100] Converting the voice data to be recognized into text data;
[0101] Performing keyword detection on the text data;
[0102] If there are keywords in the text data associated with a preset voice command, then perform the step of extracting the voice features corresponding to the voice data to be recognized.
[0103] Optionally, a thesaurus containing emergency keywords (such as "help", "danger", "stop", etc.) can be established. When the recognized voice is determined to be the voice of a passenger, keyword detection is performed on the voice. The voice data is converted into text through voice recognition technology, and then a text matching algorithm in natural language processing technology is used to determine whether the text contains emergency keywords. If it does, it is determined to be an emergency conversation. If it includes keywords related to voice execution, such as turning off the voice, playing the next song, turning on the windshield wipers, etc., then the step of extracting the voice features corresponding to the voice data to be recognized is performed.
[0104] In the embodiments of the present application, by obtaining the voice data to be recognized and extracting its corresponding voice features, and inputting these features into a pre-trained voice classification model, it is possible to accurately distinguish human voice data from other interfering sounds (such as song sounds, navigation broadcast sounds, engine noise sounds, etc.), thereby obtaining pure human voice data. The interference factors in the complex voice environment are effectively filtered, laying a foundation for accurately recognizing voice commands subsequently. And the obtained human voice data is matched with the voiceprint features of the pre-registered user. Only when the match is successful, will the voice command be further recognized and the operation be executed. Voiceprint matching, as an identity verification mechanism, can effectively exclude the voice interference of non-target users, and by using the unique voiceprint features of the target user and combining with the voice recognition model, it can better adapt to the voice characteristics of the target user, improve the recognition accuracy of diverse commands of the target user, and further increase the hit rate of commands, providing a more stable and accurate voice interaction service for users.
[0105] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0106] In one embodiment, a device for improving voice recognition accuracy is provided. The device for improving voice recognition accuracy corresponds one-to-one with the method for improving voice recognition accuracy in the above embodiments. As Figure 2 shown, the device for improving voice recognition accuracy includes a voice feature extraction unit 10, a human voice data acquisition unit 20, a voiceprint feature matching unit 30, and a voice command recognition and execution unit 40. The detailed description of each functional module is as follows:
[0107] The voice feature extraction unit 10 is used to obtain the voice data to be recognized and extract the voice features corresponding to the voice data to be recognized;
[0108] The human voice data acquisition unit 20 is used to input the voice features into a pre-trained voice classification model to obtain human voice data;
[0109] The voiceprint feature matching unit 30 is configured to match the human voice data with the voiceprint features of pre-registered users;
[0110] The voice command recognition and execution unit 40 is configured to, if the matching is successful, recognize the voice command corresponding to the human voice data and perform corresponding operations based on the voice command.
[0111] In an embodiment of the present application, the voice feature extraction unit 10 is further configured to:
[0112] Preprocess the voice data to be recognized;
[0113] Extract voice features from the preprocessed voice data to be recognized, where the voice features include any one or any combination of spectral features, Mel spectral features, rhythm features, and intonation features.
[0114] In an embodiment of the present application, the voice feature extraction unit 10 is further configured to:
[0115] Perform pre-emphasis processing on the voice data to be recognized;
[0116] Perform frame segmentation processing on the pre-emphasized voice data to be recognized to obtain multiple audio frames;
[0117] Perform windowing processing on each audio frame through a preset windowing algorithm.
[0118] In an embodiment of the present application, the voice command recognition and execution unit 40 is further configured to:
[0119] Perform security detection on the voice command;
[0120] If there is a risk in the voice command, obtain the user biometric features of the target user;
[0121] Match the user biometric features with the biometric features of the user whose voiceprint features match successfully;
[0122] If the matching is successful, perform corresponding operations based on the voice command.
[0123] In an embodiment of the present application, there are multiple pieces of the human voice data, and the voice feature extraction unit 10 is further configured to:
[0124] Recognize the voice command corresponding to each piece of human voice data;
[0125] Determine the permission priority level of the user corresponding to each piece of human voice data;
[0126] Execute the voice command based on the permission priority level.
[0127] In an embodiment of the present application, the voiceprint feature matching unit 30 is further configured to:
[0128] Extract the voiceprint feature corresponding to the human voice data;
[0129] Determine the similarity value between the extracted voiceprint feature and the voiceprint feature of the pre-registered user;
[0130] When the similarity value is greater than the preset similarity threshold, it indicates that the human voice data matches the voiceprint data of the pre-registered user successfully.
[0131] In an embodiment of the present application, the device further includes a keyword detection unit, which is configured to:
[0132] Convert the to-be-recognized voice data into text data;
[0133] Perform keyword detection on the text data;
[0134] If there is a keyword in the text data associated with a preset voice command, then execute the step of extracting the voice feature corresponding to the to-be-recognized voice data.
[0135] In an embodiment of the present application, by acquiring the to-be-recognized voice data and extracting its corresponding voice features, and inputting these features into a pre-trained voice classification model, it is possible to accurately distinguish human voice data from other interfering sounds (such as song sounds, navigation broadcast sounds, engine noise sounds, etc.), so as to obtain pure human voice data. The interference factors in the complex voice environment are effectively filtered, laying a foundation for accurately recognizing voice commands subsequently. And the obtained human voice data is matched with the voiceprint features of the pre-registered user. Only when the match is successful, will the voice command be further recognized and the operation be executed. Voiceprint matching, as an identity verification mechanism, can effectively exclude the voice interference of non-target users, and by utilizing the unique voiceprint features of the target user and combining with the voice recognition model, it can better adapt to the voice characteristics of the target user, improve the recognition accuracy of diverse commands of the target user, and further increase the hit rate of commands, providing a more stable and accurate voice interaction service for users.
[0136] For the specific limitations on the device for improving voice recognition accuracy, reference can be made to the limitations on the method for improving voice recognition accuracy in the foregoing text, which will not be elaborated here. Each module in the above device for improving voice recognition accuracy can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0137] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as shown in Figure 3 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a method for improving the accuracy of speech recognition is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0138] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the method for improving the accuracy of speech recognition as described above are implemented.
[0139] In an embodiment of the application, a readable storage medium is provided. The readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the method for improving the accuracy of speech recognition as described above are implemented.
[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0142] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for improving speech recognition accuracy, characterized in that: The method comprises: Acquire speech data to be recognized, and extract speech features corresponding to the speech data to be recognized; Inputting the speech features into a pre-trained speech classification model to obtain human voice data; Matching the human voice data with the voiceprint features of a pre-registered user; If the match is successful, the voice command corresponding to the human voice data is identified and the corresponding operation is performed based on the voice command.
2. The method for improving speech recognition accuracy according to claim 1, characterized in that: The extracting the speech features corresponding to the speech data to be recognized includes: Preprocessing the speech data to be recognized; Speech feature extraction is performed on the preprocessed speech data to be recognized, wherein the speech feature includes any one of spectrum feature, Mel spectrum feature, rhythm feature and intonation feature or any combination thereof.
3. The method for improving speech recognition accuracy as claimed in claim 2, characterized in that: The preprocessing of the to-be-recognized speech data comprises: Performing pre-emphasis processing on the speech data to be recognized; Performing frame processing on the speech data to be recognized after the pre-emphasis processing to obtain multiple audio frames; A windowing process is performed on each audio frame using a preset windowing algorithm.
4. The method for improving speech recognition accuracy according to claim 1, characterized in that: The identifying and obtaining a voice instruction corresponding to the human voice data, and performing a corresponding operation based on the voice instruction, includes: Performing a security check on the voice command; If the voice command is risky, obtaining biometric features of the target user; Matching the user's biometric features with the biometric features of the user whose voiceprint features are successfully matched; If the match is successful, the corresponding operation is performed based on the voice instruction.
5. The method for improving speech recognition accuracy according to claim 1, characterized in that: The human voice data includes a plurality of voice instructions corresponding to the human voice data, and the recognition obtains a voice instruction corresponding to the human voice data, and performs a corresponding operation based on the voice instruction: Recognize the voice command corresponding to each voice data; Determine the permission priority of each user corresponding to each voice data; Based on the permission priority, the voice command is executed.
6. The method for improving speech recognition accuracy according to claim 1, characterized in that: The matching of the human voice data with the voiceprint features of a pre-registered user includes: Extracting voiceprint features corresponding to the human voice data; Determining a similarity value between the extracted voiceprint feature and the voiceprint feature of the pre-registered user; When the similarity value is greater than a preset similarity threshold, it indicates that the human voice data matches the voiceprint data of the pre-registered user successfully.
7. The method for improving speech recognition accuracy according to claim 1, characterized in that: The step of extracting the speech features corresponding to the speech data to be recognized includes: Converting the to-be-recognized speech data into text data; Performing keyword detection on the text data; If the text data contains keywords associated with the preset voice command, the step of extracting voice features corresponding to the voice data to be recognized is performed.
8. A device for improving speech recognition accuracy, characterized in that: The device comprises: A speech feature extraction unit, used to obtain speech data to be recognized and extract speech features corresponding to the speech data to be recognized; A human voice data acquisition unit, used for inputting the voice features into a pre-trained voice classification model to obtain human voice data; A voiceprint feature matching unit, used to match the human voice data with the voiceprint features of a pre-registered user; The voice instruction recognition and execution unit is used to identify the voice instruction corresponding to the human voice data if the match is successful, and execute the corresponding operation based on the voice instruction.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executed on the processor, characterized in that: When the processor executes the computer-readable instructions, the method for improving speech recognition accuracy as described in any one of claims 1 to 7 is implemented.
10. A readable storage medium having computer readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the method for improving speech recognition accuracy as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent control method and system of electric appliance
CN122395248A