Voice rejection model generation method and device, voice rejection method and device and storage medium
By pre-training and fine-tuning the audio samples of multiple human-computer interactive devices, a voice refusal model suitable for specific devices is generated, which solves the problem of lack of universality in voice refusal in different device types, and improves training efficiency and model applicability.
Patent Information
- Application Number
- CN202311801042.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
Different types of human-computer interaction devices have differences in characteristics, noise and environmental conditions in speech recognition, resulting in the need of different data sets for model training, lack of universality, and increases the workload of model training and maintenance.
By obtaining the audio sample sets corresponding to multiple human-computer interactive devices, pre-training the voice refusal model, obtaining a second voice refusal model with strong versatility, and then fine-tuning the audio samples of the target device to generate a voice refusal model suitable for a specific device.
It improves the training efficiency of the speech-recognition model, reduces the workload of model training and maintenance, and avoids the problem that some device types are difficult to obtain sufficient sample data.
Smart Images

Figure CN120220655A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech recognition technology, and in particular, to a method for generating a speech rejection model, a speech rejection method, an apparatus, and a storage medium. Background Art
[0002] In order to improve the recognition ability of a human-machine interaction device for non-human-machine interaction instructions in continuous conversations, in related technologies, a speech rejection technology is usually used, that is, the audio data to be recognized is input into a trained speech rejection model, and then it is determined whether the audio data is used for human-machine interaction according to the result output by the speech rejection model.
[0003] However, different types of human-machine interaction devices (such as mobile phones, speakers, TVs, vehicles, etc.) may have different characteristics, noises, and environmental conditions in speech rejection. Therefore, different data sets need to be used for corresponding model training for different types of human-machine interaction devices, resulting in a lack of generality between different types of speech rejection models and increasing the workload of model training and maintenance. Summary of the Invention
[0004] To overcome the problems in the related technologies, the present disclosure provides a method for generating a speech rejection model, a speech rejection method, an apparatus, and a storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a method for generating a speech rejection model is provided, including:
[0006] Obtaining an audio sample set, where the audio sample set includes audio samples corresponding to multiple human-machine interaction devices, and the audio sample corresponding to the human-machine interaction device includes audio data and a label of the audio data, and the label represents whether the audio data is used for human-machine interaction;
[0007] Pre-training a first speech rejection model according to the audio sample set to obtain a second speech rejection model;
[0008] Fine-tuning and training the second speech rejection model according to the audio sample corresponding to the target human-machine interaction device to obtain a speech rejection model corresponding to the target human-machine interaction device.
[0009] Optionally, the pre-training includes extracting audio features of the audio samples in the audio sample set for single-modal pre-training; the fine-tuning and training includes extracting audio features and semantic features of the audio sample corresponding to the target human-machine interaction device to perform multi-modal fine-tuning and training on the second speech rejection model.
[0010] Optionally, the audio sample set further includes unsupervised training samples;
[0011] Pre-training the first voice rejection model according to the audio sample set includes:
[0012] Performing semi-supervised training on the first voice rejection model according to the audio samples corresponding to the multiple human-computer interaction devices and the unsupervised training samples.
[0013] Optionally, performing semi-supervised training on the first voice rejection model according to the audio samples corresponding to the multiple human-computer interaction devices and the unsupervised training samples includes:
[0014] Performing a first enhancement process and a second enhancement process on the unsupervised training samples to obtain a first unsupervised enhanced sample and a second unsupervised enhanced sample;
[0015] Adding pseudo-labels to the second unsupervised enhanced sample according to the model prediction results of the first unsupervised enhanced sample by the first voice rejection model;
[0016] Performing semi-supervised training on the first voice rejection model according to the audio samples corresponding to the multiple human-computer interaction devices and the second unsupervised enhanced sample with the pseudo-labels.
[0017] Optionally, performing fine-tuning training on the second voice rejection model according to the audio samples corresponding to the target human-computer interaction device includes:
[0018] Performing first fine-tuning training on the second voice rejection model according to the audio samples corresponding to the target human-computer interaction device to obtain a third voice rejection model and first fine-tuning training result data;
[0019] In the case where the first fine-tuning training result data does not meet the preset training effect, extracting target audio samples corresponding to the preset data type from the audio samples corresponding to the target human-computer interaction device;
[0020] Performing second fine-tuning training on the third voice rejection model according to the target audio samples.
[0021] Optionally, the second voice rejection model includes an audio feature extraction module, a semantic feature extraction module, and a classifier;
[0022] The fine-tuning training includes:
[0023] Performing fine-tuning training on the second voice rejection model by taking the audio samples corresponding to the target human-computer interaction device as the inputs of the audio feature extraction module and the semantic feature extraction module respectively, taking the audio features output by the audio feature extraction module and the semantic features output by the semantic feature extraction module as the inputs of the classifier, and taking the label as the output of the classifier.
[0024] Optionally, the multiple human-computer interaction devices include human-computer interaction devices of multiple different device types.
[0025] According to a second aspect of the embodiments of the present disclosure, there is provided a voice rejection recognition method, including:
[0026] Obtaining to-be-recognized audio data collected by a target human-computer interaction device;
[0027] Inputting the to-be-recognized audio data into a pre-trained voice rejection recognition model corresponding to the target human-computer interaction device to determine whether the to-be-recognized audio data is used for human-computer interaction, where the voice rejection recognition model is generated according to the voice rejection recognition model generation method provided in the first aspect of the present disclosure.
[0028] According to a third aspect of the embodiments of the present disclosure, there is provided a voice rejection recognition model generation device, including:
[0029] A first acquisition module, configured to acquire an audio sample set, where the audio sample set includes audio samples corresponding to multiple human-computer interaction devices, and the audio samples corresponding to the human-computer interaction devices include audio data and labels of the audio data, and the labels represent whether the audio data is used for human-computer interaction;
[0030] A pre-training module, configured to pre-train a first voice rejection recognition model according to the audio sample set to obtain a second voice rejection recognition model;
[0031] A model generation module, configured to fine-tune and train the second voice rejection recognition model according to the audio samples corresponding to the target human-computer interaction device to obtain a voice rejection recognition model corresponding to the target human-computer interaction device.
[0032] According to a fourth aspect of the embodiments of the present disclosure, there is provided a voice rejection recognition device, including:
[0033] A second acquisition module, configured to acquire to-be-recognized audio data collected by a target human-computer interaction device;
[0034] A voice rejection recognition module, configured to input the to-be-recognized audio data into a pre-trained voice rejection recognition model corresponding to the target human-computer interaction device to determine whether the to-be-recognized audio data is used for human-computer interaction, where the voice rejection recognition model is generated according to the voice rejection recognition model generation method provided in the first aspect of the present disclosure.
[0035] According to a fifth aspect of the embodiments of the present disclosure, there is provided a voice rejection recognition model generation device, including:
[0036] A processor;
[0037] A memory for storing processor-executable instructions;
[0038] Wherein, the processor is configured to:
[0039] Obtain an audio sample set, where the audio sample set includes audio samples corresponding to multiple human-machine interaction devices, and the audio sample corresponding to the human-machine interaction device includes audio data and a label of the audio data, and the label indicates whether the audio data is used for human-machine interaction;
[0040] Pre-train a first voice rejection recognition model according to the audio sample set to obtain a second voice rejection recognition model;
[0041] Fine-tune and train the second voice rejection recognition model according to the audio sample corresponding to the target human-machine interaction device to obtain a voice rejection recognition model corresponding to the target human-machine interaction device.
[0042] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the voice rejection recognition model generation method provided in the first aspect of the present disclosure or the steps of the voice rejection recognition method provided in the second aspect of the present disclosure are implemented.
[0043] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By pre-training the first voice rejection recognition model with an audio sample set including audio samples corresponding to multiple human-machine interaction devices, the obtained second voice rejection recognition model has strong versatility. Then, the second voice rejection recognition model is fine-tuned and trained using the audio sample corresponding to the target human-machine interaction device to obtain a voice rejection recognition model corresponding to the target human-machine interaction device. In this way, by fine-tuning the second voice model with versatility, the voice rejection recognition model corresponding to the target human-machine interaction device can be quickly obtained, improving the training efficiency of the voice rejection recognition model, reducing the workload of model training and maintenance while ensuring accuracy. Further, since pre-training is based on audio samples corresponding to multiple human-machine interaction devices, the number of samples is increased, avoiding the problem that it is difficult to obtain sufficient sample data for some specific device types of human-machine interaction devices.
[0044] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0045] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0046] Figure 1 It is a flowchart of a voice rejection recognition method in the related art.
[0047] Figure 2 It is a schematic diagram of an implementation environment shown according to an exemplary embodiment.
[0048] Figure 3 It is a flowchart of a method for generating a voice rejection recognition model shown according to an exemplary embodiment.
[0049] Figure 4 It is a flowchart of a method for obtaining training loss shown according to an exemplary embodiment.
[0050] Figure 5 It is a flowchart of a fine-tuning training method shown according to an exemplary embodiment.
[0051] Figure 6 It is a flowchart of another method for generating a voice rejection recognition model shown according to an exemplary embodiment.
[0052] Figure 7 It is a flowchart of a voice rejection recognition method shown according to an exemplary embodiment.
[0053] Figure 8 It is a block diagram of a device for generating a voice rejection recognition model shown according to an exemplary embodiment.
[0054] Figure 9 It is a block diagram of a voice rejection recognition device shown according to an exemplary embodiment.
[0055] Figure 10 It is a block diagram of a device for generating a voice rejection recognition model shown according to an exemplary embodiment.
[0056] Figure 11 It is a block diagram of another device for generating a voice rejection recognition model shown according to an exemplary embodiment. Detailed implementation manners
[0057] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0058] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and obtaining the authorization given by the owner of the corresponding device.
[0059] To improve the recognition ability of a human-computer interaction device for non-human-computer interaction instructions in continuous conversations, voice rejection recognition technology is usually used in related technologies. That is, the audio data to be recognized is input into a trained voice rejection recognition model, and then it is determined whether the audio data is for human-computer interaction according to the result output by the voice rejection recognition model. Figure 1 is a flowchart of a voice rejection recognition method in related technologies. As Figure 1 shown, the voice rejection recognition model processes the input audio data as follows: The audio data undergoes spectrum extraction operations such as framing, windowing, and fast Fourier transform to obtain spectrum features, and then audio features are obtained through a voice encoder; at the same time, an Automatic Speech Recognition (ASR) module recognizes the audio data as text; then, text features are obtained through a text encoder; and high-order features of Natural Language Processing (NLP) in the text are extracted through a Natural Language Understanding (NLU) module; finally, the above audio features, text features, and NLP high-order features are input into a classifier to obtain a voice rejection recognition result, that is, the probability that the audio data is non-human-computer interaction.
[0060] However, different types of human-computer interaction devices (such as mobile phones, speakers, TVs, vehicles, etc.) may have different characteristics, noises, and environmental conditions in voice rejection recognition. Therefore, different datasets need to be used for corresponding model training for different types of human-computer interaction devices. Although the final voice rejection recognition models have the same structure, due to being trained with different datasets separately, they lack generality, increasing the workload of model training and maintenance, and it is difficult to quickly obtain voice rejection recognition models for different types of human-computer interaction devices.
[0061] In addition, to ensure the accuracy of the voice rejection recognition model, the training process of the above voice rejection recognition model needs to rely on a large amount of labeled data. In this process, first, different types of human-computer interaction devices need to be distinguished, and the voice requests of each human-computer interaction device need to be collected and mined separately, and then it needs to be manually labeled whether the voice request is for human-computer interaction, with a large workload; moreover, in some application scenarios, it is difficult to obtain a large amount of labeled data. For example, in the scenario of in-vehicle voice assistants, since the vehicles of the models to be tested have not been launched yet and there is a lack of real user traffic, it takes a long time to accumulate a sufficient amount of labeled data, which cannot meet the R & D and iteration requirements of voice rejection recognition technology on the vehicle side. Although a certain amount of data can be obtained through means such as purchasing and dubbing, the acquisition cost of these data is high and the authenticity is low.
[0062] To this end, the present disclosure provides a method for generating a voice rejection recognition model, a voice rejection recognition method, an apparatus, and a storage medium. The specific embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0063] Figure 2 is a schematic diagram of an implementation environment shown according to an exemplary embodiment. As Figure 2 shown, the implementation environment is a human-machine dialogue scenario, which may include a terminal 01 and a server 02. Among them, the terminal 01 may be a device capable of human-machine interaction such as a smart phone, a speaker, a television, a vehicle, a tablet computer, a personal computer (PC), a notebook computer, etc. An application for collecting audio is installed on the terminal 01; the server 02 may be an application server corresponding to the above application. The terminal 01 and the server 02 are communicatively coupled and may communicate in any one of, for example, 3G, 4G, 5G, near-field communication, etc. In the actual application process, the terminal 01 collects audio signals such as the user's voice and ambient sound, and transmits the audio signal to the server 02. The voice rejection recognition model deployed in the server 02 performs non-human-machine interaction recognition on the audio signal, obtains a recognition result, and issues a rejection recognition instruction or a response instruction according to the recognition result.
[0064] Figure 3 is a flowchart of a method for generating a voice rejection recognition model shown according to an exemplary embodiment. As Figure 3 shown, the method for generating a voice rejection recognition model includes the following steps:
[0065] S101: Obtain an audio sample set, where the audio sample set includes audio samples corresponding to multiple human-machine interaction devices. The audio sample corresponding to a human-machine interaction device includes audio data and a label of the audio data, and the label indicates whether the audio data is used for human-machine interaction.
[0066] In an embodiment of the present disclosure, the audio sample corresponding to a human-machine interaction device refers to the audio data collected during the operation of the human-machine interaction device and the label of the audio data. Whether the audio data is used for human-machine interaction can be determined by whether the human-machine interaction device responds to the audio data. In another embodiment, it is also possible to manually listen to the audio data to determine whether the audio data is used for human-machine interaction and label it. In the voice rejection recognition scenario, usually, the audio data not used for human-machine interaction is labeled as 1, and the audio data used for human-machine interaction is labeled as 0.
[0067] Specifically, the audio data includes the user's voice commands, other voices of the user, ambient sound, etc. collected by the human-machine interaction device. Among them, the user's voice commands are audio data used for human-machine interaction, and other voices of the user, ambient sound, etc. are not audio data used for human-machine interaction.
[0068] S102: Pre-train the first voice rejection model according to the audio sample set to obtain the second voice rejection model.
[0069] In an embodiment of the present disclosure, the first voice rejection model can be a conventional voice rejection model structure based on a voice encoder, a text encoder, and an NLU module. This embodiment does not make specific limitations on this. The training objective of the first voice rejection model can be set according to the actual voice rejection task requirements or historical experience. Use the above audio sample set to pre-train the first voice rejection model to obtain the pre-trained voice rejection model, that is, the second voice rejection model.
[0070] S103: Fine-tune the second voice rejection model according to the audio samples corresponding to the target human-computer interaction device to obtain the voice rejection model corresponding to the target human-computer interaction device.
[0071] In an embodiment of the present disclosure, fine-tune the second voice rejection model according to the audio samples corresponding to the target human-computer interaction device in the audio sample set to fine-tune the model parameters and obtain the voice rejection model corresponding to the target human-computer interaction device. This voice rejection model can better adapt to the specific environment and voice characteristics of the target human-computer interaction device to ensure the accuracy of voice rejection.
[0072] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: Pre-train the first voice rejection model through an audio sample set containing audio samples corresponding to multiple human-computer interaction devices, so that the obtained second voice rejection model has strong generality. Then, fine-tune the second voice rejection model using the audio samples corresponding to the target human-computer interaction device to obtain the voice rejection model corresponding to the target human-computer interaction device. In this way, fine-tuning the second voice model with generality can quickly obtain the voice rejection model corresponding to the target human-computer interaction device, improving the training efficiency of the voice rejection model, reducing the workload of model training and maintenance while ensuring accuracy. Further, since pre-training is based on audio samples corresponding to multiple human-computer interaction devices, the sample quantity is increased, avoiding the problem that it is difficult to obtain sufficient sample data for some specific device types of human-computer interaction devices.
[0073] As an alternative implementation, the multiple human-computer interaction devices include multiple human-computer interaction devices of different device types, such as: smart phones, speakers, TVs, vehicles, etc. In another embodiment, the multiple human-computer interaction devices may also include multiple human-computer interaction devices of the same device type.
[0074] As an alternative implementation, the pre-training includes extracting audio features of audio samples in the audio sample set for unimodal pre-training; the fine-tuning training includes extracting audio features and semantic features of audio samples corresponding to the target human-computer interaction device for multimodal fine-tuning training of the second voice rejection recognition model.
[0075] In an embodiment of the present disclosure, the pre-training is a unimodal training process, and the fine-tuning training is a multimodal training process. The inventor found through research that although the usage scenarios of human-computer interaction devices of different device types are different, the voice commands for human-computer interaction by users usually have certain commonalities in audio features, such as the spectral distribution of sound, pitch change, etc. Therefore, extracting audio features of audio samples in the audio sample set for unimodal pre-training can enable the second voice rejection recognition model to perform general audio learning and thus have strong generality. Then, extracting audio features and semantic features of audio samples corresponding to the target human-computer interaction device for multimodal fine-tuning training of the second voice rejection recognition model enables the finally obtained voice rejection recognition model to accurately perform audio learning and semantic learning for the target human-computer interaction device to adapt to the data distribution and voice signal differences of the target human-computer interaction device.
[0076] In addition, the embodiments of the present disclosure can also be adaptively applied to the training scenarios of other multimodal models, where the multimodal models can be image recognition models, video understanding models, fusion models, etc.
[0077] As an alternative implementation, the audio sample set further includes unsupervised training samples; S102 may include: performing semi-supervised training on the first voice rejection recognition model according to audio samples corresponding to multiple human-computer interaction devices and the unsupervised training samples.
[0078] In an embodiment of the present disclosure, the pre-training process can be semi-supervised learning, that is, performing semi-supervised training on the first voice rejection recognition model according to audio samples corresponding to multiple human-computer interaction devices and the unsupervised training samples. In this way, by using unsupervised training samples to increase the data scale of training samples, the problem that it is difficult to obtain sufficient sample data for certain specific device types of human-computer interaction devices can be further avoided.
[0079] In another embodiment of the present disclosure, extracting audio features of audio samples in the audio sample set and audio features in the unsupervised training samples, and performing semi-supervised training on the first voice rejection recognition model to obtain the second voice rejection recognition model. In this way, a large amount of unsupervised audio data can be utilized to improve the learning and understanding ability of the second voice rejection recognition model for audio features, reduce the dependence on semantic features, and improve the generality and accuracy of the second voice rejection recognition model.
[0080] As another alternative implementation, data augmentation processing can be performed on the audio samples corresponding to multiple human-computer interaction devices to further increase the sample data volume; for example, the data augmentation processing can include audio mixing, audio splicing, audio speed change, SpecAugment (a data augmentation method for speech recognition data at the Mel spectrogram level), etc.
[0081] As an alternative implementation, according to the audio samples corresponding to multiple human-computer interaction devices and the unsupervised training samples, semi-supervised training is performed on the first voice rejection model, including: performing first augmentation processing and second augmentation processing on the unsupervised training samples to obtain first unsupervised augmented samples and second unsupervised augmented samples; adding pseudo-labels to the second unsupervised augmented samples according to the model prediction results of the first voice rejection model for the first unsupervised augmented samples; performing semi-supervised training on the first voice rejection model according to the audio samples corresponding to multiple human-computer interaction devices and the second unsupervised augmented samples with pseudo-labels.
[0082] Figure 4 It is a flowchart of a method for obtaining training loss shown according to an exemplary embodiment. As Figure 4 shown, first perform first augmentation processing on the unsupervised training samples, such as weak augmentation processing, to obtain first unsupervised augmented samples, and perform second augmentation processing on the unsupervised training samples, such as strong augmentation processing, to obtain second unsupervised augmented samples. Then, input the first unsupervised augmented samples and the second unsupervised augmented samples into the first voice rejection model respectively for model prediction to obtain corresponding model prediction results a and model prediction results b. Then add pseudo-labels to the second unsupervised augmented samples according to the model prediction results a of the first unsupervised augmented samples, where non-confident labels can be filtered when adding pseudo-labels. Then, calculate the cross-entropy according to the model prediction results b of the second unsupervised augmented samples and the pseudo-labels of the second unsupervised augmented samples to obtain the unsupervised loss; then, input the supervised training samples (i.e., the audio samples corresponding to multiple human-computer interaction devices) into the first voice rejection model to calculate the supervised loss, and use the sum of the unsupervised loss and the supervised loss as the training loss of the first voice rejection model.
[0083] As an alternative implementation, S103 can include: performing first fine-tuning training on the second voice rejection model according to the audio samples corresponding to the target human-computer interaction device to obtain a third voice rejection model and first fine-tuning training result data; in the case where the first fine-tuning training result data does not meet the preset training effect, extracting target audio samples corresponding to the preset data type from the audio samples corresponding to the target human-computer interaction device; performing second fine-tuning training on the third voice rejection model according to the target audio samples.
[0084] In an embodiment of the present disclosure, Figure 5It is a flowchart of a fine-tuning training method shown according to an exemplary embodiment. As Figure 5 shown, it can be achieved through the following steps:
[0085] S1031: Perform first fine-tuning training on the second voice rejection model according to the audio samples corresponding to the target human-computer interaction device.
[0086] Among them, the first fine-tuning training can be device adaptive adjustment, so that the differential features existing between the target human-computer interaction device and other human-computer interaction devices can be learned, for example, ASR confidence, vertical domain, etc., to adapt to the data distribution and voice signal differences of the target human-computer interaction device.
[0087] S1032: Determine whether the first fine-tuning training result data obtained by the first fine-tuning training meets the preset training effect.
[0088] Among them, the preset training effect can be set according to the actual voice rejection task requirements or historical experience; in the case where the first fine-tuning training result data does not meet the preset training effect, execute S1033 below; in the case where the first fine-tuning training result data meets the preset training effect, execute S1035 below.
[0089] S1033: Extract the target audio samples corresponding to the preset data type from the audio samples corresponding to the target human-computer interaction device.
[0090] Among them, the preset data type can be the data type corresponding to the core vertical domain of the target human-computer interaction device. For example, when the target human-computer interaction device is a vehicle, its core vertical domain is navigation; when the target human-computer interaction device is a TV, its core vertical domain is video; when the target human-computer interaction device is a speaker, its core vertical domain is voice.
[0091] S1034: Perform second fine-tuning training on the third voice rejection model obtained by the first fine-tuning training according to the target audio samples to obtain the voice rejection model corresponding to the target human-computer interaction device.
[0092] Among them, the second fine-tuning training can be device fine-tuning, so that the effect of the voice rejection model on the preset data type can be improved, and the voice rejection model can better adapt to the usage scenario corresponding to the target human-computer interaction device.
[0093] S1035: Use the third voice rejection model obtained by the first fine-tuning training as the voice rejection model corresponding to the target human-computer interaction device.
[0094] As an alternative implementation, the voice rejection recognition model includes an audio feature extraction module, a semantic feature extraction module, and a classifier; the above fine-tuning training may include: by using the audio samples as the inputs of the audio feature extraction module and the semantic feature extraction module respectively, taking the audio features output by the audio feature extraction module and the semantic features output by the semantic feature extraction module as the inputs of the classifier, and taking the labels as the outputs of the classifier, to perform fine-tuning training on the voice rejection recognition model.
[0095] Exemplarily, in the audio feature extraction module, the input audio samples can be framed, windowed, and subjected to fast Fourier transform to obtain spectral features (with a dimension of 500*128), and the spectral features are input into a voice encoder, which can be an encoder based on Conformer (a voice recognition model based on a convolutional enhanced encoder-decoder architecture), and a voice coding vector (with a dimension of 768) is output, that is, the audio features;
[0096] In the semantic feature extraction module, the input audio samples can be recognized as text (with a dimension of 20 after word segmentation) through an ASR module, and then the text is input into a text encoder, which can be an encoder based on word2vec+CNN (a word vector conversion model), and a text coding vector (with a dimension of 256) is output, that is, the text features; at the same time, NLP high-order features (with a dimension of 34) in the above text are extracted through an NLU module; the text features and the NLP high-order features are output as semantic features;
[0097] Finally, the audio features and the semantic features are concatenated, and the concatenation result (with a dimension of 1058) is input into the classifier, which can be a fully connected layer, and a label (with a dimension of 2) is output. When the label is 1, it indicates that the audio data in the above audio sample is a non-human-computer interaction instruction, and when the label is 0, it indicates that the audio data in the above audio sample is a human-computer interaction instruction.
[0098] Figure 6 It is a flowchart of another method for generating a voice rejection recognition model shown according to an exemplary embodiment. As Figure 6As shown, first, obtain the audio samples and unsupervised training samples corresponding to the human-computer interaction devices of multiple device types. Among them, the human-computer interaction devices of multiple device types can be speakers, TVs, and vehicles. Then, input the above audio samples and unsupervised training samples into the voice rejection recognition model for single-modal pre-training of the voice rejection recognition model. Specifically, the audio data in the above audio samples and unsupervised training samples are subjected to spectrum extraction operations such as framing, windowing, and fast Fourier transform to obtain spectrum features, and then audio features are obtained through a voice encoder. Then, input the audio samples corresponding to the vehicle into the pre-trained voice rejection recognition model for the first fine-tuning training of the voice rejection recognition model. Specifically, the audio data in the audio samples corresponding to the vehicle are subjected to spectrum extraction operations such as framing, windowing, and fast Fourier transform to obtain spectrum features, and then audio features are obtained through a voice encoder; at the same time, the ASR module recognizes the audio data as text; then, text features are obtained through a text encoder; and NLP high-order features in the text are extracted through the NLU module; finally, the above audio features, text features, and NLP high-order features are input into a classifier to obtain a voice rejection recognition result, where the voice rejection recognition result does not meet the preset training effect. Finally, filter out the audio samples related to navigation from the audio samples corresponding to the vehicle, and input the audio samples into the voice rejection recognition model after the first fine-tuning training for the second fine-tuning training of the voice rejection recognition model to obtain the voice rejection recognition model corresponding to the vehicle. Specifically, the process of the second fine-tuning training is the same as that of the above first fine-tuning training process.
[0099] To more clearly show that the improved voice rejection recognition model in this embodiment can accurately identify non-human-recognized audio data, the present application also conducts an experimental comparison between the target voice rejection recognition models corresponding to the human-computer interaction devices of multiple device types and the existing voice rejection recognition models in the related art, as follows:
[0100] In this experimental comparison, target voice rejection recognition models corresponding to vehicles, speakers, and TVs are generated respectively using the method of the present application. Among them, the data volumes of the training sets and test sets of each target voice rejection recognition model are as follows: the training set of the target voice rejection recognition model corresponding to the vehicle includes 610,000 supervised sample data and 500,000 unsupervised sample data, and its test set includes 11,700 test data; the training set of the target voice rejection recognition model corresponding to the speaker includes 260,000 supervised sample data and 1,090,000 unsupervised sample data, and its test set includes 14,200 test data; the training set of the target voice rejection recognition model corresponding to the TV includes 190,000 supervised sample data and 590,000 unsupervised sample data, and its test set includes 7,900 test data.
[0101] The comparison of the effects of the target voice rejection recognition models corresponding to the human-computer interaction devices of multiple device types and the existing voice rejection recognition models in the related art is shown in the following table:
[0102] According to the above effect comparison, compared with the existing voice rejection recognition model in the related art, the target voice rejection recognition model obtained by the method of this application can maintain the accuracy of voice rejection recognition, and the recall rate has been improved for the target voice rejection recognition models corresponding to human-computer interaction devices of different device types. It is easy to know that the above target voice rejection recognition model can identify more correct positive examples, that is, audio data that is not a human-computer interaction instruction, can reduce the risk of missed judgment, and has a good voice rejection recognition effect.
[0103] Figure 7 is a flowchart of a voice rejection recognition method shown according to an exemplary embodiment. As Figure 7 shown, the voice rejection recognition method includes the following steps:
[0104] S201: Obtain the audio data to be recognized collected by the target human-computer interaction device.
[0105] In an embodiment of the present disclosure, the audio data to be recognized collected by the target human-computer interaction device is obtained. Exemplarily, the target human-computer interaction device may be a smart phone, a speaker, a television, a vehicle, etc.
[0106] S202: Input the audio data to be recognized into the voice rejection recognition model corresponding to the target human-computer interaction device that has been pre-trained, and determine whether the audio data to be recognized is used for human-computer interaction.
[0107] Among them, the voice rejection recognition model is generated according to the above voice rejection recognition model generation method.
[0108] Figure 8 is a block diagram of a voice rejection recognition model generation device shown according to an exemplary embodiment. Referring to Figure 8 , the voice rejection recognition model generation device 300 may include a first acquisition module 301, a pre-training module 302, and a model generation module 303.
[0109] The first acquisition module 301 is configured to acquire an audio sample set, the audio sample set includes audio samples corresponding to multiple human-computer interaction devices, the audio sample corresponding to the human-computer interaction device includes audio data and a label of the audio data, and the label represents whether the audio data is used for human-computer interaction;
[0110] The pre-training module 302 is configured to pre-train the first voice rejection recognition model according to the audio sample set to obtain a second voice rejection recognition model;
[0111] The model generation module 303 is configured to fine-tune and train the second voice rejection recognition model according to the audio sample corresponding to the target human-computer interaction device to obtain a voice rejection recognition model corresponding to the target human-computer interaction device.
[0112] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By pre-training the first voice rejection model with an audio sample set including audio samples corresponding to multiple human-computer interaction devices, the obtained second voice rejection model has strong versatility. Then, fine-tuning the second voice rejection model with the audio samples corresponding to the target human-computer interaction device to obtain a voice rejection model corresponding to the target human-computer interaction device. In this way, fine-tuning the second voice model with versatility can quickly obtain a voice rejection model corresponding to the target human-computer interaction device, improving the training efficiency of the voice rejection model, reducing the workload of model training and maintenance while ensuring accuracy. Further, since pre-training is based on audio samples corresponding to multiple human-computer interaction devices, the sample quantity is increased, avoiding the problem that it is difficult to obtain sufficient sample data for some specific types of human-computer interaction devices.
[0113] Optionally, the pre-training includes extracting audio features of the audio samples in the audio sample set for single-modal pre-training; the fine-tuning training includes extracting audio features and semantic features of the audio samples corresponding to the target human-computer interaction device for multi-modal fine-tuning training of the second voice rejection model.
[0114] Optionally, the audio sample set further includes unsupervised training samples;
[0115] The pre-training module 302 includes:
[0116] A first semi-supervised training sub-module, configured to perform semi-supervised training on the first voice rejection model according to the audio samples corresponding to the multiple human-computer interaction devices and the unsupervised training samples.
[0117] Optionally, the first semi-supervised training sub-module includes:
[0118] An enhancement sub-module, configured to perform a first enhancement process and a second enhancement process on the unsupervised training samples to obtain a first unsupervised enhanced sample and a second unsupervised enhanced sample;
[0119] A labeling sub-module, configured to add pseudo-labels to the second unsupervised enhanced sample according to the model prediction result of the first unsupervised enhanced sample by the first voice rejection model;
[0120] A second semi-supervised training sub-module, configured to perform semi-supervised training on the first voice rejection model according to the audio samples corresponding to the multiple human-computer interaction devices and the second unsupervised enhanced sample with the pseudo-labels.
[0121] Optionally, the model generation module 303 includes:
[0122] A first fine-tuning training submodule is configured to perform a first fine-tuning training on the second speech rejection model according to the audio sample corresponding to the target human-computer interaction device, to obtain a third speech rejection model and first fine-tuning training result data;
[0123] A sample extraction submodule is configured to extract a target audio sample corresponding to a preset data type from the audio sample corresponding to the target human-computer interaction device when the first fine-tuning training result data does not meet the preset training effect;
[0124] The second fine-tuning training submodule is configured to perform second fine-tuning training on the third speech rejection model according to the target audio sample.
[0125] Optionally, the second speech rejection model includes an audio feature extraction module, a semantic feature extraction module and a classifier;
[0126] The fine-tuning training comprises:
[0127] The second speech rejection model is fine-tuned and trained by using the audio samples corresponding to the target human-computer interaction device as the input of the audio feature extraction module and the semantic feature extraction module, respectively, using the audio features output by the audio feature extraction module and the semantic features output by the semantic feature extraction module as the input of the classifier, and using the label as the output of the classifier.
[0128] Optionally, the multiple human-computer interaction devices include multiple human-computer interaction devices of different device types.
[0129] Regarding the speech rejection model generating device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0130] Figure 9 is a block diagram of a voice rejection device according to an exemplary embodiment. Figure 9 The voice rejection device 400 may include a second acquisition module 401 and a voice rejection module 402.
[0131] The second acquisition module 401 is configured to acquire the audio data to be recognized collected by the target human-computer interaction device;
[0132] The speech rejection module 402 is configured to input the audio data to be recognized into a pre-trained speech rejection model corresponding to the target human-computer interaction device to determine whether the audio data to be recognized is used for human-computer interaction, wherein the speech rejection model is generated according to the above-mentioned speech rejection model generation method provided in the present disclosure.
[0133] Regarding the voice rejection recognition device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0134] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the above-mentioned voice rejection recognition model generation method or the steps of the above-mentioned voice rejection recognition method provided by the present disclosure are implemented.
[0135] Figure 10 FIG. 7 is a block diagram of a device 800 for generating a voice rejection recognition model according to an exemplary embodiment. For example, the device 800 may be a smart phone, a speaker, a television, a vehicle, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0136] Referring to Figure 10 , the device 800 may include one or more of the following components: a first processing component 802, a first memory 804, a first power supply component 806, a multimedia component 808, an audio component 810, a first input / output interface 812, a sensor component 814, and a communication component 816.
[0137] The first processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The first processing component 802 may include one or more first processors 820 to execute instructions to complete all or part of the steps of the above-mentioned voice rejection recognition model generation method. In addition, the first processing component 802 may include one or more modules to facilitate the interaction between the first processing component 802 and other components. For example, the first processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the first processing component 802.
[0138] The first memory 804 is configured to store various types of data to support the operation of the device 800. Examples of these data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The first memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0139] The first power supply component 806 provides power for various components of the device 800. The first power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0140] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0141] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the first memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0142] The first input / output interface 812 provides an interface between the first processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0143] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0144] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0145] In an exemplary embodiment, the device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-mentioned voice rejection recognition model generation method.
[0146] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 804 including instructions, and the above instructions can be executed by a first processor 820 of the device 800 to complete the steps of the above-mentioned voice rejection recognition model generation method or the steps of the above-mentioned voice rejection method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0147] In another exemplary embodiment, there is also provided a computer program product that includes a computer program executable by a programmable device, the computer program having code portions for performing the above-described voice rejection recognition model generation method or the above-described voice rejection recognition method when executed by the programmable device.
[0148] Figure 11 FIG. 4 is a block diagram of another apparatus 1900 for voice rejection recognition model generation according to an exemplary embodiment. For example, apparatus 1900 may be provided as a server. Referring to Figure 11 FIG. 4, apparatus 1900 includes a second processing component 1922, which further includes one or more processors, and memory resources represented by a second memory 1932 for storing instructions executable by the second processing component 1922, such as application programs. The application programs stored in the second memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the second processing component 1922 is configured to execute instructions to perform the above-described voice rejection recognition model generation method.
[0149] Apparatus 1900 may also include a second power component 1926 configured to perform power management of apparatus 1900, a wired or wireless network interface 1950 configured to connect apparatus 1900 to a network, and a second input / output interface 1958. Apparatus 1900 may operate based on an operating system stored in memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0150] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0151] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for generating a voice rejection recognition model, characterized in that, Including: Obtain an audio sample set, where the audio sample set includes audio samples corresponding to multiple human-computer interaction devices. The audio samples corresponding to the human-computer interaction devices include audio data and labels of the audio data, and the labels represent whether the audio data is for human-computer interaction; Pre-train a first voice rejection recognition model according to the audio sample set to obtain a second voice rejection recognition model; Fine-tune and train the second voice rejection recognition model according to the audio samples corresponding to the target human-computer interaction device to obtain a voice rejection recognition model corresponding to the target human-computer interaction device.
2. The method according to claim 1, wherein The pre-training includes extracting audio features of the audio samples in the audio sample set for single-modal pre-training; the fine-tuning training includes extracting audio features and semantic features of the audio samples corresponding to the target human-computer interaction device to perform multi-modal fine-tuning training on the second voice rejection recognition model.
3. The method according to claim 1, characterized in that, The audio sample set further includes unsupervised training samples; The pre-training the first voice rejection recognition model according to the audio sample set includes: Perform semi-supervised training on the first voice rejection recognition model according to the audio samples corresponding to the multiple human-computer interaction devices and the unsupervised training samples.
4. The method according to claim 3, wherein The performing semi-supervised training on the first voice rejection recognition model according to the audio samples corresponding to the multiple human-computer interaction devices and the unsupervised training samples includes: Perform a first enhancement process and a second enhancement process on the unsupervised training samples to obtain a first unsupervised enhanced sample and a second unsupervised enhanced sample; Add pseudo-labels to the second unsupervised enhanced sample according to the model prediction result of the first unsupervised enhanced sample by the first voice rejection recognition model; Perform semi-supervised training on the first voice rejection recognition model according to the audio samples corresponding to the multiple human-computer interaction devices and the second unsupervised enhanced sample with the pseudo-labels.
5. The method according to claim 1, wherein The fine-tuning and training the second voice rejection recognition model according to the audio samples corresponding to the target human-computer interaction device includes: Perform a first fine-tuning training on the second voice rejection recognition model according to the audio samples corresponding to the target human-computer interaction device to obtain a third voice rejection recognition model and first fine-tuning training result data; In the case that the first fine-tuning training result data does not meet the preset training effect, extract target audio samples corresponding to a preset data type from the audio samples corresponding to the target human-computer interaction device; Perform a second fine-tuning training on the third voice rejection recognition model according to the target audio samples.
6. The method according to claim 2, characterized in that The second voice rejection recognition model includes an audio feature extraction module, a semantic feature extraction module, and a classifier; The fine-tuning training includes: Fine-tune and train the second voice rejection recognition model by using the audio samples corresponding to the target human-computer interaction device as the inputs of the audio feature extraction module and the semantic feature extraction module respectively, using the audio features output by the audio feature extraction module and the semantic features output by the semantic feature extraction module as the inputs of the classifier, and using the labels as the outputs of the classifier.
7. The method according to any one of claims 1-6, characterized in that, The multiple human-computer interaction devices include human-computer interaction devices of multiple different device types.
8. A voice rejection recognition method, characterized in that, Including: Obtain the audio data to be recognized collected by the target human-computer interaction device; Input the audio data to be recognized into the pre-trained voice rejection model corresponding to the target human-computer interaction device to determine whether the audio data to be recognized is used for human-computer interaction, where the voice rejection model is generated according to the voice rejection model generation method described in any one of claims 1-7.
9. A voice rejection recognition model generation device, characterized in that, Comprising: A first acquisition module configured to acquire an audio sample set, the audio sample set including audio samples corresponding to multiple human-computer interaction devices, the audio sample corresponding to the human-computer interaction device including audio data and a label of the audio data, the label characterizing whether the audio data is used for human-computer interaction; A pre-training module configured to pre-train a first voice rejection model according to the audio sample set to obtain a second voice rejection model; A model generation module configured to fine-tune and train the second voice rejection model according to the audio sample corresponding to the target human-computer interaction device to obtain a voice rejection model corresponding to the target human-computer interaction device.
10. A voice rejection recognition device, characterized in that, Comprising: A second acquisition module configured to acquire the audio data to be recognized collected by the target human-computer interaction device; A voice rejection module configured to input the audio data to be recognized into the pre-trained voice rejection model corresponding to the target human-computer interaction device to determine whether the audio data to be recognized is used for human-computer interaction, where the voice rejection model is generated according to the voice rejection model generation method described in any one of claims 1-7.
11. A voice rejection recognition model generation device, characterized in that Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: Acquire an audio sample set, the audio sample set including audio samples corresponding to multiple human-computer interaction devices, the audio sample corresponding to the human-computer interaction device including audio data and a label of the audio data, the label characterizing whether the audio data is used for human-computer interaction; Pre-train a first voice rejection model according to the audio sample set to obtain a second voice rejection model; Fine-tune and train the second voice rejection model according to the audio sample corresponding to the target human-computer interaction device to obtain a voice rejection model corresponding to the target human-computer interaction device.
12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method described in any one of claims 1-8 are implemented.