A model training method, wake-up method, device, and storage medium

By outputting features at the middle layer of the speech recognition model assist in training the speech wake-up model, and combining joint training between the cloud and the terminal, the problems of poor effect and low wake-up rate in the prior art are solved, achieving higher wake-up accuracy and efficiency.

CN117594046BActive Publication Date: 2025-07-04MOORE THREADS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311360531.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-19
Publication Date
2025-07-04
Estimated Expiration
2043-10-19

AI Technical Summary

Technical Problem

The existing methods of awakening on the terminal and awakening on the cloud have poor results in the voice wake-up model, with low wake-up rate, and failed to effectively use the wake-up model on the cloud to assist in improving the effect of the wake-up model on the terminal.

Method used

When the wake-up words are included in the sample data, the features output from the intermediate layer of the speech recognition model are input to the first speech wake-up model for training, and the parameters are adjusted in combination with the distance loss function, and the speech wake-up model of the cloud and terminal are jointly trained to optimize the parameters of the wake-up model.

Benefits of technology

Improves the wake-up rate and accuracy of the wake-up result of the voice wake-up model, reducing wake-up latency performance and computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117594046B_ABST
    Figure CN117594046B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, a wake-up method, an apparatus, and a storage medium. Among them, the model training method may include: obtaining sample data; wherein the sample data includes voice data and text data corresponding to the voice data; inputting the sample data into a speech recognition model to train the speech recognition model to obtain a first feature output by an intermediate layer of the speech recognition model; when the sample data contains a wake-up word, inputting the first feature into a first voice wake-up model to train the first voice wake-up model. Through the model training method provided by the present disclosure, the effect of the first voice wake-up model can be improved, the wake-up rate of the first voice wake-up model can be increased, and the wake-up result of the first voice wake-up model can be made more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech technology, and in particular, to a model training method, a wake-up method, a device, and a storage medium. Background Art

[0002] In the speech interaction scenario of intelligent hardware, whether the terminal device can be accurately woken up by voice (Keyword Spotting, KWS) is a crucial step. Currently, a method of combining on-device wake-up and cloud wake-up is often used to wake up the terminal device by voice. A lightweight and low-power speech wake-up model can be deployed on the device, and then a larger and higher-power but better-recognizing speech wake-up model can be deployed on the cloud.

[0003] Since the speech wake-up model on the cloud has high power consumption, it can be default not to be turned on; since the speech wake-up model on the device has low power consumption, it can always be turned on. If the on-device speech wake-up model finds that the current probability of being woken up is greater than a certain threshold (for example, 0.9), it can be directly woken up; at the same time, the on-device speech wake-up model can maintain another threshold (for example, 0.8). If it is found that the probability of a wake-up request is between the two thresholds (that is, between 0.8 and 0.9), the wake-up request can be sent to the speech wake-up model on the cloud for secondary wake-up.

[0004] However, the existing method of combining on-device wake-up and cloud wake-up still has problems of poor speech wake-up model effect and low wake-up rate. Summary of the Invention

[0005] In view of this, the present disclosure provides a model training method, a wake-up method, a device, and a storage medium, which can improve the effect of the speech wake-up model, increase the wake-up rate of the speech wake-up model, and make the wake-up result of the speech wake-up model more accurate.

[0006] According to one aspect of the present disclosure, a model training method is provided. The method includes: obtaining sample data; where the sample data includes speech data and text data corresponding to the speech data; inputting the sample data into a speech recognition model to train the speech recognition model to obtain a first feature output by an intermediate layer of the speech recognition model; and when the sample data contains a wake-up word, inputting the first feature into a first speech wake-up model to train the first speech wake-up model.

[0007] In a possible implementation, the first speech wake-up model is located in the cloud, and the method further includes: when the sample data contains a wake-up word, training a second speech wake-up model according to the sample data and the first speech wake-up model, where the second speech wake-up model is located in the terminal.

[0008] In a possible implementation, when the wake-up word is included in the sample data, training the second voice wake-up model according to the sample data and the first voice wake-up model includes: inputting the sample data into the second voice wake-up model to obtain a second probability distribution output by the second voice wake-up model; adjusting the parameters of the second voice wake-up model according to a distance loss function so that the second probability distribution output by the second voice wake-up model with adjusted parameters approaches the first probability distribution output by the first voice wake-up model; where the distance loss function is used to determine the loss of the second probability distribution relative to the first probability distribution.

[0009] In a possible implementation, the training process of the speech recognition model, the first voice wake-up model, and the second voice wake-up model includes: when the wake-up word is included in the sample data, determining a total loss function according to a first loss function, a second loss function, a third loss function, and the distance loss function; updating the speech recognition model, the first voice wake-up model, and the second voice wake-up model according to the total loss function; where the first loss function is the loss function of the speech recognition model; the second loss function is the loss function of the first voice wake-up model; and the third loss function is the loss function of the second voice wake-up model.

[0010] In a possible implementation, when the wake-up word is included in the sample data, determining the total loss function according to the first loss function, the second loss function, the third loss function, and the distance loss function includes: performing a weighted sum of the first loss function, the second loss function, the third loss function, and the distance loss function to obtain the total loss function.

[0011] According to another aspect of the present disclosure, a wake-up method is provided, the method including: receiving an input voice; when the terminal is unable to determine whether to perform wake-up based on the voice, inputting the voice into a speech recognition model in the cloud to obtain a second feature output by an intermediate layer of the speech recognition model; and inputting the second feature into a first voice wake-up model in the cloud to determine whether to perform wake-up.

[0012] In a possible implementation, before inputting the speech to a speech recognition model in the cloud to obtain a second feature output by an intermediate layer of the speech recognition model when the terminal cannot determine whether to perform wake-up based on the speech, the method further includes: inputting the speech to a second speech wake-up model of the terminal to obtain a first wake-up probability; if the first wake-up probability is greater than a first threshold, waking up the terminal; if the first wake-up probability is less than a second threshold, not waking up the terminal; if the first wake-up probability is between the first threshold and the second threshold, determining that the terminal cannot determine whether to perform wake-up based on the speech; wherein, the first threshold is greater than the second threshold.

[0013] In a possible implementation, when it is determined to perform wake-up after inputting the second feature to a first speech wake-up model in the cloud, the method further includes: inputting the second feature to a module after the intermediate layer of the speech recognition model to perform speech recognition.

[0014] According to another aspect of the present disclosure, there is provided a model training apparatus, the apparatus includes: an acquisition module, configured to acquire sample data; wherein, the sample data includes speech data and text data corresponding to the speech data; a speech recognition model training module, configured to input the sample data to a speech recognition model to train the speech recognition model to obtain a first feature output by an intermediate layer of the speech recognition model; a first speech wake-up model training module, configured to input the first feature to a first speech wake-up model to train the first speech wake-up model when the sample data includes a wake-up word.

[0015] In a possible implementation, the first speech wake-up model is located in the cloud, and the apparatus further includes: a second speech wake-up model training module, configured to train a second speech wake-up model based on the sample data and the first speech wake-up model when the sample data includes a wake-up word, and the second speech wake-up model is located in the terminal.

[0016] In a possible implementation, the second speech wake-up model training module is further configured to: input the sample data to the second speech wake-up model to obtain a second probability distribution output by the second speech wake-up model; adjust parameters of the second speech wake-up model according to a distance loss function to make the second probability distribution output by the second speech wake-up model with adjusted parameters approach a first probability distribution output by the first speech wake-up model; wherein, the distance loss function is used to determine a loss of the second probability distribution relative to the first probability distribution.

[0017] In a possible implementation, the device further includes: a total loss function determination module, configured to determine a total loss function according to a first loss function, a second loss function, a third loss function, and the distance loss function when the sample data contains a wake-up word; the speech recognition model training module is further configured to update the speech recognition model according to the total loss function; the first voice wake-up model training module is further configured to update the first voice wake-up model according to the total loss function; the second voice wake-up model training module is further configured to update the second voice wake-up model according to the total loss function; wherein, the first loss function is the loss function of the speech recognition model; the second loss function is the loss function of the first voice wake-up model; the third loss function is the loss function of the second voice wake-up model.

[0018] In a possible implementation, the total loss function determination module is further configured to: perform a weighted sum on the first loss function, the second loss function, the third loss function, and the distance loss function to obtain a total loss function.

[0019] According to another aspect of the present disclosure, a wake-up device is provided, the device includes: a receiving module, configured to receive an input voice; a second feature acquisition module, configured to, when the terminal cannot determine whether to perform wake-up according to the voice, input the voice into a speech recognition model in the cloud to obtain a second feature output by an intermediate layer of the speech recognition model; a first wake-up module, configured to input the second feature into a first voice wake-up model in the cloud to determine whether to perform wake-up.

[0020] In a possible implementation, the device further includes a second wake-up module, configured to: before the second feature acquisition module inputs the voice into a speech recognition model in the cloud to obtain a second feature output by an intermediate layer of the speech recognition model when the terminal cannot determine whether to perform wake-up according to the voice, input the voice into a second voice wake-up model of the terminal to obtain a first wake-up probability; if the first wake-up probability is greater than a first threshold, wake up the terminal; if the first wake-up probability is less than a second threshold, do not wake up the terminal; if the first wake-up probability is between the first threshold and the second threshold, determine that the terminal cannot determine whether to perform wake-up according to the voice; wherein, the first threshold is greater than the second threshold.

[0021] In a possible implementation, the device further includes a speech recognition module, configured to: when the first wake-up module inputs the second feature into a first voice wake-up model in the cloud and determines to perform wake-up, input the second feature into a module after an intermediate layer of the speech recognition model to perform speech recognition.

[0022] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above-mentioned model training method or wake-up method when executing the instructions stored in the memory.

[0023] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, the computer program instructions implement the above-mentioned model training method or wake-up method when executed by a processor.

[0024] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned model training method or wake-up method.

[0025] In the model training method provided by the present disclosure, when the sample data contains a wake-up word, the first feature output by the intermediate layer of the speech recognition model is input into the first voice wake-up model to train the first voice wake-up model; since the feature extraction ability of the speech recognition model is stronger than that of the first voice wake-up model, and the features extracted by the speech recognition model can contain semantic information, using the features output by the intermediate layer of the speech recognition model to train the first voice wake-up model can improve the effect of the first voice wake-up model, increase the wake-up rate of the first voice wake-up model, and make the wake-up result of the first voice wake-up model more accurate.

[0026] Other features and aspects of the present disclosure will become clear according to the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings included in and constituting a part of this specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0028] Figure 1 A flowchart showing a model training method according to an embodiment of the present disclosure.

[0029] Figure 2 A schematic diagram showing the training process of a speech recognition model according to an embodiment of the present disclosure.

[0030] Figure 3 A schematic diagram showing the training process of a first voice wake-up model according to an embodiment of the present disclosure.

[0031] Figure 4A flowchart showing a model training method according to an embodiment of the present disclosure.

[0032] Figure 5 A schematic diagram showing joint training of a speech recognition model, a first speech wake-up model, and a second speech wake-up model according to an embodiment of the present disclosure.

[0033] Figure 6 A flowchart showing a wake-up method according to an embodiment of the present disclosure.

[0034] Figure 7 A schematic diagram showing wake-up of a terminal according to an embodiment of the present disclosure.

[0035] Figure 8 A schematic structural diagram showing a model training device according to an embodiment of the present disclosure.

[0036] Figure 9 A schematic structural diagram showing a wake-up device according to an embodiment of the present disclosure.

[0037] Figure 10 A schematic structural diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Description of Specific Embodiments

[0038] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0039] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0040] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0041] Speech is the main medium of communication in daily life and also plays a very important role in speech interaction. Speech interaction is the process of one or more information transmissions through natural language communication between humans and machines. For example, when we book a flight through the intelligent assistant Siri or play a song through a smart speaker, it is a process of speech interaction.

[0042] In the voice interaction scenario of smart hardware, a normal process is as follows: The user can wake up the terminal device (such as a mobile phone) through a specific wake-up word (such as "Hi Siri"). The terminal device can turn on the microphone to record the user's speech in real time and send the audio to the Automatic Speech Recognition (ASR) system on the terminal or in the cloud for recognition. The recognition result can be sent to downstream tasks for subsequent processing, such as natural language understanding, etc. Finally, the terminal device can play the information to be fed back through speech synthesis. Among them, whether the terminal device can be accurately voice-woken up is the first and crucial step in the voice interaction link.

[0043] In the related art, a method of combining on-device wake-up and cloud wake-up is often used to perform voice wake-up on the terminal device. In a related art, the terminal may include a first wake-up model for performing a first wake-up on the wake-up speech input to the first wake-up model, determining a first wake-up result, and outputting the wake-up speech to the cloud; the cloud may include a second wake-up model for performing a second wake-up on the wake-up speech input to the second wake-up model and determining a second wake-up result.

[0044] In another related art, the acoustic features of the wake-up speech can be analyzed by using a preset acoustic model and a preset wake-up word recognition network of the intelligent terminal to obtain the confidence of the acoustic features of the wake-up speech relative to the acoustic features of the preset wake-up word; then it is determined whether the confidence coefficient falls within a preset medium confidence coefficient range. If so, the wake-up speech is uploaded to the remote server; and it is determined whether the language features obtained by analyzing the wake-up speech through the language model match the language features of the preset wake-up word. The method of combining on-device wake-up and cloud wake-up in the related art has the following problems: 1. The on-device voice wake-up model and the cloud voice wake-up model are essentially two different models. In the related art, it is not considered that they can be jointly trained during the training process, and the effect of the cloud voice wake-up model can be used to assist in improving the effect of the on-device voice wake-up model during the training process; 2. The voice wake-up model essentially belongs to a type of speech recognition model, only the modeling units are different. In the related art, it is not considered that a large amount of data for speech recognition can be used to assist in improving the effect of the voice wake-up model; 3. It is not considered that part of the computational workload of the cloud voice wake-up model can be optimized together with the computational workload of speech recognition.

[0045] In view of this, the embodiments of the present disclosure provide a model training method, which can improve the effect of the voice wake-up model and make the wake-up result of the voice wake-up model more accurate.

[0046] Figure 1 The flowchart showing a model training method according to an embodiment of the present disclosure is asFigure 1 As shown, the method may include:

[0047] S101. Obtain sample data; wherein, the sample data includes speech data and text data corresponding to the speech data.

[0048] Exemplarily, speech recognition sample data and speech wake-up sample data may be obtained respectively. The speech recognition sample data may include speech data and corresponding text data, and the speech wake-up sample data may include speech data with wake-up words and corresponding text data as well as speech data without wake-up words and corresponding text data. Exemplarily, the number of speech recognition sample data may be greater than the number of speech wake-up sample data. As an example, the number of speech recognition sample data may be much greater than the number of speech wake-up sample data. For example, the number of speech recognition sample data may be ten times the number of speech wake-up sample data.

[0049] S102. Input the sample data into a speech recognition model, train the speech recognition model, and obtain a first feature output by an intermediate layer of the speech recognition model.

[0050] Exemplarily, the speech recognition model may include a first sub-model, a second sub-model, and a decoder. The input of the first sub-model is the speech data in a large amount of speech recognition sample data, and its output is speech features with a certain high-dimensional representation. The input of the second sub-model is the speech features with a high-dimensional representation, and the output is text related to semantics. The number of layers of the first sub-model may be less than the number of layers of the second sub-model. For example, the number of layers of the first sub-model may be three, and the number of layers of the second sub-model may be six. Each layer of the neural network module may be an attention mechanism module. For example, it may be a Conformer module, and Conformer is a convolution-enhanced Transformer. Exemplarily, the first feature may be the feature output by the first sub-model.

[0051] Figure 2 Show a schematic diagram of the training process of a speech recognition model according to an embodiment of the present disclosure, as Figure 2As shown in the figure, the speech recognition model may include a first sub-model, a second sub-model, and a speech recognition decoder (ASR Decorder); among them, the number of layers of the first sub-model may be three, and the number of layers of the second sub-model may be five. Each layer of the neural network module may be a conformer module with convolutional enhancement, that is, a Conformer module. Sample data may be input into the speech recognition model for training. The sample data may include speech data and corresponding text data; after passing through the first sub-model, the speech data may output features of a certain high dimension modeled from the speech signal (i.e., the first feature). After the first feature passes through the second sub-model, it may output text related to semantics. After passing through the speech recognition decoder, the result of speech recognition may be output. The loss function of the speech recognition model may be denoted as the first loss function (i.e., Figure 2 Loss1 in

[0052] ). The first loss may be calculated using the first loss function based on the input text data and the speech recognition result output by the speech recognition model. According to the first loss, the parameters of the speech recognition model may be adjusted. When the preset training conditions are met, for example, when the value of the first loss is less than the preset threshold, the training may be ended to obtain a trained speech recognition model.

[0053] S103. When the wake-up word is included in the sample data, input the first feature into the first speech wake-up model to train the first speech wake-up model.

[0054] Exemplarily, the first speech wake-up model may be pre-trained using speech wake-up sample data.

[0055] Exemplarily, when the sample data input into the speech recognition model is speech wake-up sample data (i.e., when the wake-up word is included in the sample data), the first feature output by the first sub-model in the speech recognition model may be input into the pre-trained first speech wake-up model to continue training the first speech wake-up model.

[0056] Exemplarily, the first feature may be input into the first speech wake-up model after dimensional transformation.

[0057] Figure 3Schematic diagram showing the training process of a first voice wake-up model according to an embodiment of the present disclosure Figure 3 The voice recognition model in Figure 2 can refer to the relevant description of the voice recognition model in Figure 3 As shown, when the sample data input to the voice recognition model is voice wake-up sample data, the first feature output by the first sub-model in the voice recognition model can be input into the first voice wake-up model to train the first voice wake-up model; since the first feature is a high-dimensional feature, the dimension of the first feature can be transformed, and after the dimension of the first feature is transformed into a feature dimension suitable for input into the first voice wake-up model, it is then input.

[0058] The first voice wake-up model can include a first voice wake-up sub-model and a keyword spotting decoder (KWSDecorder). The first feature after dimension transformation first passes through the first voice wake-up sub-model and then through the keyword spotting decoder, and the wake-up probability for voice wake-up (i.e., the probability that the input voice data contains a wake-up word) can be output. The loss function of the first voice wake-up model can be denoted as the second loss function (i.e., Figure 3 Loss2 in ), and the parameters of the first voice wake-up model can be adjusted according to the second loss function. When the preset training conditions are met, for example, when the value of the second loss function is less than a preset threshold, the training can be ended to obtain the trained first voice wake-up model.

[0059] In this way, by inputting the features output by the intermediate layer of the voice recognition model into the first voice wake-up model for training, the features learned by the voice recognition model from a large number of sample data trainings can be used to assist in improving the effect of the first voice wake-up model, increasing the wake-up rate of the first voice wake-up model, and making the wake-up result more accurate.

[0060] In the embodiment of the present disclosure, when the sample data contains a wake-up word, the first feature output by the intermediate layer of the voice recognition model is input into the pre-trained first voice wake-up model to continue training the first voice wake-up model. Since the feature extraction ability of the intermediate layer of the voice recognition model is stronger than that of the pre-trained first voice wake-up model, and the features extracted by the intermediate layer of the voice recognition model can contain semantic information, therefore, using the first feature extracted by the intermediate layer of the voice recognition model to train the first voice wake-up model can improve the effect of the first voice wake-up model, increase the wake-up rate of the first voice wake-up model, and make the wake-up result of the first voice wake-up model more accurate.

[0061] Figure 4 Flowchart showing a model training method according to an embodiment of the present disclosure. As Figure 4 shown, the method may include:

[0062] S401. Obtain sample data; wherein, the sample data includes voice data and text data corresponding to the voice data.

[0063] S402. Input the sample data into a speech recognition model, train the speech recognition model, and obtain a first feature output by an intermediate layer of the speech recognition model.

[0064] Steps S401 - S402 are the same as steps S101 - S102 above Figure 1 and will not be elaborated here.

[0065] S403. When the sample data contains a wake - up word, input the first feature into a first voice wake - up model and train the first voice wake - up model.

[0066] Among them, the first voice wake - up model can be located in the cloud.

[0067] Exemplarily, the steps for training the first voice wake - up model can refer to step S103 above Figure 1 and will not be elaborated here.

[0068] S404. When the sample data contains a wake - up word, train a second voice wake - up model according to the sample data and the first voice wake - up model.

[0069] Among them, the second voice wake - up model can be located at the terminal.

[0070] Exemplarily, the second voice wake - up model can be pre - trained using voice wake - up sample data.

[0071] In a possible implementation manner, when inputting the sample data into the speech recognition model in step S402 above, the same sample data can be input into a pre - trained second voice wake - up model to continue training the second voice wake - up model. When the input sample data is voice wake - up sample data (i.e., when the sample data contains a wake - up word), the second voice wake - up model can be trained according to the sample data and the first voice wake - up model.

[0072] Exemplarily, the process of training the second voice wake - up model can include:

[0073] (1) Input the sample data into the second voice wake - up model to obtain a second probability distribution output by the second voice wake - up model.

[0074] Input the voice wake-up sample data into the second voice wake-up model, and the second voice wake-up model can output the wake-up probability for voice wake-up (i.e., the probability that the input voice data contains the wake-up word). During the training process of the second voice wake-up model, the posterior probability distribution output by the second voice wake-up model can be obtained, denoted as the second probability distribution. At the same time, the loss function of the second voice wake-up model can be obtained, denoted as the third loss function.

[0075] (2) Adjust the parameters of the second voice wake-up model according to the distance loss function, so that the second probability distribution output by the second voice wake-up model with adjusted parameters is close to the first probability distribution output by the first voice wake-up model; wherein, the distance loss function is used to determine the loss of the second probability distribution relative to the first probability distribution.

[0076] During the training process of the first voice wake-up model, the posterior probability distribution output by the first voice wake-up model can be denoted as the first probability distribution. The distance loss function can be calculated based on the first probability distribution and the second probability distribution. As an example, the Kullback-Leibler (KL) divergence between the first probability distribution and the second probability distribution can be calculated to obtain the distance loss function, and the calculation process of the KL divergence can refer to related technologies.

[0077] The parameters of the second voice wake-up model can be adjusted according to the third loss function; at the same time, the parameters of the second voice wake-up model can also be adjusted according to the distance loss function, so that the second probability distribution output by the second voice wake-up model with adjusted parameters is close to the first probability distribution. In this way, since the second voice wake-up model is deployed to the terminal, the model size and power consumption of the second voice wake-up model need to be considered during training, while the first voice wake-up model is deployed to the cloud, the number of parameters and the model scale of the first voice wake-up model can be increased, and the effect of the first voice wake-up model can be better than that of the second voice wake-up model; during the training process of the second voice wake-up model, by making the second probability distribution close to the first probability distribution, the first voice wake-up model can be used to further optimize the second voice wake-up model and improve the effect of the second voice wake-up model.

[0078] In an embodiment of the present disclosure, when the sample data includes a wake word, the first feature output by the intermediate layer of the speech recognition model is input into the first voice wake-up model located in the cloud to train the first voice wake-up model; according to the sample data and the first voice wake-up model located in the cloud, the second voice wake-up model located in the terminal is trained, which can improve the effect of the first voice wake-up model, and further optimize the second voice wake-up model by using the first voice wake-up model, improve the effect of the second voice wake-up model, thereby improving the wake-up rate of the first voice wake-up model and the second voice wake-up model, and making the wake-up results of the first voice wake-up model and the second voice wake-up model more accurate.

[0079] In one embodiment, the speech recognition model, the first voice wake-up model, and the second voice wake-up model can be jointly trained.

[0080] Figure 5 The figure shows a schematic diagram of jointly training the speech recognition model, the first voice wake-up model, and the second voice wake-up model according to an embodiment of the present disclosure. Figure 5 The speech recognition model in Figure 2 can refer to the relevant description of the speech recognition model in Figure 5 The first voice wake-up model in Figure 3 can refer to the relevant description of the first voice wake-up model in Figure 5 As shown in Figure 5 , the speech recognition model and the first voice wake-up model can be located in the cloud, and the second voice wake-up model can be located in the terminal. The sample data can be input into the speech recognition model and the second voice wake-up model at the same time to train the speech recognition model and the second voice wake-up model. The sample data can include speech recognition sample data and voice wake-up sample data. When the input sample data is speech recognition sample data, the parameters of the speech recognition model can be updated according to the first loss function of the speech recognition model (i.e., Loss1 in Figure 5 ), and the parameters of the first voice wake-up model and the second voice wake-up model do not need to be updated.

[0081] When the input sample data is voice wake-up sample data, the first feature output by the first sub-model of the speech recognition model can be input into the first voice wake-up model after dimensional transformation to train the first voice wake-up model. In this way, the features learned by the speech recognition model from a large number of sample data trainings can be used to assist in improving the effect of the first voice wake-up model.

[0082] The loss function of the first voice wake-up model is denoted as the second loss function (i.e., Loss2 in Figure 5 ), and the posterior probability distribution output by the first voice wake-up model is denoted as the first probability distribution (i.e., Figure 5P) in it; the second voice wake-up model can be trained according to the input voice wake-up sample data. The second voice wake-up model can include a second voice wake-up sub-model and a voice wake-up decoder (KWSDecorder). After the voice wake-up sample data passes through the second voice wake-up sub-model and then through the voice wake-up decoder, the wake-up probability for voice wake-up can be output (i.e., the probability that the input voice wake-up sample data contains a wake-up word).

[0083] The loss function of the second voice wake-up model is denoted as the third loss function (i.e., Figure 5 Loss3 in it), and the posterior probability distribution output by the second voice wake-up model is denoted as the second probability distribution; the KL divergence between the first probability distribution and the second probability distribution can be calculated to obtain a distance loss function, and the parameters of the second voice wake-up model can be adjusted according to the distance loss function to make the second probability distribution close to the first probability distribution; in this way, the first voice wake-up model in the cloud can be used to further improve the effect of the second voice wake-up model on the terminal.

[0084] When the input sample data is voice wake-up sample data, the total loss function can be determined according to the first loss function, the second loss function, the third loss function, and the distance loss function. As an example, the first loss function, the second loss function, the third loss function, and the distance loss function can be weighted and summed to obtain the total loss function. Those skilled in the art can adjust the weights according to the actual training situation to determine the total loss function, so as to achieve better training effects. Or, the weights can also be adjustable training parameters, and specific weights can be obtained after training.

[0085] The voice recognition model, the first voice wake-up model, and the second voice wake-up model can be updated according to the total loss function. When the preset training conditions are met, for example, when the value of the total loss function is less than a preset threshold, the training can be stopped to obtain the trained voice recognition model, the first voice wake-up model, and the second voice wake-up model.

[0086] In this way, by combining the training of the voice recognition model, the first voice wake-up model, and the second voice wake-up model, using the first voice wake-up model in the cloud to assist in improving the effect of the second voice wake-up model on the terminal, and using the voice recognition model to assist in improving the effect of the first voice wake-up model, the wake-up rates of the first voice wake-up model and the second voice wake-up model can be improved, making the wake-up results of the first voice wake-up model and the second voice wake-up model more accurate.

[0087] The above describes the model training method provided by the embodiments of the present disclosure. Next, the wake-up method provided by the embodiments of the present disclosure will be described from the application side.

[0088] Figure 6The flowchart of a wake-up method according to an embodiment of the present disclosure is shown. As Figure 6 shown, the method may include:

[0089] S601. Receive the input voice.

[0090] Exemplarily, the user may input voice to the terminal through an audio acquisition device such as a microphone, and the terminal may receive the voice input by the user.

[0091] Exemplarily, the second voice wake-up model may be trained according to the model training method shown in steps S401 to S404 above, and the second voice wake-up model may be deployed on the terminal, for example, on terminal devices such as mobile phones and in-vehicle chips; those skilled in the art may preset the first threshold and the second threshold of the second voice wake-up model according to related technologies. For example, the first threshold may be set to α Figure 4 according to the development dataset, and the second threshold may be set to α high ; where the first threshold is greater than the second threshold; after the terminal receives the voice input by the user, the voice may be input into the second voice wake-up model of the terminal to obtain the first wake-up probability P1 output by the second voice wake-up model (i.e., the probability of waking up the terminal according to the voice); if the first wake-up probability is greater than the first threshold (i.e., P1>α low ), the terminal may be woken up; if the first wake-up probability is less than the second threshold (P1<α high ), the terminal is not woken up; if the first wake-up probability is between the first threshold and the second threshold (i.e., α low ≤P1≤α low ), it is determined that it is impossible to determine whether to wake up the terminal according to the voice. In this case, the following steps may be performed. high

[0092] S602. When the terminal cannot determine whether to wake up according to the voice, input the voice into the voice recognition model in the cloud to obtain the second feature output by the intermediate layer of the voice recognition model.

[0093] Exemplarily, the voice recognition model may be trained according to the model training method shown in steps S101 to S102 above or Figure 1 in steps S401 to S402 above, and the voice recognition model may be deployed in the cloud; when the terminal cannot determine whether to wake up according to the input voice, the voice may be input into the voice recognition model in the cloud to obtain the second feature output by the intermediate layer of the voice recognition model. Figure 4

[0094] S603. Input the second feature into the first voice wake-up model in the cloud to determine whether to wake up. ​​

[0095] Exemplarily, it can be based on the above Figure 1 steps S101 to S103 in or Figure 4 the model training method shown in steps S401 to S403 in to train the first voice wake-up model, and the first voice wake-up model can be deployed in the cloud; those skilled in the art can preset the third threshold of the first voice wake-up model according to related technologies. For example, the third threshold can be set to β according to the development dataset high ; the second feature can be input into the first voice wake-up model in the cloud to obtain the second wake-up probability P2 output by the first voice wake-up model (that is, the probability of waking up the terminal according to this voice); if the second wake-up probability is greater than the third threshold (that is, P2>β high ), it is determined to wake up the terminal; otherwise, the terminal is not woken up. This wake-up process can be called secondary wake-up. In this way, when the second voice wake-up model of the terminal cannot determine whether to wake up the terminal according to the input voice, the second feature output by the intermediate layer of the speech recognition model is input into the first voice wake-up model in the cloud, and the first voice wake-up model in the cloud is used to determine whether to wake up the terminal, which can improve the wake-up rate and ensure the wake-up latency performance.

[0096] In a possible implementation manner, when the second feature is input into the first voice wake-up model in the cloud and it is determined to wake up, the method may further include: inputting the second feature into the module after the intermediate layer of the speech recognition model for speech recognition.

[0097] Exemplarily, after the first voice wake-up model in the cloud determines to wake up the terminal and wakes up the terminal, the second feature can be input into the module after the intermediate layer of the speech recognition model for speech recognition, and the recognition result can be sent to the downstream task for subsequent processing. As an example, the module after the intermediate layer of the speech recognition model may include a second sub-model and a speech recognition decoder. The input of the second sub-model is the second feature, and the output is text related to semantics. After this text passes through the speech recognition decoder, the result of speech recognition can be output.

[0098] Since the first voice wake-up model in the cloud has used a part of the speech recognition model (that is, the intermediate layer of the speech recognition model) for forward inference during the process of determining whether to wake up, after waking up, the speech does not need to pass through the speech recognition model from the beginning, but can continue to be recognized by the module after the intermediate layer of the speech recognition model from the output of the intermediate layer of the speech recognition model (that is, the second feature); this can reduce the computational amount and power consumption of part of the speech recognition.

[0099] In the case where the terminal cannot determine whether to wake up according to the input voice, the voice is input into the voice recognition model in the cloud, and the second feature output by the middle layer of the voice recognition model is obtained; the second feature is input into the first voice wake-up model in the cloud to determine whether to wake up; the wake-up rate can be improved, the latency performance of wake-up can be guaranteed, and part of the computational overhead in the voice recognition process after wake-up can be reduced.

[0100] Figure 7 Fig. shows a schematic diagram of waking up a terminal according to an embodiment of the present disclosure. Figure 7 The voice recognition model in can refer to Figure 2 the relevant description of the voice recognition model in; as Figure 7 shown, the voice input by the user can be input into the second voice wake-up model deployed on the terminal to obtain the first wake-up probability P1; the first threshold α of the second voice wake-up model can be preset according to the development dataset high and the second threshold α low , where α high >α low ; if P1>α high , the terminal can be directly woken up, so that the terminal can be quickly woken up; if P1<α low , no operation can be performed; if α low ≤P1≤α high , the input voice can be input into the voice recognition model deployed in the cloud to obtain the second feature output by the first sub-model of the voice recognition model, and the second feature is input into the first voice wake-up model in the cloud after dimensional transformation to obtain the second wake-up probability P2; the third threshold β of the first voice wake-up model can be preset according to the development dataset high , if P2>β high , the terminal is woken up, otherwise the terminal is not woken up; in this way, by inputting the second feature output by the first sub-model of the voice recognition model into the first voice wake-up model in the cloud and using the first voice wake-up model in the cloud to determine whether to wake up the terminal, the wake-up rate can be improved and the latency performance of wake-up can be guaranteed. After the terminal is woken up, the voice can be input into the voice recognition model for recognition. Since the first voice wake-up model in the cloud uses the first sub-model of the voice recognition model for forward inference in the process of determining whether to wake up, after waking up, the voice does not need to pass through the voice recognition model from the beginning, but can continue from the output of the first sub-model through the second sub-model and the voice recognition decoder to obtain the voice recognition result for sending to the downstream task for subsequent processing; in this way, in the case of secondary wake-up, part of the computational overhead in the voice recognition process in the voice interaction link can be reduced.

[0101] Based on the same inventive concept as the embodiments of the above model training method, an embodiment of the present disclosure further provides a model training apparatus, which can be used to execute the technical solutions described in the embodiments of the above model training method. For example, it can execute each step of the model training method shown in the following Figure 1 or Figure 4 .

[0102] Figure 8 FIG. shows a schematic structural diagram of a model training apparatus according to an embodiment of the present disclosure. As shown in Figure 8 , the apparatus may include: an acquisition module 801, configured to acquire sample data; wherein the sample data includes speech data and text data corresponding to the speech data; a speech recognition model training module 802, configured to input the sample data into a speech recognition model, train the speech recognition model, and obtain first features output by an intermediate layer of the speech recognition model; a first speech wake-up model training module 803, configured to, when the sample data includes a wake-up word, input the first features into a first speech wake-up model, and train the first speech wake-up model.

[0103] In a possible implementation manner, the first speech wake-up model is located in the cloud, and the apparatus further includes: a second speech wake-up model training module, configured to, when the sample data includes a wake-up word, train a second speech wake-up model according to the sample data and the first speech wake-up model, where the second speech wake-up model is located in a terminal.

[0104] In a possible implementation manner, the second speech wake-up model training module is further configured to: input the sample data into the second speech wake-up model to obtain a second probability distribution output by the second speech wake-up model; adjust parameters of the second speech wake-up model according to a distance loss function, so that the second probability distribution output by the second speech wake-up model with adjusted parameters is close to a first probability distribution output by the first speech wake-up model; wherein the distance loss function is used to determine a loss of the second probability distribution relative to the first probability distribution.

[0105] In a possible implementation, the apparatus further includes: a total loss function determination module, configured to determine a total loss function according to a first loss function, a second loss function, a third loss function, and the distance loss function when the sample data includes a wake-up word; the speech recognition model training module 802 is further configured to update the speech recognition model according to the total loss function; the first voice wake-up model training module 803 is further configured to update the first voice wake-up model according to the total loss function; the second voice wake-up model training module is further configured to update the second voice wake-up model according to the total loss function; wherein, the first loss function is the loss function of the speech recognition model; the second loss function is the loss function of the first voice wake-up model; the third loss function is the loss function of the second voice wake-up model.

[0106] In a possible implementation, the total loss function determination module is further configured to: perform a weighted sum of the first loss function, the second loss function, the third loss function, and the distance loss function to obtain a total loss function.

[0107] The model training apparatus provided in the embodiments of the present disclosure trains the first voice wake-up model by inputting the first feature output by the intermediate layer of the speech recognition model into the first voice wake-up model when the sample data includes a wake-up word; since the feature extraction ability of the speech recognition model is stronger than that of the first voice wake-up model, and the features extracted by the speech recognition model can include semantic information, training the first voice wake-up model using the features output by the intermediate layer of the speech recognition model can improve the effect of the first voice wake-up model, increase the wake-up rate of the first voice wake-up model, and make the wake-up result of the first voice wake-up model more accurate.

[0108] Based on the same inventive concept as the above-mentioned wake-up method embodiments, the embodiments of the present disclosure further provide a wake-up apparatus, which can be used to execute the technical solutions described in the above-mentioned wake-up method embodiments. For example, it can execute each step of the wake-up method shown above. Figure 6 The steps of the shown wake-up method.

[0109] Figure 9 A schematic structural diagram of a wake-up apparatus according to an embodiment of the present disclosure is shown, as Figure 9 shown, the apparatus may include: a receiving module 901, configured to receive input speech; a second feature obtaining module 902, configured to input the speech into a speech recognition model in the cloud to obtain a second feature output by the intermediate layer of the speech recognition model when the terminal cannot determine whether to perform wake-up according to the speech; a first wake-up module 903, configured to input the second feature into a first voice wake-up model in the cloud to determine whether to perform wake-up.

[0110] In a possible implementation, the device further includes a second wake-up module, configured to: before the second feature acquisition module 902 inputs the speech to a speech recognition model in the cloud to obtain a second feature output by an intermediate layer of the speech recognition model when the terminal cannot determine whether to wake up based on the speech, input the speech to a second speech wake-up model of the terminal to obtain a first wake-up probability; if the first wake-up probability is greater than a first threshold, wake up the terminal; if the first wake-up probability is less than a second threshold, do not wake up the terminal; if the first wake-up probability is between the first threshold and the second threshold, determine that the terminal cannot determine whether to wake up based on the speech; wherein, the first threshold is greater than the second threshold.

[0111] In a possible implementation, the device further includes a speech recognition module, configured to: when the first wake-up module 903 inputs the second feature to a first speech wake-up model in the cloud and determines to wake up, input the second feature to a module after an intermediate layer of the speech recognition model for speech recognition.

[0112] The wake-up device provided by the embodiments of the present disclosure, when the terminal cannot determine whether to wake up based on the input speech, inputs the speech to a speech recognition model in the cloud to obtain a second feature output by an intermediate layer of the speech recognition model; inputs the second feature to a first speech wake-up model in the cloud to determine whether to wake up; can improve the wake-up rate, ensure the latency performance of waking up, and can reduce part of the computational overhead in the speech recognition process after waking up.

[0113] The embodiments of the present disclosure further propose a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented. Exemplarily, the steps of the above Figure 1 or Figure 4 shown model training method can be executed, or the steps of the above Figure 6 shown wake-up method can be executed. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0114] The embodiments of the present disclosure further propose an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory. Exemplarily, the steps of the above Figure 1 or Figure 4 shown model training method can be executed, or the steps of the above Figure 6 shown wake-up method can be executed.

[0115] Embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method. Exemplarily, the steps of the above Figure 1 or Figure 4 shown model training method can be executed, or the steps of the above Figure 6 shown wake-up method can be executed.

[0116] Figure 10 FIG. 10 shows a schematic structural diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 10 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method. Exemplarily, the steps of the above Figure 1 or Figure 4 shown model training method can be executed, or the steps of the above Figure 6 shown wake-up method can be executed.

[0117] The electronic device 1900 may further include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0118] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions. The above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method. Exemplarily, the steps of the above Figure 1 or Figure 4 shown model training method can be executed, or the steps of the above Figure 6 shown wake-up method can be executed.

[0119] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present disclosure.

[0120] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0121] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a respective computing / processing device, or may be downloaded to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0122] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0123] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0124] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0125] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0126] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0127] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A model training method, characterized in that, The method includes: Obtaining sample data; wherein, the sample data includes voice data and text data corresponding to the voice data; Inputting the sample data into a speech recognition model to train the speech recognition model, and obtaining a first feature output by an intermediate layer of the speech recognition model; the speech recognition model is used to recognize text corresponding to speech; the first feature extracted by the intermediate layer contains semantic information; the first feature includes: features output after the voice data in the sample data passes through a first sub-model of the speech recognition model; When the sample data contains a wake-up word, inputting the first feature into a first voice wake-up model to train the first voice wake-up model; the first voice wake-up model is used to determine whether to perform wake-up.

2. The method according to claim 1, wherein The first voice wake-up model is located in the cloud, and the method further includes: When the sample data contains a wake-up word, training a second voice wake-up model according to the sample data and the first voice wake-up model, and the second voice wake-up model is located at the terminal.

3. The method according to claim 2, wherein The step of, when the sample data contains a wake-up word, training the second voice wake-up model according to the sample data and the first voice wake-up model includes: Inputting the sample data into the second voice wake-up model to obtain a second probability distribution output by the second voice wake-up model; Adjusting parameters of the second voice wake-up model according to a distance loss function, so that the second probability distribution output by the second voice wake-up model with adjusted parameters approaches the first probability distribution output by the first voice wake-up model; wherein, the distance loss function is used to determine the loss of the second probability distribution relative to the first probability distribution.

4. The method according to claim 3, wherein The training process of the speech recognition model, the first voice wake-up model, and the second voice wake-up model includes: When the sample data contains a wake-up word, determining a total loss function according to a first loss function, a second loss function, a third loss function, and the distance loss function; Updating the speech recognition model, the first voice wake-up model, and the second voice wake-up model according to the total loss function; wherein, the first loss function is the loss function of the speech recognition model; the second loss function is the loss function of the first voice wake-up model; the third loss function is the loss function of the second voice wake-up model.

5. The method according to claim 4, wherein The step of, when the sample data contains a wake-up word, determining a total loss function according to a first loss function, a second loss function, a third loss function, and the distance loss function includes: Performing weighted summation on the first loss function, the second loss function, the third loss function, and the distance loss function to obtain a total loss function.

6. A wake-up method, characterized in that, The method includes: Receiving input voice; In the case where the terminal cannot determine whether to perform wake-up based on the voice, the voice is input into a voice recognition model in the cloud to obtain a second feature output by an intermediate layer of the voice recognition model; the voice recognition model is used to recognize the text corresponding to the voice; the second feature extracted by the intermediate layer contains semantic information; the second feature includes: the feature output after the voice in the voice passes through a first sub-model of the voice recognition model. The second feature is input into a first voice wake-up model in the cloud to determine whether to perform wake-up.

7. The method according to claim 6, characterized in that Before the voice is input into the voice recognition model in the cloud to obtain the second feature output by the intermediate layer of the voice recognition model in the case where the terminal cannot determine whether to perform wake-up based on the voice, it further includes: The voice is input into a second voice wake-up model of the terminal to obtain a first wake-up probability. If the first wake-up probability is greater than a first threshold, the terminal is woken up. If the first wake-up probability is less than a second threshold, the terminal is not woken up. If the first wake-up probability is between the first threshold and the second threshold, it is determined that the terminal cannot determine whether to perform wake-up based on the voice. Wherein, the first threshold is greater than the second threshold.

8. The method according to claim 6, characterized in that, In the case where the second feature is input into the first voice wake-up model in the cloud and it is determined to perform wake-up, the method further includes: The second feature is input into a module after the intermediate layer of the voice recognition model to perform voice recognition.

9. A model training device, characterized in that, The device includes: An acquisition module, configured to acquire sample data; wherein, the sample data includes voice data and text data corresponding to the voice data. A voice recognition model training module, configured to input the sample data into a voice recognition model to train the voice recognition model to obtain a first feature output by an intermediate layer of the voice recognition model; the voice recognition model is used to recognize the text corresponding to the voice; the first feature extracted by the intermediate layer contains semantic information; the first feature includes: the feature output after the voice data in the sample data passes through a first sub-model of the voice recognition model. A first voice wake-up model training module, configured to input the first feature into a first voice wake-up model to train the first voice wake-up model in the case where the sample data includes a wake-up word; the first voice wake-up model is used to determine whether to perform wake-up.

10. A wake-up device, characterized in that, The device includes: A receiving module, configured to receive the input voice. A second feature acquisition module, configured to input the voice into a voice recognition model in the cloud to obtain a second feature output by an intermediate layer of the voice recognition model in the case where the terminal cannot determine whether to perform wake-up based on the voice; the voice recognition model is used to recognize the text corresponding to the voice; the second feature extracted by the intermediate layer contains semantic information; the second feature includes: the feature output after the voice passes through a first sub-model of the voice recognition model. A first wake-up module, configured to input the second feature into a first voice wake-up model in the cloud to determine whether to perform wake-up.

11. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, when the processor is configured to execute the instructions, it implements the method described in any one of claims 1-5, or implements the method described in any one of claims 6-8.

12. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, it implements the method described in any one of claims 1-5, or implements the method described in any one of claims 6-8.

Citation Information

Patent Citations

  • Voice wake-up method and device

    CN107622770A

  • Model training method and device, voice wake-up method and device, equipment and medium

    CN116645960A