Training method and device of wake-up model, and wake-up method and device

CN122531364APending Publication Date: 2026-08-07JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINGDONG TECH HLDG CO LTD
Filing Date
2026-05-09
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明提供一种唤醒模型的训练方法、唤醒方法及装置,用以解决现有技术中降低最终微调得到的唤醒模型的准确性的缺陷

Benefits of technology

[0019]The wake-up model training method, wake-up method, and apparatus provided by this invention involve inputting the first negative samples (not including wake-up words) into an initial wake-up model to obtain the posterior probability of each articulation unit in each frame of the first negative samples output by the initial wake-up model. For each frame, the first posterior probability of each target articulation unit in the wake-up word is obtained from the posterior probabilities of all articulation units in the frame. Based on the comparison between each first posterior probability and a selection threshold, if the first negative sample is determined to be a negative sample for model fine-tuning, it is identified as a target negative sample. Finally, the initial wake-up model is fine-tuned based on each target negative sample to obtain the wake-up model. It can be seen that this invention selects target negative samples for model fine-tuning based on the comparison between the posterior probability output by the initial wake-up model and the selection threshold. This selection method is relatively objective, enabling the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the final fine-tuned wake-up model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531364A_ABST
    Figure CN122531364A_ABST
Patent Text Reader

Abstract

The application provides a training method and device of a wake-up model, and a wake-up method, and relates to the technical field of wake-up. The method comprises the following steps: obtaining a first negative sample, inputting the first negative sample into an initial wake-up model to obtain the posterior probability of each pronunciation unit in each frame of the first negative sample; obtaining the first posterior probability of each target pronunciation unit in a wake-up word from the posterior probabilities of all the pronunciation units in the frame; determining the first negative sample as a target negative sample when it is determined that the first negative sample is a negative sample participating in model fine-tuning based on the comparison result of each first posterior probability and a screening threshold; and fine-tuning the initial wake-up model based on each target negative sample to obtain a wake-up model. The application screens the target negative sample participating in model fine-tuning based on the comparison result of the posterior probability output by the initial wake-up model and the screening threshold, the screening method is relatively objective, the initial wake-up model can perceive its own weaknesses, and thus the accuracy of the wake-up model obtained through final fine-tuning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wake-up technology, and in particular to a method for training a wake-up model, a wake-up method, and a device. Background Technology

[0002] With the rapid development of deep learning technology and hardware, users have placed higher demands on wake-up performance, requiring higher wake-up rates, lower latency, and fewer false wake-ups. However, due to cost constraints or the development of new features, the resources actually allocated to wake-up have not only failed to increase but have actually decreased. Therefore, how to deliver a superior wake-up system with even more limited resources has become a key research area.

[0003] In related technologies, an initial wake-up model is usually trained using positive and negative samples, and then fine-tuned based on randomly stacked negative samples to obtain the final wake-up model.

[0004] However, the aforementioned technologies rely on randomly stacked negative samples to fine-tune the initial wake-up model. Randomly stacked negative samples are subjective and cannot detect the model's weaknesses, thus reducing the accuracy of the final fine-tuned wake-up model. Summary of the Invention

[0005] This invention provides a training method, a wake-up method, and an apparatus for a wake-up model, which addresses the shortcomings of existing technologies that reduce the accuracy of the wake-up model obtained through final fine-tuning.

[0006] This invention provides a training method for a wake-up model, comprising the following steps.

[0007] Obtain at least one first negative sample, wherein the first negative sample does not include a wake word; For each of the first negative samples, the first negative sample is input into the initial wake-up model to obtain the posterior probability of each vocal unit in each frame of the first negative sample output by the initial wake-up model; For each frame, the first posterior probability of each target pronunciation unit in the wake word is obtained from the posterior probabilities of all pronunciation units in the frame. If, based on the comparison results between each of the first posterior probabilities and the screening threshold, the first negative sample is determined to be a negative sample participating in model fine-tuning, then the first negative sample is determined as the target negative sample. The initial wake-up model is fine-tuned based on each of the target negative samples to obtain the wake-up model.

[0008] According to a training method for a wake-up model provided by the present invention, based on the comparison results of each first posterior probability and a screening threshold, the first negative sample is determined as a negative sample participating in model fine-tuning, including: For each target articulation unit, a second posterior probability greater than the filtering threshold is determined from the first posterior probabilities of the target articulation unit in all the frames; Determine the target position of the target frame corresponding to the second posterior probability in the first negative sample; If the pronunciation of the target position in the first negative sample is different from the pronunciation of the target pronunciation unit, the first negative sample is determined to be a negative sample participating in model fine-tuning.

[0009] According to a training method for a wake-up model provided by the present invention, the step of fine-tuning the initial wake-up model based on each of the target negative samples to obtain a wake-up model includes: For each target negative sample, if the first posterior probability of the target vocal unit of the target negative sample is greater than or equal to the wake-up threshold, a first weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample after setting the first weight to obtain the wake-up model; If the first posterior probability of the target vocal unit of the target negative sample is less than the wake-up threshold and greater than the screening threshold, a second weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample with the second weight set to obtain the wake-up model, wherein the second weight is less than the first weight.

[0010] According to a training method for a wake-up model provided by the present invention, obtaining at least one first negative sample includes: The at least one first negative sample is obtained based on the rules covering the equalization of the pronunciation units.

[0011] According to a training method for a wake-up model provided by the present invention, the method further includes: Obtain a second negative sample and a positive sample, wherein the second negative sample does not include the wake word, and the positive sample includes the wake word; Based on the second negative sample and the positive sample, the model structure is trained to obtain the initial wake-up model.

[0012] The present invention provides a wake-up method, comprising the following steps.

[0013] Acquire the target speech; The target speech is input into the wake-up model to obtain the recognition result output by the wake-up model; The wake-up is performed if the recognition result indicates that the target speech includes a wake-up word. The wake-up model is trained based on any of the methods described above.

[0014] The present invention also provides a training apparatus for a wake-up model, comprising: The first acquisition unit is used to acquire at least one first negative sample, wherein the first negative sample does not include a wake word; The first recognition unit is used to input the first negative sample into the initial wake-up model for each first negative sample, and obtain the posterior probability of each sound unit in each frame of the first negative sample output by the initial wake-up model. The first determining unit is configured to, for each frame, obtain the first posterior probability of each target pronunciation unit in the wake word from the posterior probabilities of all pronunciation units in the frame; The second determining unit is used to determine the first negative sample as the target negative sample when the first negative sample is determined to be a negative sample participating in model fine-tuning based on the comparison results of each of the first posterior probabilities and the screening threshold. The fine-tuning unit is used to fine-tune the initial wake-up model based on each of the target negative samples to obtain the wake-up model.

[0015] The present invention also provides a wake-up device, comprising: The second acquisition unit is used to acquire the target speech; The second recognition unit is used to input the target speech into the wake-up model and obtain the recognition result output by the wake-up model; A wake-up unit is used to wake up the target speech if the recognition result indicates that a wake-up word is included in the target speech; The wake-up model is trained based on any of the methods described above.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a training method for a wake-up model as described above, or to implement a wake-up method as described above.

[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for a wake-up model as described above, or implements a wake-up method as described above.

[0018] The present invention also provides a computer program product, comprising a computer program that, when executed by a processor, implements a training method for a wake-up model as described above, or implements a wake-up method as described above.

[0019] The wake-up model training method, wake-up method, and apparatus provided by this invention involve inputting the first negative samples (not including wake-up words) into an initial wake-up model to obtain the posterior probability of each articulation unit in each frame of the first negative samples output by the initial wake-up model. For each frame, the first posterior probability of each target articulation unit in the wake-up word is obtained from the posterior probabilities of all articulation units in the frame. Based on the comparison between each first posterior probability and a selection threshold, if the first negative sample is determined to be a negative sample for model fine-tuning, it is identified as a target negative sample. Finally, the initial wake-up model is fine-tuned based on each target negative sample to obtain the wake-up model. It can be seen that this invention selects target negative samples for model fine-tuning based on the comparison between the posterior probability output by the initial wake-up model and the selection threshold. This selection method is relatively objective, enabling the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the final fine-tuned wake-up model. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the training method for waking up a model in related technologies.

[0022] Figure 2 This is one of the flowcharts illustrating the training method for the wake-up model provided in this embodiment of the invention.

[0023] Figure 3 This is the second flowchart illustrating the training method for the wake-up model provided in this embodiment of the invention.

[0024] Figure 4 This is a flowchart of the training method for the wake-up model provided in this embodiment of the invention.

[0025] Figure 5 This is a flowchart illustrating the wake-up method provided in an embodiment of the present invention.

[0026] Figure 6 This is a schematic diagram of the structure of the training device for the wake-up model provided in an embodiment of the present invention.

[0027] Figure 7 This is a schematic diagram of the wake-up device provided in an embodiment of the present invention.

[0028] Figure 8 This is a schematic diagram of the physical structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0030] Figure 1 This is a flowchart illustrating the training method for wake-up models in related technologies, such as... Figure 1 As shown, the initial wake-up model is trained based on positive samples and basic negative samples to obtain the initial wake-up model. The initial wake-up model is then fine-tuned based on stacked negative samples to obtain the first-decision model and the second-decision model. The user's voice input is input into the first-decision model and the second-decision model. Wake-up is performed when the output of the first-decision model indicates that the user's voice input includes the wake-up word, and the output of the second-decision model also indicates that the user's voice input includes the wake-up word.

[0031] However, the aforementioned technologies rely on randomly stacked negative samples to fine-tune the initial wake-up model. Randomly stacked negative samples are subjective and cannot detect the model's weaknesses, thus reducing the accuracy of the final fine-tuned wake-up model.

[0032] Based on this, the present invention proposes a training method for a wake-up model. Based on the comparison results between the posterior probability output by the initial wake-up model and the screening threshold, target negative samples for model fine-tuning are selected. The screening method is relatively objective and enables the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the final fine-tuned wake-up model.

[0033] The following is combined with Figures 2 to 4 The present invention describes a training method for a wake-up model. The subject executing this wake-up model training method can be an electronic device such as a terminal, toy, in-vehicle device, robot, tablet computer, computer, or server, or it can be a wake-up model training device installed in the electronic device. This wake-up model training device can be implemented through software, hardware, or a combination of both.

[0034] Figure 2 This is one of the flowcharts illustrating the training method for the wake-up model provided in this embodiment of the invention, such as... Figure 2 As shown, the training method for this wake-up model includes the following steps: Step 201: Obtain at least one first negative sample, which does not include a wake word.

[0035] Exemplarily, the first negative sample can be determined based on audio that does not include the wake word and is randomly selected from a voice database. Here, not including the wake word means other audio that does not include any information in the wake word, or audio that includes a certain character or word in the wake word but does not include the entire wake word.

[0036] It should be noted that the number of the first negative samples can be determined based on the type of the modeling unit of the model. Different modeling units can use different numbers of the first negative samples, and the specific number can be determined based on empirical values. The present invention does not limit this.

[0037] Step 202: For each of the first negative samples, input the first negative sample into the initial wake-up model to obtain the posterior probability of each pronunciation unit in each frame of the first negative sample output by the initial wake-up model.

[0038] Exemplarily, when all the first negative samples are obtained, input the first negative samples into a pre-trained initial wake-up model. The initial wake-up model can also be called a baseline model to obtain the posterior probability of each pronunciation unit in each frame of the first negative sample output by the initial wake-up model. In the initial wake-up model based on character modeling, the pronunciation unit represents a character; in the initial wake-up model based on phoneme modeling, the pronunciation unit represents a phoneme; in the initial wake-up model based on untoned pinyin modeling, the pronunciation unit represents untoned pinyin. [[ID=']]

[0039] Step 203: For each of the frames, in the posterior probabilities of all the pronunciation units of the frame, obtain the first posterior probability of each target pronunciation unit in the wake word.

[0040] Exemplarily, when obtaining the posterior probability of each pronunciation unit in each frame of the first negative sample, for each frame, obtain the first posterior probability of each target pronunciation unit in the wake word in the posterior probabilities of all the pronunciation units of the frame. Taking the wake word "Hello, Xiaobao" as an example, and the modeling unit is untoned pinyin, then "Hello, Xiaobao" includes four target pronunciation units. It is necessary to obtain the first posterior probability of the target pronunciation unit of "ni", the first posterior probability of the target pronunciation unit of "hao", the first posterior probability of the target pronunciation unit of "xiao", and the first posterior probability of the target pronunciation unit of "bao" in the posterior probabilities of all the pronunciation units of the frame.

[0041] Step 204: When it is determined that the first negative sample is a negative sample participating in model fine-tuning based on the comparison result of each of the first posterior probabilities and the screening threshold, determine the first negative sample as the target negative sample.

[0042] For example, after obtaining the first posterior probability of each target pronunciation unit in the wake word, the first posterior probability of each target pronunciation unit in the wake word is compared with the screening threshold to obtain the comparison result. Then, based on the comparison result, it is determined whether the first negative sample needs to be used as a negative sample to participate in model fine-tuning. When it is determined that the first negative sample needs to be used as a negative sample to participate in model fine-tuning, the first negative sample is determined as the target negative sample.

[0043] It should be noted that when the wake word includes multiple target pronunciation units, each target pronunciation unit can correspond to a single filtering threshold, or each can correspond to a different filtering threshold. When each target pronunciation unit corresponds to a different filtering threshold, for each target pronunciation unit, it is necessary to compare the first posterior probability of that target pronunciation unit with the filtering threshold corresponding to that target pronunciation unit to obtain the comparison result. The specific filtering threshold can be set based on empirical values ​​or requirements, and this invention does not limit it in this regard.

[0044] Step 205: Fine-tune the initial wake-up model based on each of the target negative samples to obtain the wake-up model.

[0045] For example, after selecting target negative samples from all first negative samples, these target negative samples are added to the training set. The initial wake-up model is then fine-tuned based on the training set until the convergence condition is met, and the wake-up model is finally obtained. The convergence condition can be reaching a preset number of iterations, the wake-up rate being greater than a preset wake-up rate, the false wake-up rate being less than a preset value, etc.

[0046] The wake-up model training method provided by this invention involves inputting the first negative samples (not including wake-up words) into the initial wake-up model to obtain the posterior probability of each articulation unit in each frame of the first negative samples output by the initial wake-up model. For each frame, the first posterior probability of each target articulation unit in the wake-up word is obtained from the posterior probabilities of all articulation units in the frame. Based on the comparison between each first posterior probability and a selection threshold, if the first negative sample is determined to be a negative sample for model fine-tuning, it is identified as the target negative sample. Finally, the initial wake-up model is fine-tuned based on each target negative sample to obtain the wake-up model. It can be seen that this invention selects target negative samples for model fine-tuning based on the comparison between the posterior probability output by the initial wake-up model and the selection threshold. This selection method is relatively objective, enabling the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the final fine-tuned wake-up model. Furthermore, this invention does not require determining whether to wake up based on the outputs of the primary and secondary decision models, but only based on the trained wake-up model, thereby reducing equipment overhead and improving the stability of the wake-up system.

[0047] In one embodiment, the determination of the first negative sample as a negative sample participating in model fine-tuning based on the comparison results of each of the first posterior probabilities and the screening threshold in step 204 above can be achieved in the following way: For each target pronunciation unit, a second posterior probability greater than the filtering threshold is determined from the first posterior probabilities of the target pronunciation unit in all frames; the target position of the target frame corresponding to the second posterior probability is determined in the first negative sample; if the pronunciation of the target position in the first negative sample is different from the pronunciation of the target pronunciation unit, the first negative sample is determined to be a negative sample participating in model fine-tuning.

[0048] For example, for each target articulation unit in the wake word, find the second posterior probability greater than the filtering threshold from the first posterior probability of the target articulation unit in all frames. Then, obtain the target position of the target frame corresponding to the second posterior probability in the first negative sample. Compare the pronunciation of the target position in the first negative sample with the pronunciation of the target articulation unit in the wake word. If the pronunciation of the target position in the first negative sample is different from the pronunciation of the target articulation unit in the wake word, it means that the information in the wake word does not exist in the first negative sample, but the initial wake-up model will be falsely activated (erroneously triggered). Therefore, the first negative sample is determined as a negative sample to participate in model fine-tuning. If the pronunciation of the target position in the first negative sample is the same as the pronunciation of the target articulation unit in the wake word, it means that the information in the wake word exists in the first negative sample. It is not a false activation by the initial wake-up model, but a correct perception by the initial wake-up model. Therefore, the first negative sample is discarded and is not determined as a negative sample to participate in model fine-tuning.

[0049] It should be noted that for each target pronunciation unit in the wake word, the following steps are performed: among the first posterior probabilities of the target pronunciation units in all frames, determine the second posterior probability that is greater than the screening threshold; determine the target position of the target frame corresponding to the second posterior probability in the first negative sample; if the pronunciation of the target position in the first negative sample is different from the pronunciation of the target pronunciation unit, it is considered that the first negative sample does not match the target pronunciation unit in the wake word; after traversing all target pronunciation units in the wake word, if the first negative sample does not match at least one target pronunciation unit in the wake word, then the first negative sample is determined to be a negative sample participating in model fine-tuning.

[0050] In this embodiment, based on the comparison between the posterior probability and the screening threshold, and the verification of the pronunciation at the corresponding position, negative samples that participate in model fine-tuning are selected from all the first negative samples. This achieves objective screening of negative samples that participate in model fine-tuning. The negative samples obtained after screening are the most accurate negative sample data for suppressing false wake-ups, thereby improving the accuracy of the wake-up model finally trained.

[0051] In one embodiment, step 205 above fine-tunes the initial wake-up model based on each of the target negative samples to obtain a wake-up model, which can be implemented in the following way: For each target negative sample, if the first posterior probability of the target vocal unit of the target negative sample is greater than or equal to the wake-up threshold, a first weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample with the first weight set to obtain the wake-up model.

[0052] If the first posterior probability of the target vocal unit of the target negative sample is less than the wake-up threshold and greater than the screening threshold, a second weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample with the second weight set to obtain the wake-up model, wherein the second weight is less than the first weight.

[0053] For example, for each target negative sample, the first posterior probability of the target vocal unit of the target negative sample is compared with the wake-up threshold. If the first posterior probability of the target vocal unit of the target negative sample is greater than or equal to the wake-up threshold, it means that the target negative sample is more likely to be erroneously triggered by the initial wake-up model. Therefore, a larger first weight needs to be set for the target negative sample. Then, the target negative sample with the first weight is added to the training set. The initial wake-up model is fine-tuned based on the training set. Of course, the training set can also include some positive samples until the convergence condition is met, and finally the trained wake-up model is obtained.

[0054] If the first posterior probability of the target vocal unit of the target negative sample is less than the wake-up threshold but greater than the screening threshold, it means that the target negative sample will be erroneously triggered by the initial wake-up model, but the probability is small. Therefore, it is necessary to set a small second weight for the target negative sample and then add the target negative sample with the second weight to the training set. The initial wake-up model is then fine-tuned based on the training set until the convergence condition is met, and finally the trained wake-up model is obtained.

[0055] In this embodiment, when the first posterior probability of the target vocal unit of the target negative sample is greater than or equal to the wake-up threshold, a larger first weight is set for the target negative sample; when the first posterior probability of the target vocal unit of the target negative sample is less than the wake-up threshold but greater than the screening threshold, a smaller second weight is set for the target negative sample, so as to improve the false wake-up suppression effect of the trained wake-up model.

[0056] In one embodiment, step 201 above, which obtains at least one first negative sample, can be implemented in the following way: The at least one first negative sample is obtained based on the rules covering the equalization of the pronunciation units.

[0057] For example, when selecting the first negative sample, it can be based on the rule of covering balanced pronunciation units. All the first negative samples obtained in the end are combined into a negative sample set, so that all the pronunciations in the final negative sample set can cover all common pronunciations. Moreover, the frequency of use of all common pronunciations is distributed as evenly as possible. For example, the frequency of use of the pronunciation 'a' is higher than that of the pronunciation 'y'. This means that the number of first negative samples including the pronunciation 'a' should be greater than the number of first negative samples including the pronunciation 'y'.

[0058] In this embodiment, based on the rules covering the equalization pronunciation unit, at least one first negative sample is obtained, so that the coverage of the first negative sample is more extensive and accurate, and all pronunciations that generate false wake-ups are filtered out as much as possible, so as to further improve the accuracy of the wake-up model.

[0059] In one embodiment, Figure 3 This is a second schematic flowchart of the training method for the wake-up model provided in this embodiment of the invention, as shown below. Figure 3 As shown, prior to step 201 above, the training method for this wake-up model further includes the following steps: Step 301: Obtain a second negative sample and a positive sample, wherein the second negative sample does not include the wake word, and the positive sample includes the wake word.

[0060] For example, audio including the wake word can be recorded, and then scene simulation amplification can be performed on the audio including the wake word to obtain multiple audio samples containing the wake word. The spectral features corresponding to these audio samples are then identified as positive samples. The second negative sample can be determined based on audio samples randomly selected from a speech database that do not include the wake word. That is, the spectral features corresponding to the audio samples that do not include the wake word are identified as the second negative sample. Here, "not including the wake word" refers to other audio that does not include any information from the wake word, or audio that includes a word or phrase from the wake word but does not include the entire wake word. The second negative sample can be a different sample from the first negative sample.

[0061] It should be noted that the number of the second negative sample and the number of positive sample can be determined based on the type of the modeling unit of the model and the content of the wake word (e.g., a wake word composed of four completely different characters, a wake word composed of reduplicated words), and this invention does not limit this.

[0062] Step 302: Based on the second negative sample and the positive sample, train the model structure to obtain the initial wake-up model.

[0063] For example, when the second negative sample and positive sample are obtained, they are input into the model structure to obtain the first confidence score of the positive sample including the wake-up word and the second confidence score of the second negative sample including the wake-up word. A first loss is constructed based on the first confidence score of the positive sample including the wake-up word and the label of the positive sample including the wake-up word. A second loss is constructed based on the second confidence score of the second negative sample including the wake-up word and the label of the second negative sample not including the wake-up word. The model parameters of the model structure are adjusted based on the first loss and the second loss to finally obtain the initial wake-up model. The initial wake-up model meets the wake-up rate target in the test set (the wake-up rate is greater than the preset wake-up rate), but the false wake-up effect is poor. Therefore, it is necessary to fine-tune the initial wake-up model based on the selected target negative sample and some positive samples to finally obtain the wake-up model.

[0064] It should be noted that when the modeling unit of the model structure is greater than one, the number of parameters is large. The model structure can be pre-trained, and then fine-tuned based on the pre-trained model to obtain the initial wake-up model. This invention does not limit this.

[0065] In this embodiment, the model structure is pre-trained based on the second negative sample and positive sample to obtain an initial wake-up model. This facilitates subsequent fine-tuning of the initial wake-up model based on the selected target negative sample and some positive samples, ultimately obtaining the wake-up model and improving its accuracy.

[0066] Figure 4 This is a flowchart of the training method for the wake-up model provided in this embodiment of the invention, as follows: Figure 4 As shown, it includes the following steps: Step 401: Obtain the second negative sample and positive sample.

[0067] The second negative sample does not include the wake word, while the positive sample does include the wake word.

[0068] Step 402: Based on the second negative sample and the positive sample, train the model structure to obtain the initial wake-up model.

[0069] Step 403: Obtain at least one first negative sample.

[0070] The first negative sample did not include a wake word.

[0071] Step 404: For each of the first negative samples, input the first negative sample into the initial wake-up model to obtain the posterior probability of each vocal unit in each frame of the first negative sample output by the initial wake-up model.

[0072] Step 405: For each frame, obtain the first posterior probability of each target pronunciation unit in the wake word from the posterior probabilities of all pronunciation units in the frame.

[0073] Step 406: For each target articulation unit, determine a second posterior probability greater than the filtering threshold from the first posterior probabilities of the target articulation unit in all the frames.

[0074] Step 407: Determine the target position of the target frame corresponding to the second posterior probability in the first negative sample.

[0075] Step 408: If the pronunciation of the target position in the first negative sample is different from the pronunciation of the target pronunciation unit, determine the first negative sample as the target negative sample to participate in model fine-tuning.

[0076] Step 409: Fine-tune the initial wake-up model based on each of the target negative samples to obtain the wake-up model.

[0077] The wake-up model training method provided by this invention involves inputting the first negative samples (not including wake-up words) into the initial wake-up model to obtain the posterior probability of each articulation unit in each frame of the first negative samples output by the initial wake-up model. For each frame, the first posterior probability of each target articulation unit in the wake-up word is obtained from the posterior probabilities of all articulation units in the frame. Based on the comparison between each first posterior probability and a selection threshold, if the first negative sample is determined to be a negative sample for model fine-tuning, it is identified as the target negative sample. Finally, the initial wake-up model is fine-tuned based on each target negative sample to obtain the wake-up model. It can be seen that this invention selects target negative samples for model fine-tuning based on the comparison between the posterior probability output by the initial wake-up model and the selection threshold. This selection method is relatively objective, enabling the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the final fine-tuned wake-up model.

[0078] Figure 5 This is a flowchart illustrating the wake-up method provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the wake-up method includes the following steps: Step 501: Obtain the target speech.

[0079] For example, the target speech could be the speech input by the user through a microphone.

[0080] Step 502: Input the target speech into the wake-up model to obtain the recognition result output by the wake-up model.

[0081] For example, when the target speech input by the user is obtained, the target speech is input into the wake-up model to obtain the recognition result output by the wake-up model. This recognition result is used to characterize the confidence level of the target speech including the wake word.

[0082] Step 503: If the recognition result indicates that the target speech includes a wake-up word, then wake up the speech.

[0083] The wake-up model is trained based on any of the methods described above.

[0084] For example, the confidence level of the target speech containing the wake word in the recognition result is compared with a preset confidence level. If the confidence level of the target speech containing the wake word is greater than or equal to the preset confidence level, it is determined that the target speech contains the wake word, and the electronic device is woken up; if the confidence level of the target speech containing the wake word is less than the preset confidence level, it is determined that the target speech does not contain the wake word, and the electronic device is not woken up.

[0085] The wake-up method provided in this invention inputs the target speech into a wake-up model, obtains the recognition result output by the wake-up model, and performs wake-up when the recognition result indicates that the target speech includes a wake-up word. Since the target negative samples required for fine-tuning the wake-up model are obtained by comparing the posterior probability output by the initial wake-up model with a selection threshold, this selection method is relatively objective. It allows the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the target negative samples and further improving the accuracy of the wake-up model.

[0086] The training apparatus for the wake-up model provided by the present invention will be described below. The training apparatus for the wake-up model described below can be referred to in correspondence with the training method for the wake-up model described above.

[0087] Figure 6 This is a schematic diagram of the structure of the training device for the wake-up model provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the training device 600 for the wake-up model includes a first acquisition unit 601, a first recognition unit 602, a first determination unit 603, a second determination unit 604, and a fine-tuning unit 605; wherein: The first acquisition unit 601 is used to acquire at least one first negative sample, wherein the first negative sample does not include a wake word; The first recognition unit 602 is used to input the first negative sample into the initial wake-up model for each first negative sample, and obtain the posterior probability of each sound unit in each frame of the first negative sample output by the initial wake-up model. The first determining unit 603 is configured to, for each frame, obtain the first posterior probability of each target pronunciation unit in the wake word from the posterior probabilities of all pronunciation units in the frame. The second determining unit 604 is used to determine the first negative sample as the target negative sample when the first negative sample is determined to be a negative sample participating in model fine-tuning based on the comparison results of each of the first posterior probabilities and the screening threshold. The fine-tuning unit 605 is used to fine-tune the initial wake-up model based on each of the target negative samples to obtain the wake-up model.

[0088] The wake-up model training device provided by this invention, for each acquired first negative sample that does not include a wake-up word, inputs the first negative sample into the initial wake-up model, obtains the posterior probability of each articulation unit in each frame of the first negative sample output by the initial wake-up model, and for each frame, obtains the first posterior probability of each target articulation unit in the wake-up word from the posterior probabilities of all articulation units in the frame. Based on the comparison results of each first posterior probability and a screening threshold, if the first negative sample is determined to be a negative sample participating in model fine-tuning, it is identified as the target negative sample. Finally, the initial wake-up model is fine-tuned based on each target negative sample to obtain the wake-up model. It can be seen that this invention selects target negative samples for model fine-tuning based on the comparison results of the posterior probability output by the initial wake-up model and the screening threshold. The selection method is relatively objective, enabling the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the final fine-tuned wake-up model.

[0089] Based on any of the above embodiments, the second determining unit 604 is specifically used for: For each target articulation unit, a second posterior probability greater than the filtering threshold is determined from the first posterior probabilities of the target articulation unit in all the frames; Determine the target position of the target frame corresponding to the second posterior probability in the first negative sample; If the pronunciation of the target position in the first negative sample is different from the pronunciation of the target pronunciation unit, the first negative sample is determined to be a negative sample participating in model fine-tuning.

[0090] Based on any of the above embodiments, the fine-tuning unit 605 is specifically used for: For each target negative sample, if the first posterior probability of the target vocal unit of the target negative sample is greater than or equal to the wake-up threshold, a first weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample after setting the first weight to obtain the wake-up model; If the first posterior probability of the target vocal unit of the target negative sample is less than the wake-up threshold and greater than the screening threshold, a second weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample with the second weight set to obtain the wake-up model, wherein the second weight is less than the first weight.

[0091] Based on any of the above embodiments, the first acquisition unit 601 is specifically used for: The at least one first negative sample is obtained based on the rules covering the equalization of the pronunciation units.

[0092] Based on any of the above embodiments, the training device 600 for the wake-up model further includes: The third acquisition unit is used to acquire a second negative sample and a positive sample, wherein the second negative sample does not include the wake-up word, and the positive sample includes the wake-up word; The training unit is used to train the model structure based on the second negative sample and the positive sample to obtain the initial wake-up model.

[0093] Figure 7 This is a schematic diagram of the wake-up device provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the wake-up device 700 includes a second acquisition unit 701, a second identification unit 702, and a wake-up unit 703; wherein: The second acquisition unit 701 is used to acquire the target speech; The second recognition unit 702 is used to input the target speech into the wake-up model and obtain the recognition result output by the wake-up model; The wake-up unit 703 is used to wake up the target speech when the recognition result indicates that the target speech includes a wake-up word. The wake-up model is trained based on any of the methods described above.

[0094] The wake-up device provided in this embodiment of the invention inputs target speech into a wake-up model, obtains the recognition result output by the wake-up model, and performs wake-up when the recognition result indicates that the target speech includes a wake-up word. Since the target negative samples required for fine-tuning the wake-up model are obtained by comparing the posterior probability output by the initial wake-up model with a selection threshold, this selection method is relatively objective. It allows the initial wake-up model to perceive its own weaknesses, thereby improving the accuracy of the target negative samples and further improving the accuracy of the wake-up model.

[0095] Figure 8 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of the present invention, such as... Figure 8As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a training method for a wake-up model, the method including: acquiring at least one first negative sample, wherein the first negative sample does not include a wake-up word; For each of the first negative samples, the first negative sample is input into the initial wake-up model to obtain the posterior probability of each vocal unit in each frame of the first negative sample output by the initial wake-up model; For each frame, the first posterior probability of each target pronunciation unit in the wake word is obtained from the posterior probabilities of all pronunciation units in the frame. If, based on the comparison results between each of the first posterior probabilities and the screening threshold, the first negative sample is determined to be a negative sample participating in model fine-tuning, then the first negative sample is determined as the target negative sample. The initial wake-up model is fine-tuned based on each of the target negative samples to obtain the wake-up model.

[0096] Alternatively, you can use the following method: Acquire the target speech; The target speech is input into the wake-up model to obtain the recognition result output by the wake-up model; The wake-up is performed if the recognition result indicates that the target speech includes a wake-up word. The wake-up model is trained based on any of the methods described above.

[0097] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0098] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, the computer program being executed by a processor, the computer being able to execute the training method of the wake-up model provided by the above methods, the method including: acquiring at least one first negative sample, the first negative sample not including a wake-up word; For each of the first negative samples, the first negative sample is input into the initial wake-up model to obtain the posterior probability of each vocal unit in each frame of the first negative sample output by the initial wake-up model; For each frame, the first posterior probability of each target pronunciation unit in the wake word is obtained from the posterior probabilities of all pronunciation units in the frame. If, based on the comparison results between each of the first posterior probabilities and the screening threshold, the first negative sample is determined to be a negative sample participating in model fine-tuning, then the first negative sample is determined as the target negative sample. The initial wake-up model is fine-tuned based on each of the target negative samples to obtain the wake-up model.

[0099] Alternatively, when the program instructions are executed by the computer, the computer can implement the following method: Acquire the target speech; The target speech is input into the wake-up model to obtain the recognition result output by the wake-up model; The wake-up is performed if the recognition result indicates that the target speech includes a wake-up word. The wake-up model is trained based on any of the methods described above.

[0100] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for the wake-up model provided by the above methods, the method comprising: acquiring at least one first negative sample, wherein the first negative sample does not include a wake-up word; For each of the first negative samples, the first negative sample is input into the initial wake-up model to obtain the posterior probability of each vocal unit in each frame of the first negative sample output by the initial wake-up model; For each frame, the first posterior probability of each target pronunciation unit in the wake word is obtained from the posterior probabilities of all pronunciation units in the frame. If, based on the comparison results between each of the first posterior probabilities and the screening threshold, the first negative sample is determined to be a negative sample participating in model fine-tuning, then the first negative sample is determined as the target negative sample. The initial wake-up model is fine-tuned based on each of the target negative samples to obtain the wake-up model.

[0101] Alternatively, when the computer program is executed by the processor, it implements the following method: Acquire the target speech; The target speech is input into the wake-up model to obtain the recognition result output by the wake-up model; The wake-up is performed if the recognition result indicates that the target speech includes a wake-up word. The wake-up model is trained based on any of the methods described above.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A training method for a wake-up model, characterized in that, include: Obtain at least one first negative sample, wherein the first negative sample does not include a wake word; For each of the first negative samples, the first negative sample is input into the initial wake-up model to obtain the posterior probability of each vocal unit in each frame of the first negative sample output by the initial wake-up model; For each frame, the first posterior probability of each target pronunciation unit in the wake word is obtained from the posterior probabilities of all pronunciation units in the frame. If, based on the comparison results between each of the first posterior probabilities and the screening threshold, the first negative sample is determined to be a negative sample participating in model fine-tuning, then the first negative sample is determined as the target negative sample. The initial wake-up model is fine-tuned based on each of the target negative samples to obtain the wake-up model.

2. The training method for the wake-up model according to claim 1, characterized in that, Based on the comparison results of each of the first posterior probabilities and the screening threshold, the first negative sample is determined to be a negative sample participating in model fine-tuning, including: For each target articulation unit, a second posterior probability greater than the filtering threshold is determined from the first posterior probabilities of the target articulation unit in all the frames; Determine the target position of the target frame corresponding to the second posterior probability in the first negative sample; If the pronunciation of the target position in the first negative sample is different from the pronunciation of the target pronunciation unit, the first negative sample is determined to be a negative sample participating in model fine-tuning.

3. The training method for the wake-up model according to claim 1, characterized in that, The step of fine-tuning the initial wake-up model based on each of the target negative samples to obtain the wake-up model includes: For each target negative sample, if the first posterior probability of the target vocal unit of the target negative sample is greater than or equal to the wake-up threshold, a first weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample after setting the first weight to obtain the wake-up model; If the first posterior probability of the target vocal unit of the target negative sample is less than the wake-up threshold and greater than the screening threshold, a second weight is set for the target negative sample, and the initial wake-up model is fine-tuned based on the target negative sample with the second weight set to obtain the wake-up model, wherein the second weight is less than the first weight.

4. The training method for the wake-up model according to claim 1, characterized in that, The acquisition of at least one first negative sample includes: The at least one first negative sample is obtained based on the rules covering the equalization of the pronunciation units.

5. The training method for the wake-up model according to any one of claims 1-4, characterized in that, The method further includes: Obtain a second negative sample and a positive sample, wherein the second negative sample does not include the wake word, and the positive sample includes the wake word; Based on the second negative sample and the positive sample, the model structure is trained to obtain the initial wake-up model.

6. A wake-up method, characterized in that, include: Acquire the target speech; The target speech is input into the wake-up model to obtain the recognition result output by the wake-up model; The wake-up is performed if the recognition result indicates that the target speech includes a wake-up word. The wake-up model is trained based on the method described in any one of claims 1-5.

7. A training device for a wake-up model, characterized in that, include: The first acquisition unit is used to acquire at least one first negative sample, wherein the first negative sample does not include a wake word; The first recognition unit is used to input the first negative sample into the initial wake-up model for each first negative sample, and obtain the posterior probability of each sound unit in each frame of the first negative sample output by the initial wake-up model. The first determining unit is configured to, for each frame, obtain the first posterior probability of each target pronunciation unit in the wake word from the posterior probabilities of all pronunciation units in the frame; The second determining unit is used to determine the first negative sample as the target negative sample when the first negative sample is determined to be a negative sample participating in model fine-tuning based on the comparison results of each of the first posterior probabilities and the screening threshold. The fine-tuning unit is used to fine-tune the initial wake-up model based on each of the target negative samples to obtain the wake-up model.

8. A wake-up device, characterized in that, include: The second acquisition unit is used to acquire the target speech; The second recognition unit is used to input the target speech into the wake-up model and obtain the recognition result output by the wake-up model; A wake-up unit is used to wake up the target speech if the recognition result indicates that a wake-up word is included in the target speech; The wake-up model is trained based on the method described in any one of claims 1-5.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the training method of the wake-up model as described in any one of claims 1 to 5, or implements the wake-up method as described in claim 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the training method of the wake-up model as described in any one of claims 1 to 5, or implements the wake-up method as described in claim 6.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the wake-up model as described in any one of claims 1 to 5, or implements the wake-up method as described in claim 6.