Wake-up processing method, device, facility, and computer storage medium
The wake-up processing method separately trains multiple wake-up words using a single model to avoid crosstalk, enhancing recognition accuracy and reducing false wake-up events in voice devices.
Patent Information
- Application Number
- JP2024531560
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-06
- Filing Date
- 2022-03-23
- Publication Date
- 2025-09-24
- Estimated Expiration
- 2042-03-23
AI Technical Summary
Existing voice recognition technologies face issues with false wake-up rates due to crosstalk between different wake-up words, which are often trained simultaneously, leading to misrecognition and increased false wake-up events.
A wake-up processing method that uses a single wake-up model to separately train multiple wake-up words, obtaining confidence levels and thresholds for each word, and triggers events based on these comparisons to avoid crosstalk and improve recognition accuracy.
This approach reduces false wake-up rates and improves development efficiency by preventing interference between wake-up words, optimizing user experience and reducing processing workload.
Smart Images

Figure 0007743630000002 
Figure 0007743630000003 
Figure 0007743630000004
Abstract
Description
[Technical Field]
[0001] This disclosure claims priority to a Chinese patent application filed on August 6, 2021, bearing application number "202110904169.X" and entitled "Wake-up processing method, device, equipment and computer storage medium," the entire contents of which are incorporated herein by reference. The present disclosure relates to the field of voice recognition technology, and in particular to a wake-up processing method, device, equipment, and computer storage medium. [Background technology]
[0002] With the development of voice recognition technology, smart homes are becoming a trend, and voice devices are becoming more and more prevalent in people's daily lives. Currently, many users generally have a variety of voice devices in their homes, and before voice control of the voice devices, they need to perform a wake-up operation for the voice devices.
[0003] However, in related technologies, these voice devices may generally need to recognize multiple wake-up words, and these different wake-up words are trained simultaneously, which makes it easy for crosstalk to occur between different wake-up words, further causing false wake-up problems and increasing the false wake-up rate of voice recognition. Summary of the Invention
[0004] The present disclosure aims to provide a wake-up processing method, device, equipment, and computer storage medium that can avoid the possibility of wake-up word crosstalk occurring when different wake-up words are trained simultaneously, and reduce the false wake-up rate of audio equipment.
[0005] To achieve the above objectives, the technical solution of the present disclosure is realized as follows:
[0006] According to a first aspect, an embodiment of the present disclosure is a wake-up processing method applied to an audio device, comprising: obtaining audio to be recognized; processing each of the audios to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds, wherein the at least two sets of training data are obtained by training at least two sets of wake-up word training sets through the wake-up model, respectively; and triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and their corresponding confidence thresholds.
[0007] In some embodiments, the step of obtaining audio to be recognized comprises: collecting data using a sound collecting device to obtain initial sound data; and pre-processing the initial speech data to obtain the audio to be recognized.
[0008] In some embodiments, each set of training data includes model parameters and a confidence threshold, and the step of processing the audio to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence levels and corresponding confidence thresholds includes: The method includes a step of processing each of the audio to be recognized using the wake-up model and model parameters in the at least two sets of training data to obtain at least two confidence levels, and obtaining confidence level thresholds corresponding to each of the at least two confidence levels from the at least two sets of training data.
[0009] In some embodiments, the at least two sets of training data include a first set of training data including first model parameters and a first confidence threshold, and a second set of training data including second model parameters and a second confidence threshold; The step of processing each of the audios to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds includes: processing the audio to be recognized using the wake-up model and the first model parameters in the first set of training data to obtain a first confidence level; and determining the first confidence level threshold corresponding to the first confidence level from the first set of training data; processing the audio to be recognized using the wake-up model and the second model parameters in the second set of training data to obtain a second confidence level; and determining the second confidence level threshold corresponding to the second confidence level from the second set of training data.
[0010] In some embodiments, triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold comprises: If the first reliability is greater than or equal to the first reliability threshold, or if the second reliability is greater than or equal to the second reliability threshold, triggering a wake-up event of the audio facility.
[0011] In some embodiments, the wake-up event includes a first wake-up event and / or a second wake-up event, the first wake-up event having an association with a wake-up word corresponding to the first set of training data, and the second wake-up event having an association with a wake-up word corresponding to the second set of training data.
[0012] In some embodiments, triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold comprises: Triggering the first wake-up event of the audio equipment when the first reliability is greater than or equal to the first reliability threshold and the second reliability is less than the second reliability threshold; or triggering the second wake-up event of the audio equipment when the second reliability is greater than or equal to the second reliability threshold and the first reliability is less than the first reliability threshold; or If the first reliability is greater than or equal to the first reliability threshold and the second reliability is greater than or equal to the second reliability threshold, a first value by which the first reliability exceeds the first reliability threshold and a second value by which the second reliability exceeds the second reliability threshold are calculated, and a target wake-up event of the audio equipment is triggered by the first value and the second value.
[0013] In some embodiments, the step of triggering a target wake-up event of the audio facility according to the first value and the second value comprises: If the first value is greater than or equal to the second value, the target wake-up event is determined to be the first wake-up event and is triggered; or If the first value is smaller than the second value, the target wake-up event is determined to be the second wake-up event and is triggered.
[0014] In some embodiments, the method further comprises: obtaining the at least two wake-up word training sets; and training the wake-up model using the at least two wake-up word training sets to obtain the at least two sets of training data, each set of training data including model parameters and a confidence threshold.
[0015] In some embodiments, the step of obtaining at least two wake-up word training sets comprises: obtaining an initial training set including at least two wake-up words; and grouping the initial training set according to different wake-up words to obtain the at least two wake-up word training sets.
[0016] According to a second aspect, an embodiment of the present disclosure is a wake-up processing device applied to an audio device, the wake-up processing device comprising: an acquisition unit arranged to acquire audio awaiting recognition; a processing unit for processing each of the audios to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds, the at least two sets of training data being arranged such that at least two sets of wake-up word training sets are respectively trained through the wake-up model; a trigger unit arranged to trigger a wake-up event of the audio equipment based on a comparison result between the at least two reliabilities and their corresponding reliability thresholds.
[0017] According to a third aspect, an embodiment of the present disclosure provides an audio device including a memory and a processor, the memory is used to store a computer program executable by the processor; The processor, during execution of the computer program, provides audio facilities used to carry out the method of any one of the first aspects.
[0018] According to a fourth aspect, an embodiment of the present disclosure provides a computer storage medium storing a computer program that, when executed by at least one processor, implements the method according to any one of the first aspects.
[0019] The embodiments of the present disclosure provide a wake-up processing method, apparatus, device, and computer storage medium, which includes: receiving audio to be recognized; respectively processing the audio to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence levels and corresponding confidence thresholds; the at least two sets of training data are obtained by respectively training at least two sets of wake-up word training sets using the wake-up model; and determining a device to be woken up based on a comparison result between the at least two confidence levels and the corresponding confidence thresholds. In this way, by separately training multiple wake-up words using the same wake-up model, it is possible to avoid the possibility of crosstalk with other wake-up words and achieve better recognition results with a smaller training amount. It is also possible to separate the training data for wake-up words so that they do not interfere with each other, which improves development efficiency and also reduces the false wake-up rate of the voice device when recognizing multiple wake-up words simultaneously. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a schematic flow diagram of a wake-up processing method provided by an embodiment of the present disclosure. [Figure 2] FIG. 10 is a schematic flow diagram of another wake-up processing method provided by an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram of a training process for a wake-up model provided by an embodiment of the present disclosure. [Figure 4] FIG. 2 is a detailed schematic flow chart of a wake-up processing method provided by an embodiment of the present disclosure. [Figure 5] FIG. 2 is a schematic diagram illustrating the configuration of a wake-up processing device provided by an embodiment of the present disclosure. [Figure 6] FIG. 2 is a schematic diagram of a specific hardware structure of the audio equipment provided by the embodiments of the present disclosure; DETAILED DESCRIPTION OF THE INVENTION
[0021] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. It should be understood that the specific embodiments described herein are only for the purpose of illustrating the relevant disclosure, and are not intended to limit the disclosure. In addition, for ease of description, only the relevant parts of the disclosure are shown.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. The terms used herein are used only for the purpose of describing examples of this disclosure and are not intended to limit this disclosure.
[0023] In the following description, "some embodiments" are described that describe a subset of all possible embodiments, but it should be understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other where there is no contradiction.
[0024] It should be noted that the terms "first," "second," and "third" used in connection with the embodiments of the present disclosure are used only to distinguish between similar objects and do not represent a particular ordering of the objects. It should be understood that "first," "second," and "third" can be interchanged with respect to a particular order or order where permissible, to allow the embodiments of the present disclosure described herein to be performed in an order other than that illustrated or described herein.
[0025] In practical applications, currently available wake-up model speech recognition methods are divided into two types: (1) a type in which multiple wake-up words are trained using the same model, and (2) a type in which multiple models are used to train different wake-up words.
[0026] However, in technical solutions such as (1), multiple wake-up words are trained using the same model, which makes crosstalk between different wake-up words more likely to occur due to their similarity. Considering issues of wake-up response speed and storage space, the training set must not be too large, which makes it easy for sounds between different wake-up words to be misrecognized and false wake-up complaints to occur. In technical solutions such as (2), model loading takes time, and switching delays are significant, making it impossible to achieve solutions that recognize multiple wake-up words simultaneously. In short, in related technologies, these voice devices may need to recognize multiple wake-up words, and these different wake-up words are trained simultaneously, which makes it easy for crosstalk to occur between different wake-up words, further increasing false wake-up issues and increasing the false wake-up rate of voice recognition.
[0027] Thus, an embodiment of the present disclosure provides a wake-up processing method, the basic idea of which is to obtain audio to be recognized, and process the audio to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence levels and corresponding confidence thresholds, the at least two sets of training data being obtained by respectively training at least two sets of wake-up word training sets using the wake-up model, and determine a device to be woken up based on a comparison result between the at least two confidence levels and the corresponding confidence thresholds. In this way, by using the same wake-up model to train multiple wake-up words separately, it is possible to avoid the possibility of crosstalk with other wake-up words and achieve better recognition results with a smaller training amount. It is also possible to separate the training data for wake-up words so that they do not interfere with each other, thereby improving development efficiency and reducing the false wake-up rate of the voice device when recognizing multiple wake-up words simultaneously.
[0028] Hereinafter, each embodiment of the present disclosure will be described in detail in conjunction with the drawings.
[0029] Example 1 1, which shows a schematic flow diagram of a wake-up processing method provided by an embodiment of the present disclosure. As shown in FIG. 1, the method may include steps S101 to S103.
[0030] In step S101, audio waiting to be recognized is acquired.
[0031] The wake-up processing method according to the embodiment of the present disclosure is applied to a wake-up processing device or a voice device integrated with the wake-up processing device, where the voice device can have a voice dialogue with a user and can be any device that requires voice wake-up, such as any common home appliance, including a voice air conditioner, a voice water heater, a voice rice cooker, or a voice microwave oven, but is not limited thereto.
[0032] In addition, since the audio equipment can have a voice dialogue with the user, data collection can also be performed by a sound collection device at that time. Therefore, in some embodiments, in step S101, the step of obtaining the audio to be recognized can be collecting data using a sound collecting device to obtain initial sound data; and pre-processing the initial speech data to obtain the audio to be recognized.
[0033] In the embodiments of the present disclosure, the sound collecting device may be an audio collector such as a microphone, a microphone, etc. Specifically, real-time data collection by the microphone can obtain initial voice data from a user, and then the initial voice data can be pre-processed to obtain audio to be recognized.
[0034] It should be noted that the initial voice data in the embodiments of the present disclosure can include user sound information, but if it is only environmental sound, it is not related to wake-up recognition and is outside the scope of the discussion of the present embodiment, so no further mention will be made here. That is, the initial voice data can be uttered by the user, for example, "Mi-chan Mi-chan." After the voice equipment acquires the sound information, it preprocesses the information.
[0035] Here, the pre-processing can include two aspects: an end point detection process and a pre-emphasis process, which will be described in detail below.
[0036] In one possible embodiment, the end point detection process means finding the start and end points of the command audio, and intercepting a sound segment of several consecutive frames from the sound information, and setting the preceding frames as audio to be recognized according to the order of the sound segments. Specifically, the number of frames to be set as audio to be recognized can be determined based on the length of the set wake-up word. For example, a specific time length can be preset according to the number of characters in the wake-up word, and sound segments within this time length can be determined as audio to be recognized. The specific time length can be adjusted according to actual circumstances, and is not limited to this embodiment.
[0037] Alternatively, the number of frames of audio waiting to be recognized can be determined based on the length of time that null data is detected between two consecutive sound segments. For example, in actual use, if a user first invokes a wake-up word, and then invokes the remaining voice command after a pause of a few seconds, the segment before the null data can be the audio waiting to be recognized.
[0038] For example, taking a voice-activated air conditioner as an example, in relation to the above-mentioned embodiment, the voice-activated air conditioner receives an audio segment of "Mi-chan Mi-chan" through a sound collecting device, and the preset time length of the wake-up word "Mi-chan Mi-chan" is 2 seconds. In the endpoint detection process, it is necessary to intercept the number of frames corresponding to the previous time length of 2 seconds as the audio to be recognized. Alternatively, if there is a blank section between two sentences in the audio segment of "Mi-chan Mi-chan, turn up the temperature" received by the voice-activated air conditioner through a sound collecting device, and the audio information in the blank section is null data, the number of frames before the null data in this blank section can be set as the audio to be recognized.
[0039] In another possible embodiment, the pre-emphasis process means emphasizing the high frequency part of the audio and increasing the high frequency resolution, and after the sound information is obtained, using an audio recognition method to extract environmental sound information and audio information from the sound information, remove noise interference, increase the high frequency resolution, and obtain clear human sound information.
[0040] In addition, the embodiments of the present disclosure can also use the audio to be recognized that contains environmental sound to train a wake-up model with noise. Specifically, environmental sound information can be extracted from the audio to be recognized and then sent to a server as training data. The sound pressure level of the environmental sound information can be used as a feature parameter in further training the wake-up model. A noisy training method can be applied to the wake-up model. The recognition process of the wake-up model can adjust corresponding parameters according to different sizes of environmental sound information, such as adjusting a corresponding confidence threshold, so that the wake-up model can be adapted to different usage scenarios.
[0041] In step S102, the audio to be recognized is processed using a wake-up model and at least two sets of training data, respectively, to obtain at least two confidence measures and corresponding confidence thresholds.
[0042] In an embodiment of the present disclosure, the at least two sets of training data are obtained by respectively training at least two sets of wake-up word training sets through a wake-up model, where each set of training data may include model parameters and a confidence threshold (wherein the confidence threshold is also referred to as a "wake-up threshold").
[0043] Correspondingly, in some embodiments, in step S102, the step of processing the audio to be recognized using a wake-up model and at least two sets of training data, respectively, to obtain at least two confidence measures and corresponding confidence thresholds, includes: The method may include processing each of the audio to be recognized using the wake-up model and model parameters in the at least two sets of training data to obtain at least two confidence levels, and obtaining confidence level thresholds corresponding to each of the at least two confidence levels from the at least two sets of training data.
[0044] The at least two wake-up word training sets are obtained by grouping based on different wake-up words, i.e., each wake-up word corresponds to one wake-up word training set, and the at least two sets of training data are obtained by training the at least two wake-up word training sets using a wake-up model, respectively, i.e., there is a correspondence between the training data and the wake-up words, and the at least two sets of training data each correspond to one wake-up word. For example, assuming that wake-up word A and wake-up word B exist, one wake-up word A training set and one wake-up word B training set are obtained, and training data for wake-up word A and training data for wake-up word B are obtained through training using the wake-up model.
[0045] In this way, the process of processing the audio to be recognized recognizes the audio to be recognized using training data corresponding to at least two wake-up words, respectively, and can obtain the confidence that the audio to be recognized will be obtained under the model parameters of each set of training data. Unlike related art in which different wake-up words are trained simultaneously, in the embodiments of the present disclosure, both the recognition process and the training process realize separate processing using different wake-up words, preventing crosstalk between different wake-up words in the recognition results and reducing the false wake-up rate during use. Furthermore, by separately training and recognizing wake-up words in this way, the workload of the processor is greatly reduced, the response time is also reduced, and the user experience is optimized.
[0046] Illustratively, taking a voice air conditioner as an example, assuming that the audio received by the voice air conditioner waiting to be recognized is "Mi-chan Mi-chan", the wake-up model will combine with training data corresponding to at least two wake-up words (e.g., "Mi-chan Mi-chan" and "Mi-chan, hello") to determine the confidence corresponding to each wake-up word, and finally obtain the confidence corresponding to each of the two wake-up words "Mi-chan Mi-chan" and "Mi-chan, hello".
[0047] Although a voice recognition module may be built into the voice equipment to recognize the audio waiting to be recognized, it is also possible to connect the voice equipment to a server, perform voice recognition via the server, and feed back the specific results to the voice equipment for use as input, thereby preventing wake-up word crosstalk between multiple voice equipment, and it is understood that the specific method can be adjusted according to actual circumstances. It is also understood that the wake-up word may be any predetermined character, and is not limited to this embodiment.
[0048] In addition, for the audio to be recognized, the embodiment of the present disclosure can also perform a text conversion process on the audio to be recognized to obtain audio-text information before processing it using a wake-up model and at least two sets of training data, and then match the audio-text information using a character matching or semantic matching method to determine at least one keyword or key phrase, and then process it using a wake-up model and at least two sets of training data, which is not described here.
[0049] Additionally, in an embodiment of the present disclosure, the wake-up model and reliability threshold are pre-set in the audio equipment by factory settings, and there is an initial wake-up model and reliability threshold when the audio equipment is first powered on and used, which can be trained and updated during subsequent use to better suit the user's usage scenario, again without any limitation here.
[0050] Furthermore, assuming that the at least two sets of training data herein are two sets, the at least two sets of training data include a first set of training data including first model parameters and a first confidence threshold, and a second set of training data including second model parameters and a second confidence threshold.
[0051] Correspondingly, in some embodiments, the step of processing the audio to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds may include: processing the audio to be recognized using the wake-up model and the first model parameters in the first set of training data to obtain a first confidence level; and determining the first confidence level threshold corresponding to the first confidence level from the first set of training data; processing the audio to be recognized using the wake-up model and the second model parameters in the second set of training data to obtain a second confidence level; and determining the second confidence level threshold corresponding to the second confidence level from the second set of training data.
[0052] That is, assuming that wake-up word A and wake-up word B exist, after obtaining training data for wake-up word A and training data for wake-up word B, the wake-up model and the training data for wake-up word A can be used to process the audio to be recognized to obtain the confidence level and corresponding confidence threshold for wake-up word A, and the wake-up model and the training data for wake-up word B can be used to process the audio to be recognized to obtain the confidence level and corresponding confidence threshold for wake-up word B, thereby obtaining the confidence levels and confidence thresholds corresponding to these two sets of training data, which can then be compared to determine the wake-up event to be triggered.
[0053] In step S103, a wake-up event of the audio equipment is triggered based on a comparison result between the at least two confidence levels and the corresponding confidence level thresholds.
[0054] In addition, after at least two confidence levels and their corresponding confidence thresholds are obtained, these at least two confidence levels may be compared with their corresponding confidence thresholds, and then a wake-up event for the audio equipment is triggered based on the comparison result.
[0055] Specifically, taking the case of two wake-up words as an example, the at least two confidence levels only include a first confidence level and a second confidence level. In some embodiments, in step S103, triggering a wake-up event of the audio device based on a comparison result between the at least two confidence levels and the corresponding confidence level thresholds includes: The method may include triggering a wake-up event of the audio equipment when the first reliability is equal to or greater than the first reliability threshold, or when the second reliability is equal to or greater than the second reliability threshold.
[0056] In an embodiment of the present disclosure, the wake-up event may include a first wake-up event and / or a second wake-up event, wherein the first wake-up event has an association relationship with a wake-up word corresponding to the first set of training data, and the second wake-up event has an association relationship with a wake-up word corresponding to the second set of training data.
[0057] In some embodiments, when the at least two confidence levels include a first confidence level and a second confidence level, triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold value includes: Triggering the first wake-up event of the audio equipment when the first reliability is greater than or equal to the first reliability threshold and the second reliability is less than the second reliability threshold; or triggering the second wake-up event of the audio equipment when the second reliability is greater than or equal to the second reliability threshold and the first reliability is less than the first reliability threshold; or The method may further include the step of: if the first reliability is greater than or equal to the first reliability threshold and the second reliability is greater than or equal to the second reliability threshold, calculating a first value by which the first reliability exceeds the first reliability threshold and a second value by which the second reliability exceeds the second reliability threshold, and triggering a target wake-up event of the audio equipment according to the first value and the second value.
[0058] If both of the two reliabilities are equal to or greater than the corresponding reliability threshold, then it is necessary to calculate a first value where the first reliability exceeds the first reliability threshold and a second value where the second reliability exceeds the second reliability threshold. In some embodiments, the step of triggering a target wake-up event of the audio equipment according to the first value and the second value includes: If the first value is greater than or equal to the second value, the target wake-up event is determined to be the first wake-up event and is triggered; or If the first value is smaller than the second value, the target wake-up event is determined to be the second wake-up event and is triggered.
[0059] For example, when a sound collection device of equipment waiting to be woken up receives audio waiting to be recognized with the content "Mi-san," the wake-up model obtains a confidence level corresponding to "Mi-chan Mi-chan" and a confidence level corresponding to "Mi-san" based on the training data corresponding to the wake-up word "Mi-chan Mi-chan" and the training data corresponding to the wake-up word "Mi-san," respectively, compares the two confidence levels with the corresponding confidence thresholds, and if the confidence level of "Mi-chan Mi-chan" is equal to or greater than the corresponding confidence threshold, the wake-up event corresponding to "Mi-chan Mi-chan" is the target wake-up event; otherwise, if the confidence level of "Mi-san" is equal to or greater than the corresponding confidence threshold, the wake-up event corresponding to "Mi-san" is the target wake-up event.
[0060] In addition, under special circumstances, if the reliability of "Mr. Mi, Mr. Mi" and "Ms. Mi" are both equal to or greater than the corresponding reliability threshold, the amount by which the reliability exceeds the reliability threshold can be compared, and the wake-up event corresponding to the wake-up word whose reliability exceeds the reliability threshold by a larger amount among the two wake-up words can be determined as the target wake-up event. In this way, after the target wake-up event is determined, the audio device can be made to execute the target wake-up event to perform the corresponding wake-up operation.
[0061] It should be noted that different wake-up words, for example, wake-up words with different pronunciations but the same meaning, can generate the same wake-up command, where the same wake-up command waking up the corresponding wake-up event can be applied to the wake-up process of different wake-up words in a single audio device, or to a cascaded audio central control system consisting of multiple audio devices. The embodiments of the present disclosure can be selected as needed according to the situation, and are not limited thereto.
[0062] An embodiment of the present disclosure provides a voice processing method applicable to a voice equipment. The method includes: receiving audio to be recognized; using a wake-up model and at least two sets of training data to process the audio to be recognized; obtaining at least two confidence levels and corresponding confidence thresholds; the at least two sets of training data are obtained by training at least two sets of wake-up word training sets using the wake-up model; and triggering a wake-up event for the voice equipment based on a comparison result of the at least two confidence levels and the corresponding confidence thresholds. In this way, by separately training multiple wake-up words using the same wake-up model, it is possible to avoid crosstalk with other wake-up words and achieve better recognition results with a smaller training amount. Furthermore, it is possible to separate the training data for wake-up words so that they do not interfere with each other, thereby improving development efficiency and reducing the false wake-up rate of the voice equipment when multiple wake-up words are recognized simultaneously.
[0063] Example 2 Based on the same inventive idea as the above-mentioned embodiment, referring to Figure 2, a schematic flow diagram of another wake-up processing method provided by the embodiment of the present disclosure is shown. As shown in Figure 2, the method includes: Step S201: obtaining an initial training set including at least two wake-up words; Step S202: grouping the initial training set according to different wake-up words to obtain the at least two wake-up word training sets; Step S203 includes training the wake-up model using the at least two sets of wake-up word training sets to obtain the at least two sets of training data.
[0064] In the embodiments of the present disclosure, the wake-up model may be a neural network model. A neural network (NN) is a complex network system formed by extensively interconnecting a large number of simple processing units (called "neurons"), which reflects many basic characteristics of human brain function and is a highly complex nonlinear dynamic learning system. Neural networks have the capabilities of massive parallelism, distributed memory and processing, self-organization, self-adaptation, and self-learning, and are particularly suitable for dealing with imprecise and ambiguous information processing problems that require simultaneous consideration of multiple factors and conditions. Here, the wake-up model may be a deep neural network (DNN) model. Specifically, the wake-up model here may include a DNN structural design and a mathematical model of each neuron.
[0065] In the embodiment of the present disclosure, each set of training data may include at least model parameters and a confidence threshold. Specifically, the training data may include optimal parameters (abbreviated as "model parameters") obtained after training in the DNN and a confidence threshold.
[0066] It should also be noted that the embodiments of the present disclosure may use wake-up models in which multiple wake-up words are separately trained to obtain corresponding training data, thereby realizing a method in which the wake-up word data are separated and do not interfere with each other. Furthermore, since multi-model systems use multiple models for separate training, it takes time to load the wake-up models in the later use process, which results in serious recognition delays due to the switching process. However, the technical solution provided by the embodiments of the present disclosure is different from the multi-model solution in that different wake-up words are used to train through the same wake-up model, thereby reducing the problem of delays during the wake-up process.
[0067] It should also be noted that by dividing the training sets by grouping them based on different wake-up words, training the training sets corresponding to different wake-up words separately, and storing the obtained data independently, it is possible to train a model with a limited training set, thereby achieving the technical effect of no crosstalk occurring between wake-up words, and avoiding crosstalk between wake-up words, unlike in the case of a model that trains wake-up words simultaneously to wake up.
[0068] Furthermore, if a new wake-up word needs to be added, in some embodiments, the method trains the wake-up model according to a set of wake-up word training sets corresponding to the new wake-up words to obtain a new set of training data.
[0069] In other words, since the embodiments of the present disclosure realize data separation, retraining is only required for newly added wake-up words, and does not affect existing wake-up words, thereby improving development efficiency even when new wake-up words are added.
[0070] In other words, for a wake-up model that is already in use, if new wake-up words need to be added, the existing model can be trained using new wake-up words according to the training method of the above embodiment, where the already-used wake-up model is the wake-up model in the above technical solution, and the wake-up model is continuously trained, and the wake-up model continuously learns new wake-up words, so that the product can be continuously updated to meet new user needs.
[0071] For example, taking wake-up word A and wake-up word B as an example, refer to Figure 3, which is a schematic diagram of the training process of the wake-up model provided by the embodiment of the present disclosure. In Figure 3, there are two sets of wake-up word training sets, for example, a wake-up word A training set and a wake-up word B training set. The wake-up model can be trained using the wake-up word A training set to obtain training data for wake-up word A, and the wake-up model can be trained using the wake-up word B training set to obtain training data for wake-up word B.
[0072] Specifically, in the embodiment of the present disclosure, two wake-up word training sets are used to train the same wake-up model, thereby obtaining two sets of training data. Here, the input initial training set is divided into different groups according to different wake-up words, and each set of wake-up word training sets is used sequentially as input data to train the wake-up model until training of all wake-up word training sets is completed. Note that each set of wake-up word training sets is stored by separating the training data obtained after training the wake-up model by different wake-up words.
[0073] For example, the wake-up word training set is divided into groups of different wake-up words, where the different wake-up words may be wake-up words with completely different meanings or wake-up words with the same meaning but in different dialects, such as the Cantonese wake-up words "xiu mei xiu mei" and "xiang mei xiang mei." Thus, when using a wake-up model for recognition processing, multiple pieces of input information can be input to obtain one output piece of information. Here, the input information is audio to be recognized. The wake-up model can have a built-in voice recognition module that recognizes the audio to be recognized and outputs a corresponding wake-up event. Alternatively, a voice recognition module can be installed in the audio equipment to obtain the audio to be recognized from the sound information, perform voice recognition on the audio to be recognized, and output a corresponding wake-up event. In the embodiments of the present disclosure, the specific manner of inputting information can be selected according to actual needs and is not limited in any way.
[0074] An embodiment of the present disclosure provides a voice processing method applied to a voice system, which includes obtaining at least two wake-up word training sets, training the wake-up model using the at least two wake-up word training sets, and obtaining at least two sets of training data, each set of training data including model parameters and a confidence threshold. In this way, it is possible to avoid the possibility of wake-up word crosstalk occurring when different wake-up words are trained simultaneously, achieve mutual separation of wake-up words and mutual separation of training data, and reduce the false wake-up rate of the voice system when multiple wake-up words are recognized simultaneously.
[0075] Example 3 Based on the same inventive idea as the above-mentioned embodiment, referring to Figure 4, a detailed schematic flow diagram of the wake-up processing method provided by the embodiment of the present disclosure is shown. Take the example of having wake-up word A and wake-up word B, as shown in Figure 4, this method includes: Step S401: a microphone collects audio in real time, and obtains audio to be recognized through front-end pre-processing; step S402, processing the audio to be recognized using training data of a wake-up model and a wake-up word A to obtain a confidence A and a corresponding confidence threshold A; step S403, processing the audio to be recognized using training data of a wake-up model and a wake-up word B to obtain a confidence B and a corresponding confidence threshold B; step S404, determining whether reliability A≧reliability threshold A or reliability B≧reliability threshold B; Step S405: if the determination result is YES, triggering a wake-up event for the audio equipment.
[0076] If the determination result in step S404 is YES, step S405 may be executed, and after waking up the audio equipment, the process may return to step S401 to continue collecting the next audio; if the determination result is NO, the process may return to step S401 to continue collecting the next audio.
[0077] In the embodiments of the present disclosure, a single wake-up model is adopted, and different wake-up words are trained separately to obtain independent training data, thereby realizing that the training data of the wake-up words are separated and do not interfere with each other.
[0078] In one possible embodiment, the process involved is as follows:
[0079] (1) Multiple wake-up word designs use the same wake-up model.
[0080] (2) The wake-up model is trained using the wake-up word A training set to obtain training data for wake-up word A.
[0081] (3) The wake-up model is trained using the wake-up word B training set to obtain training data for wake-up word B.
[0082] (4) The wake-up model, the training data of wake-up word A, and the training data of wake-up word B are stored in the voice module for wake-up recognition.
[0083] (5) The microphone collects audio in real time, and after front-end processing, audio is obtained ready for recognition.
[0084] (6) Process the audio to be recognized using the training data of the wake-up model and wake-up word A to obtain a confidence A and a confidence threshold A.
[0085] (7) Process the audio to be recognized using the training data of the wake-up model and the wake-up word B to obtain a confidence B and a corresponding confidence threshold B.
[0086] (8) If reliability A≧reliability threshold A or reliability B≧reliability threshold B, a wake-up event is triggered.
[0087] In another possible embodiment, the following method can be adopted for the processing step (8).
[0088] (1) If reliability A≧reliability threshold A and reliability B<reliability threshold B, a wake-up event of A is triggered.
[0089] (2) If reliability A<reliability threshold A and reliability B≧reliability threshold B, a wake-up event of B is triggered.
[0090] (3) If reliability A≧reliability threshold A and reliability B≧reliability threshold B, a comprehensive judgment is made based on the percentage value exceeding the reliability threshold, and a wake-up event is triggered.
[0091] In addition, in the solutions of the related art, when multiple wake-up words are trained simultaneously, crosstalk occurs between different wake-up words, that is, when the training level is insufficient, the two wake-up words may sound blurred in the ambient noise, resulting in misidentification. Especially when overlapping words exist (e.g., "Miss Mei, Miss Mei" and "Miss Mei"), it is necessary to design a model and use a large number of training sets to distinguish each wake-up word, and due to the limited hardware storage resources and the requirement for wake-up response speed, it is relatively difficult to solve the crosstalk problem. In this embodiment, each set of wake-up word training sets is used for individual training, thereby eliminating the possibility of crosstalk with other wake-up words and achieving better recognition results with less training.
[0092] Using the wake-up words of Mi-chan Mi-chan in Cantonese (Siyu Mei Siyu Mei) and standard dialect (Shou Mei Shou Mei) as an example, Table 1 compares the model test data when the wake-up words are trained simultaneously and when they are trained separately.
[0093] [Table 1]
[0094] As can be seen from the model test data in Table 1 above, the solution of the embodiments of the present disclosure can achieve a small false wake-up rate when multiple wake-up words are recognized simultaneously, while the solutions of the related art train wake-up words simultaneously, and the enterprise standard requirement for false wake-up testing is no more than three times per 24 hours. For the embodiments of the present disclosure, if wake-up words are trained individually, false wake-up testing can be performed no more than once per 72 hours. Furthermore, when new wake-up words are added, the embodiments of the present disclosure separate the data, so re-training is only required for the newly added wake-up words, without affecting existing wake-up words, and development efficiency can be improved.
[0095] The embodiments of the present disclosure provide a wake-up processing method, and the above-mentioned embodiments are described in detail to specifically realize the embodiments. As can be seen, the technical solution of the above-mentioned embodiments can avoid the possibility of wake-up word crosstalk occurring when different wake-up words are trained simultaneously, realize that the training data for the wake-up words are separated and do not interfere with each other, and can also reduce the false wake-up rate of the audio equipment when multiple wake-up words are recognized simultaneously.
[0096] Example 4 Based on the same inventive idea as the above-mentioned embodiment, referring to Figure 5, there is shown a schematic diagram illustrating the configuration of a wake-up processing device provided by the embodiment of the present disclosure. As shown in Figure 5, the wake-up processing device 50 may include an acquisition unit 501, a processing unit 502, and a trigger unit 503, among which: an acquisition unit 501 arranged to acquire audio awaiting recognition; The processing unit 502 processes the audio to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence levels and corresponding confidence thresholds, respectively, where the at least two sets of training data are obtained by respectively training at least two sets of wake-up word training sets through the wake-up model; The trigger unit 503 is arranged to trigger a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and respective corresponding confidence thresholds.
[0097] In some embodiments, the acquisition unit 501 is specifically configured to acquire initial audio data by data acquisition by a sound acquisition device, and pre-process the initial audio data to obtain the audio to be recognized.
[0098] In some embodiments, each set of training data includes model parameters and a confidence threshold, and correspondingly, the processing unit 502 is specifically configured to process the audio to be recognized using the wake-up model and the model parameters in the at least two sets of training data, respectively, to obtain at least two confidence levels, and to obtain confidence thresholds corresponding to each of the at least two confidence levels from the at least two sets of training data.
[0099] In some embodiments, the at least two sets of training data include a first set of training data including first model parameters and a first confidence threshold, and a second set of training data including second model parameters and a second confidence threshold; correspondingly, the processing unit 502 is specifically configured to process the audio to be recognized using the wake-up model and the first model parameters in the first set of training data to obtain a first confidence level and determine the first confidence threshold corresponding to the first confidence level from the first set of training data; process the audio to be recognized using the wake-up model and the second model parameters in the second set of training data to obtain a second confidence level and determine the second confidence threshold corresponding to the second confidence level from the second set of training data.
[0100] In some embodiments, the trigger unit 503 is specifically configured to trigger a wake-up event of the audio equipment when the first reliability is equal to or greater than the first reliability threshold, or when the second reliability is equal to or greater than the second reliability threshold.
[0101] In some embodiments, the wake-up event includes a first wake-up event and / or a second wake-up event, the first wake-up event having an association with a wake-up word corresponding to the first set of training data, and the second wake-up event having an association with a wake-up word corresponding to the second set of training data.
[0102] In some embodiments, the trigger unit 503 is specifically configured to: trigger the first wake-up event of the voice equipment when the first reliability is equal to or greater than the first reliability threshold and the second reliability is less than the second reliability threshold; or trigger the second wake-up event of the voice equipment when the second reliability is equal to or greater than the second reliability threshold and the first reliability is less than the first reliability threshold; or calculate a first value by which the first reliability exceeds the first reliability threshold and a second value by which the second reliability exceeds the second reliability threshold when the first reliability is equal to or greater than the first reliability threshold and the second reliability is equal to or greater than the second reliability threshold, and trigger a target wake-up event of the voice equipment according to the first value and the second value.
[0103] In some embodiments, the trigger unit 503 is further configured to determine that the target wake-up event is the first wake-up event and triggered if the first value is greater than or equal to the second value, or to determine that the target wake-up event is the second wake-up event and triggered if the first value is less than the second value.
[0104] In some embodiments, the acquiring unit 501 is further configured to acquire the at least two wake-up word training sets; The processing unit 502 further trains the wake-up model using the at least two wake-up word training sets to obtain the at least two sets of training data, each set of training data being arranged to include model parameters and a confidence threshold.
[0105] In some embodiments, the acquiring unit 501 is further configured to acquire an initial training set, in which the initial training set includes at least two wake-up words, and group the initial training set according to different wake-up words to obtain the at least two sets of wake-up word training sets.
[0106] In this embodiment, a "unit" may be a partial circuit, a partial processor, a partial program, or software, and may be a module or a non-module. Furthermore, the components in this embodiment may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The integrated unit may be realized in the form of hardware or a software functional module.
[0107] If the integrated unit is realized in the form of a software functional module rather than being sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, an essential part or a part that contributes to the prior art, or all or a part of the technical solution, may be embodied in the form of a software product, and the software product may be stored in a storage medium and include a plurality of instructions for causing a computer (which may be a personal computer, a server, an air conditioner, or network equipment) or a processor to execute all or a part of the steps of the method described in this embodiment. Meanwhile, the above-mentioned storage medium includes various media capable of storing program code, such as a USB memory, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, and an optical disk.
[0108] Therefore, the present embodiment proposes a computer storage medium storing a wake-up processing program which, when executed by at least one processor, implements the steps of the method according to any one of the previous embodiments.
[0109] Based on the above-described configuration of the wake-up processing device 50 and the computer storage medium, FIG. 6 is a schematic diagram of a specific hardware structure of the wake-up processing device 50 provided by the embodiment of the present disclosure. As shown in FIG. 6, the wake-up processing device 50 may include a communication interface 601, a memory 602, and a processor 603, and each component is coupled to each other via a bus system 604. It should be understood that the bus system 604 is used to realize communication between these components. In addition to a data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for clarity, various buses are shown as the bus system 604 in FIG. 6. Among them, the communication interface 601 is used to receive and transmit signals in the process of transmitting and receiving information to and from other external network elements. The memory 602 is for storing computer programs executable on the processor 603. During execution of the computer program, the processor 603 Obtaining audio awaiting recognition; and processing each of the audios to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds, wherein the at least two sets of training data are obtained by respectively training at least two sets of wake-up word training sets through the wake-up model; triggering a wake-up event for the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold.
[0110] It should be understood that memory 602 in embodiments of the present disclosure may be volatile, nonvolatile, or both. Nonvolatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which acts as an external cache. By way of example and not limitation, many types of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (Direct Rambus RAM, DRRAM). Memory 602 in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0111] Alternatively, the processor 603 may be an integrated circuit chip having signal processing capabilities. In the implemented process, each step of the above method can be achieved by a hardware integrated logic circuit or software instructions in the processor 603. The processor 603 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Each method, step, and logic block diagram disclosed in the embodiments of the present disclosure can be realized or executed. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the disclosed methods in the embodiments of the present disclosure may be directly embodied as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium that is mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable and writable programmable memory, a register, etc. The storage medium is located in the memory 602, and the processor 603 reads the information in the memory 602 and completes the steps of the above method in association with the hardware.
[0112] It will be understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For a hardware implementation, the processing unit may be implemented in one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processing (DSPs), Digital Signal Processing Devices (DSP Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units, or a combination thereof, for performing the functions described in this disclosure.
[0113] For a software implementation, the techniques described herein may be implemented with modules (e.g., processes, functions, etc.) that perform the functions described herein. The software code may be stored in a memory and executed by a processor. The memory may be implemented within or external to the processor.
[0114] Optionally, as another embodiment, the processor 603 is further arranged to, during execution of the computer program, perform the steps of the method according to any one of the preceding embodiments.
[0115] It should be noted that in this disclosure, the terms "comprise," "include," or any other variation thereof is intended to cover a non-exclusive inclusion, whereby a process, method, article, or apparatus comprising a set of elements includes not only those elements but also other elements not expressly listed or inherent in such process, method, article, or apparatus. In the absence of further limitations, elements qualified by the phrase "comprises one ..." do not exclude the presence of further identical elements in a process, method, article, or apparatus that comprises the element.
[0116] The numbers of the above-mentioned embodiments of the present disclosure are merely for illustrative purposes and do not indicate superiority or inferiority of the embodiments.
[0117] The methods disclosed in the several method embodiments provided by this disclosure can be combined in any manner consistent with one another to obtain new method embodiments.
[0118] The features disclosed in the several product embodiments provided by this disclosure may be combined in any manner consistent with one another to produce novel feature embodiments.
[0119] The features disclosed in the several method or apparatus embodiments provided by this disclosure may be combined in any manner consistent with one another to produce new method or apparatus embodiments.
[0120] The above are only specific embodiments of the present disclosure, but the scope of protection of the present disclosure is not limited thereto, and all modifications or substitutions that a person skilled in the art can easily conceive within the technical scope described in the present disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be equivalent to the scope of protection of the claims.
Claims
1. 1. A wake-up processing method applied to an audio facility, the method comprising: obtaining audio to be recognized; processing each of the audios to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds, wherein the at least two sets of training data are obtained by training at least two sets of wake-up word training sets through the wake-up model, respectively; triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold; the at least two sets of training data include a first set of training data including first model parameters and a first confidence threshold, and a second set of training data including second model parameters and a second confidence threshold; processing the audio to be recognized using the wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds, processing the audio to be recognized using the wake-up model and the first model parameters in the first set of training data to obtain a first confidence level; and determining the first confidence level threshold corresponding to the first confidence level from the first set of training data; processing the audio to be recognized using the wake-up model and the second model parameters in the second set of training data to obtain a second confidence level; and determining the second confidence level threshold corresponding to the second confidence level from the second set of training data; Triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold, triggering a wake-up event of the audio facility when the first reliability is greater than or equal to the first reliability threshold or when the second reliability is greater than or equal to the second reliability threshold; The wake-up event includes a first wake-up event and / or a second wake-up event, the first wake-up event having an association with a wake-up word corresponding to the first set of training data, and the second wake-up event having an association with a wake-up word corresponding to the second set of training data; Triggering a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold, Triggering the first wake-up event of the audio equipment when the first reliability is greater than or equal to the first reliability threshold and the second reliability is less than the second reliability threshold; or Triggering the second wake-up event of the audio equipment when the second reliability is greater than or equal to the second reliability threshold and the first reliability is less than the first reliability threshold; or and if the first reliability is greater than or equal to the first reliability threshold and the second reliability is greater than or equal to the second reliability threshold, calculating a first value by which the first reliability exceeds the first reliability threshold and a second value by which the second reliability exceeds the second reliability threshold, and triggering a target wake-up event of the audio equipment according to the first value and the second value.
2. The step of obtaining audio to be recognized includes: collecting data using a sound collecting device to obtain initial sound data; and preprocessing the initial speech data to obtain the audio to be recognized.
3. Triggering a target wake-up event of the audio equipment according to the first value and the second value includes: if the first value is greater than or equal to the second value, the target wake-up event is determined to be the first wake-up event and is triggered; or 2. The method of claim 1, further comprising: if the first value is less than the second value, determining that the target wake-up event is the second wake-up event and triggering the second wake-up event.
4. The method comprises: obtaining the at least two wake-up word training sets; 2. The method of claim 1, further comprising: training the wake-up model using the at least two wake-up word training sets to obtain the at least two sets of training data, each set of training data including model parameters and a confidence threshold.
5. The step of obtaining at least two wake-up word training sets includes: obtaining an initial training set including at least two wake-up words; and b. grouping the initial training set according to different wake-up words to obtain the at least two wake-up word training sets.
6. A wake-up processing device applied to an audio equipment, the wake-up processing device comprising: an acquisition unit arranged to acquire audio awaiting recognition; a processing unit for processing each of the audios to be recognized using a wake-up model and at least two sets of training data to obtain at least two confidence measures and corresponding confidence thresholds, the at least two sets of training data being arranged such that at least two sets of wake-up word training sets are respectively trained via the wake-up model; a trigger unit configured to trigger a wake-up event of the audio equipment based on a comparison result between the at least two confidence levels and a corresponding confidence level threshold; the at least two sets of training data include a first set of training data including first model parameters and a first confidence threshold, and a second set of training data including second model parameters and a second confidence threshold; The processing unit processing the audio to be recognized using the wake-up model and the first model parameters in the first set of training data to obtain a first confidence measure; and determining the first confidence measure threshold from the first set of training data corresponding to the first confidence measure; configured to process the audio to be recognized using the wake-up model and the second model parameters in the second set of training data to obtain a second confidence measure; and determine the second confidence measure threshold corresponding to the second confidence measure from the second set of training data; The trigger unit configured to trigger a wake-up event of the audio equipment when the first reliability is equal to or greater than the first reliability threshold or when the second reliability is equal to or greater than the second reliability threshold; The wake-up event includes a first wake-up event and / or a second wake-up event, the first wake-up event having an association with a wake-up word corresponding to the first set of training data, and the second wake-up event having an association with a wake-up word corresponding to the second set of training data; The trigger unit If the first reliability is greater than or equal to the first reliability threshold and the second reliability is less than the second reliability threshold, the first wake-up event of the audio facility is triggered; or If the second reliability is greater than or equal to the second reliability threshold and the first reliability is less than the first reliability threshold, the second wake-up event of the audio facility is triggered; or if the first reliability is greater than or equal to the first reliability threshold and the second reliability is greater than or equal to the second reliability threshold, a first value by which the first reliability exceeds the first reliability threshold and a second value by which the second reliability exceeds the second reliability threshold are calculated, and the first value and the second value are arranged to trigger a target wake-up event of the audio equipment. A wake-up processing device comprising:
7. 1. An audio facility including a memory and a processor, the memory is used to store a computer program executable by the processor; An audio installation, characterized in that the processor is used to carry out the method according to any one of claims 1 to 5 during execution of the computer program.
8. A computer storage medium having stored thereon a computer program which, when executed by at least one processor, implements the method according to any one of claims 1 to 5.
9. A computer program, characterized in that when executed by at least one processor, the computer program implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Information processor, information processing method and program
JP2020042131A
Method, electronic device, home appliance network and storage medium
JP2020525850A
Hotword recognition and passive assistance
WO2020032948A1