Voice wake-up method, device, and storage medium
Patent Information
- Application Number
- CN202311369821.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-23
AI Technical Summary
[0005]本申请提供了一种语音唤醒方法、设备及存储介质,可以解决由于传统的语音唤醒方法的置信度预设阈值条件较为单一,这就导致了传统的语音唤醒方法唤醒率较低的问题
[0039]本申请的有益效果在于:获取目标设备采集的语音数据;确定所述语音数据中每个音素的音素置信度和唤醒词的第一置信度。基于所述音素置信度确定唤醒词第二置信度;在所述唤醒词第一置信度满足第一唤醒条件且所述唤醒词第二置信度满足第二唤醒条件的情况下,基于所述语音数据唤醒所述目标设备;可以解决由于传统的语音唤醒方法的置信度预设阈值条件较为单一,这就导致了传统的语音唤醒方法唤醒率较低,误唤醒高的问题。通过分别确定语音数据中的唤醒词第一置信度和唤醒词第二置信度并在唤醒词第一置信度和唤醒词第二置信度同时满足条件的情况下唤醒设备,由于唤醒词第一置信度和唤醒词第二置信度具有互补性,这样通过阈值比较可以达到降低误唤醒率并提升唤醒率的目的。
Smart Images

Figure CN117334195B_ABST
Abstract
Description
[Technical Field]
[0001] This application belongs to the field of audio data processing, specifically relating to a voice wake-up method, device, and storage medium. [Background Technology]
[0002] Voice wake-up refers to the process by which a user wakes up a device by speaking a preset wake-up word.
[0003] Traditional voice wake-up methods include: acquiring audio data; inputting the speech data from the audio data into a pre-trained wake-up model to obtain the confidence level of the wake-up word in the speech data; and determining the wake-up word in the speech data as the device wake-up word if the confidence level of the wake-up word is greater than a preset threshold.
[0004] However, because the confidence threshold conditions of traditional voice wake-up methods are relatively simple, this leads to a low wake-up rate and a high false wake-up rate. [Summary of the Invention]
[0005] This application provides a voice wake-up method, device, and storage medium, which can solve the problem that the wake-up rate of traditional voice wake-up methods is low due to the relatively simple preset confidence threshold conditions. This application provides the following technical solution.
[0006] Firstly, a voice wake-up method is provided, the method comprising:
[0007] Acquire voice data collected by the target device;
[0008] Determine the phoneme confidence level and the first confidence level of the wake word for each phoneme in the speech data.
[0009] The second confidence level of the wake word is determined based on the phoneme confidence level;
[0010] When the first confidence level of the wake word satisfies the first wake-up condition and the second confidence level of the wake word satisfies the second wake-up condition, the target device is woken up based on the voice data.
[0011] Optionally, the step of satisfying the first wake-up condition with the first confidence level of the wake-up word and satisfying the second wake-up condition with the second confidence level of the wake-up word includes:
[0012] If the first confidence level is greater than the first threshold and the second confidence level is greater than the second threshold, the target device is woken up based on the voice data;
[0013] Alternatively, if the first confidence level is greater than the third threshold and the second confidence level is greater than the fourth threshold, the target device is woken up based on the voice data; the third threshold is greater than the first threshold and the fourth threshold is less than the second threshold.
[0014] Alternatively, if the first confidence level is greater than the fifth threshold and the second confidence level is greater than the sixth threshold, the target device is woken up based on the voice data; the fifth threshold is less than the first threshold; and the sixth threshold is greater than the second threshold.
[0015] Optionally, the phoneme confidence score is obtained based on a first wake-up model; the wake-up word confidence score is obtained based on a second wake-up model.
[0016] The first threshold, the third threshold, and the fifth threshold are positive and negative test sets using wake words. The positive test set is also called the wake-up rate test set, and the negative test set is also called the false wake-up test set. They are obtained based on the output of the first wake-up model when testing the first wake-up model.
[0017] The second threshold, the fourth threshold, and the sixth threshold are obtained based on the output results of the second wake-up model when the same positive and negative test sets of wake-up words are used to test the second wake-up model.
[0018] Optionally, determining the phoneme confidence level and the first confidence level of the wake word for each phoneme in the speech data includes:
[0019] The speech data is input into a pre-trained fusion wake-up model to obtain the phoneme confidence score and the first confidence score of the wake word. The fusion wake-up model includes a feature extraction network and a first output layer and a second output layer connected to the feature extraction network. The feature extraction network is used to extract phoneme features and wake word features from the speech data. The first output layer is used to output the phoneme confidence score. The second output layer is used to output the first confidence score of the wake word.
[0020] Optionally, the training process of the fusion wake-up model includes:
[0021] Acquire training data, which includes sample speech data, phoneme labels for each phoneme corresponding to the sample speech data, and wake word labels for each wake word corresponding to the sample speech data;
[0022] The sample speech data is input into a pre-created initial network model to obtain the first training result;
[0023] The first training result and the phoneme label are input into a preset first loss function to obtain the first loss function value;
[0024] The pre-created initial network model is iteratively trained using the first loss function value to obtain the trained initial network model.
[0025] Keeping the feature extraction layer parameters of the initial network model unchanged, update the parameters of the second output layer;
[0026] The sample speech data is input into the trained initial network model to obtain the second training result;
[0027] The second training result and the wake word label are input into a preset second loss function to obtain the second loss function value;
[0028] The trained initial network model is iteratively trained using the second loss function value to obtain the trained fusion wake-up model.
[0029] Optionally, determining the phoneme confidence level and the first confidence level of the wake word for each phoneme in the speech data includes:
[0030] The speech data is input into the first wake-up model to obtain the phoneme confidence of each phoneme in the speech data;
[0031] The voice data is input into the second wake-up model to obtain the first confidence level of each wake-up word in the voice data.
[0032] Optionally, determining the second confidence level of the wake word based on the phoneme confidence level includes:
[0033] The second confidence level of the wake word is obtained by weighted averaging of the confidence levels of the phonemes.
[0034] Optionally, before acquiring the voice data collected by the target device, the method further includes:
[0035] Acquire audio data collected by the target device;
[0036] The audio data is input into the speech recognition model to obtain speech data.
[0037] In a second aspect, an electronic device is provided, the device including a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement the voice wake-up method as described in the first aspect.
[0038] Thirdly, a computer-readable storage medium is provided, wherein a program is stored therein, which, when executed by a processor, is used to implement the voice wake-up method as described in the first aspect.
[0039] The beneficial effects of this application are as follows: It acquires voice data collected by the target device; determines the phoneme confidence level and the first confidence level of the wake-up word for each phoneme in the voice data; determines the second confidence level of the wake-up word based on the phoneme confidence level; and wakes up the target device based on the voice data when the first confidence level of the wake-up word meets a first wake-up condition and the second confidence level of the wake-up word meets a second wake-up condition. This solves the problem that traditional voice wake-up methods have low wake-up rates and high false wake-ups due to the relatively simple preset threshold conditions for confidence levels. By separately determining the first and second confidence levels of the wake-up word in the voice data and waking up the device when both the first and second confidence levels of the wake-up word meet the conditions, the first and second confidence levels of the wake-up word are complementary. Therefore, by comparing thresholds, the false wake-up rate can be reduced and the wake-up rate improved.
[0040] In addition, since the target device may collect audio data that does not contain voice data during the voice data acquisition process, the target device cannot be woken up if the audio data without voice data is processed. Based on the above technical problem, this embodiment further includes: acquiring audio data collected by the target device before acquiring the voice data collected by the target device; inputting the audio data into the speech recognition model to obtain voice data. This can ensure that audio data with voice data can be processed, further improving the efficiency of voice data processing.
[0041] In addition, since the fusion wake-up model is based on the first wake-up model and the second wake-up model for fusion modeling, the parameters of the input layer and feature extraction layer of the fusion wake-up model can be fully shared with the first wake-up model and the second wake-up model. Therefore, the amount of computation is increased slightly, which can improve the efficiency of network computation.
[0042] In addition, during the training of the fusion wake-up model, the initial network model is trained first, that is, the parameters of the first wake-up model part in the fusion wake-up model are trained. After the initial network model is trained, the parameters of the second wake-up model are trained. This ensures that when the network parameters change in the future, the overall network structure can be modified with a small parameter variable. [Attached Image Description]
[0043] Figure 1 This is a flowchart of a voice wake-up method provided in one embodiment of this application;
[0044] Figure 2 This is a schematic diagram of the fusion wake-up model structure provided in one embodiment of this application;
[0045] Figure 3 This is a flowchart of a voice wake-up method provided in another embodiment of this application.
[0046] Figure 4 This is a block diagram of a voice wake-up device provided in one embodiment of this application;
[0047] Figure 5 This is a block diagram of an electronic device provided in one embodiment of this application.
Detailed Implementation Methods
[0048] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. The application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0049] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0050] In this application, unless otherwise stated, directional terms such as "upper," "lower," "top," and "bottom" are generally used in relation to the direction shown in the accompanying drawings, or in relation to the vertical, perpendicular, or gravitational direction of the component itself; similarly, for ease of understanding and description, "inner" and "outer" refer to the inner and outer contours of each component itself, but the above directional terms are not used to limit this application.
[0051] The voice wake-up method provided in this application will be described in detail below.
[0052] Optionally, the execution subject of the voice wake-up method provided in this application is an electronic device, which can be a terminal such as a computer, mobile phone, tablet computer, or camera, or it can be a server. This embodiment does not limit the implementation method of the electronic device.
[0053] This embodiment provides a voice wake-up method, such as... Figure 1 As shown, the method includes at least the following steps:
[0054] Step 101: Acquire the voice data collected by the target device.
[0055] Optionally, the target device can be an electronic device including an audio acquisition component, which can be a microphone or a microphone array, etc. This embodiment does not limit the type of audio acquisition component.
[0056] Among them, voice data refers to audio data containing speech.
[0057] Since the target device may collect audio that does not contain voice data during the voice data collection process, and waking up the target device cannot be achieved if the audio that does not contain voice data is processed. Based on the above technical problem, before acquiring the voice data collected by the target device in this embodiment, the method further includes: acquiring audio data collected by the target device; inputting the audio data into a speech recognition model to obtain the voice data.
[0058] Optionally, the audio data includes, but is not limited to: audio encoding, audio code stream, number of audio channels, audio quantization bits, audio sampling frequency, bit rate, etc. This embodiment does not limit the type of the audio data.
[0059] In one example, the speech recognition model may be a Voice Activity Detection (VAD) model. The voice activity detection model can identify and eliminate long periods of silence from the audio signal stream, so as to save channel resources without reducing service quality. In other words, the voice activity detection model can distinguish voice data and non-voice data in the audio data to obtain the voice data.
[0060] In another example, the speech recognition model may be a network recognition model trained based on a neural network, and the network recognition model can accurately identify and output voice data from the audio data.
[0061] In step 102, a phoneme confidence level of each phoneme in the voice data and a first confidence level of the wake-up word are determined.
[0062] Wherein, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the articulation movement in a syllable, one movement constitutes one phoneme, and phonemes are divided into two categories: vowels and consonants. For example, the Chinese syllable "啊 (ā)" has only one phoneme, and the syllable "代 (dài)" has two phonemes.
[0063] Wherein, confidence level is also called reliability, confidence level, or confidence coefficient. That is, when estimating population parameters through sampling, the conclusion is always uncertain due to the randomness of the sample. Therefore, a probabilistic statement method is adopted, that is, the interval estimation method in mathematical statistics, which means that within a certain allowable error range between the estimated value and the population parameter, what is the corresponding probability, and this corresponding probability is called confidence level
[0064] In one example, determining the phoneme confidence level of each phoneme in the voice data and the first confidence level of the wake-up word includes: inputting the voice data into a pre-trained fused wake-up model to obtain the phoneme confidence level and the first confidence level of the wake-up word.
[0065] The fusion wake-up model includes a feature extraction network and a first output layer and a second output layer connected to the feature extraction network. The feature extraction network is used to extract phoneme features and wake-up word features from the speech data. The first output layer outputs phoneme confidence scores, and the second output layer outputs the first confidence score of the wake-up word.
[0066] Optionally, the training process of the fusion wake-up model includes at least the following steps S1 to S8:
[0067] Step S1: Obtain training data.
[0068] The training data includes sample speech data, phoneme labels for each phoneme in the sample speech data, and wake word labels for each wake word in the sample speech data.
[0069] Step S2: Input the sample speech data into the pre-created initial network model to obtain the first training result.
[0070] Step S3: Input the first training result and phoneme label into the preset first loss function to obtain the first loss function value.
[0071] Step S4: Iteratively train the pre-created initial network model using the first loss function value to obtain the trained initial network model.
[0072] Step S5: Keep the feature extraction layer parameters of the initial network model unchanged, and update the parameters of the second output layer.
[0073] Step S6: Input the sample speech data into the trained initial network model to obtain the second training result.
[0074] Step S7: Input the second training result and the wake word label into the preset second loss function to obtain the second loss function value.
[0075] Step S8: Use the second loss function value to iteratively train the initial network model after training to obtain the trained fusion wake-up model.
[0076] In this embodiment, the fusion wake-up model combines the first wake-up model and the second wake-up model. The first and second wake-up models are trained based on neural networks. Both the first and second wake-up models include an input layer, a feature extraction layer, and an output layer. The first wake-up model is used to output phoneme confidence, and the second wake-up model is used to output wake word confidence.
[0077] according to Figure 2It can be seen that, by comparing the fused wake-up model (right) with the first wake-up model (left), the input layer and feature extraction layer of the two models have consistent structures. The fused wake-up model fuses the first recognition network and the second recognition network, and has one more output layer than the first recognition network and the second recognition network, while the parameters of the input layer and feature extraction layer of the fused wake-up model can be completely shared with those of the first recognition network and the second recognition network. Therefore, the amount of calculation only increases slightly, which can improve the efficiency of network calculation.
[0078] In addition, in the process of training the fused wake-up model, the initial network model is trained first, that is, the parameters of the first recognition network part in the fused wake-up model are trained. After the training of the initial network model is completed, the overall network is trained. This can ensure that while the subsequent network parameters change, the overall network structure can be modified with a small amount of parameter change.
[0079] In another example, determining the phoneme confidence of each phoneme in the speech data and the first confidence of the wake word comprises: inputting the speech data into a first recognition model to obtain the phoneme confidence of each phoneme in the speech data; inputting the speech data into a second recognition model to obtain the first confidence of the wake word in the speech data.
[0080] Step 103: determining a second confidence of the wake word in the speech data.
[0081] In this embodiment, determining the second confidence of the wake word based on the phoneme confidence comprises: performing weighted average on the phoneme confidence to obtain the second confidence of the wake word.
[0082] For example: when the speech data is "ni hao xiao chi", each phoneme of "ni hao xiao chi" is "n-i-h-ao-x-iao-ch-i", the confidence corresponding to each phoneme is 0.95, 0.96, 0.96, 0.95, 0.94, 0.96, 0.96, 0.94 respectively, and the weight of each phoneme is X1, X2...X8 respectively, then the weighted average of the above phoneme confidence is performed to obtain the second confidence of the wake word.
[0083] Step 104: waking up a target device based on the speech data when the first confidence of the wake word satisfies a first wake-up condition and the second confidence of the wake word satisfies a second wake-up condition.
[0084] In one example, the condition that the first confidence of the wake word satisfies a first wake-up condition and the second confidence of the wake word satisfies a second wake-up condition comprises: waking up the target device based on the speech data when the first confidence is greater than a first threshold and the second confidence is greater than a second threshold.
[0085] For example, when the first confidence level of the wake word is 0.6 and the second confidence level of the wake word is 0.6, the first threshold is 0.5 and the second threshold is 0.5. At this time, the first confidence level of the wake word is greater than the first threshold and the second confidence level of the wake word is greater than the second threshold, then the target device is woken up.
[0086] In another example, when the first confidence level of the wake word satisfies the first wake-up condition and the second confidence level of the wake word satisfies the second wake-up condition, the method includes: waking up the target device based on voice data when the first confidence level is greater than the third threshold and the second confidence level is greater than the fourth threshold; the third threshold is greater than the first threshold and the fourth threshold is less than the second threshold.
[0087] For example, when the first confidence level of the wake word is 0.9 and the second confidence level of the wake word is 0.4, the third threshold is 0.8 and the fourth threshold is 0.3. At this time, the first confidence level of the wake word is greater than the third threshold and the second confidence level of the wake word is greater than the fourth threshold, then the target device is woken up.
[0088] In this embodiment, since the first confidence level of the wake-up word is already greater than the third threshold, and the third threshold is a higher threshold obtained based on the test, it indicates that the first confidence level of the wake-up word is already high. At this time, the second confidence level of the wake-up word only needs to be greater than the lower value of the fourth threshold to wake up the target device.
[0089] In another example, when the first confidence level of the wake word meets the first wake-up condition and the second confidence level of the wake word meets the second wake-up condition, the following steps are taken: when the first confidence level is greater than the fifth threshold and the second confidence level is greater than the sixth threshold, the target device is woken up based on the voice data; the fifth threshold is less than the first threshold; and the sixth threshold is greater than the second threshold.
[0090] For example, when the first confidence level of the wake word is 0.4 and the second confidence level of the wake word is 0.9, the fifth threshold is 0.3 and the sixth threshold is 0.8. At this time, the first confidence level of the wake word is greater than the fifth threshold and the second confidence level of the wake word is greater than the sixth threshold, so the target device is woken up.
[0091] Optionally, the phoneme confidence score is obtained based on the first recognition network, and the wake word confidence score is obtained based on the second recognition network.
[0092] Accordingly, the first, third, and fifth thresholds are obtained based on the output of the first recognition network when tested using a positive and negative test set of wake-up words (the positive test set is also called the wake-up rate test set, and the negative test set is also called the false wake-up test set). The second, fourth, and sixth thresholds are obtained based on the output of the second recognition network when tested using the same positive and negative test sets of wake-up words.
[0093] In summary, the voice wake-up method provided in this embodiment,
[0094] The first confidence level of the wake word is determined. A second confidence level of the wake word is determined based on the phoneme confidence level. When the first confidence level of the wake word satisfies a first wake-up condition and the second confidence level of the wake word satisfies a second wake-up condition, the target device is woken up based on the voice data. This solves the problem that traditional voice wake-up methods have low wake-up rates and high false wake-ups due to the relatively simple preset threshold conditions for confidence levels. By separately determining the first and second confidence levels of the wake word in the voice data and waking up the device when both the first and second confidence levels are satisfied, the first and second confidence levels are complementary. Therefore, threshold comparison can reduce the false wake-up rate and improve the wake-up rate.
[0095] In addition, since the target device may collect audio data that does not contain voice data during the voice data acquisition process, the target device cannot be woken up if the audio data without voice data is processed. Based on the above technical problem, this embodiment further includes: acquiring audio data collected by the target device before acquiring the voice data collected by the target device; inputting the audio data into the speech recognition model to obtain voice data. This can ensure that audio data with voice data can be processed, further improving the efficiency of voice data processing.
[0096] In addition, since the fusion wake-up model is based on the first wake-up model and the second wake-up model, the parameters of the input layer and feature extraction layer of the fusion wake-up model can be fully shared with the first wake-up model and the second wake-up model. Therefore, the amount of computation is increased slightly, which can improve the efficiency of network output.
[0097] In addition, during the training of the fusion wake-up model, the initial network model is trained first, that is, the parameters of the first wake-up model part in the fusion wake-up model are trained. After the initial network model is trained, the entire network is trained. This ensures that when the network parameters change in the future, the overall network structure can be modified with a smaller parameter variable.
[0098] To better understand the voice wake-up method provided in this application, this embodiment illustrates the method with an example, see reference. Figure 3 The method includes at least the following steps:
[0099] Step 301: The audio acquisition module acquires external audio data.
[0100] Step 302: The VAD module processes the audio data, distinguishing it into voice data and non-voice data. If the audio data contains voice data, the voice data is input into the fusion wake-up engine for voice data processing. If the audio data does not contain voice data, step 301 will be executed again.
[0101] Step 303: The wake-up engine is integrated to decode and classify the speech data, and two confidence levels are obtained through feature extraction to double-determine whether the speech data contains a wake-up word.
[0102] Among them, the fusion wake-up engine is a multi-scale modeling fusion acoustic model that integrates the first wake-up model and the second wake-up model. The first wake-up model is used to output phoneme confidence, and the second wake-up model is used to output wake word confidence.
[0103] Step 304: After processing, the fusion wake-up model outputs phoneme confidence and wake-up word confidence. The two confidences are judged to obtain the judgment result. If the judgment result indicates that the speech data contains a wake-up word, the wake-up information is output. If the judgment result indicates that the speech data does not contain a wake-up word, step 301 is executed.
[0104] For example: Score1 represents the first confidence score of the wake word obtained based on the fusion wake-up model, and Score2 represents the second confidence score of the wake word obtained based on the fusion wake-up model. Thresh_normal1 represents the general threshold based on phoneme modeling, and Thresh_normal2 represents the general threshold based on complete wake word modeling; Thresh_high1 represents the high threshold based on phoneme modeling, and Thresh_high2 represents the high threshold based on complete wake word modeling; Thresh_low1 represents the low threshold based on phoneme modeling, and Thresh_low2 represents the low threshold based on complete wake word modeling. Multiple conditional judgments are performed, as follows:
[0105] The scores for both scales, Score1 > Thresh_normal1 and Socre2 > Thresh_normal2, are greater than their respective general thresholds.
[0106] The phoneme modeling scores of Score1 > Thresh_high1 and Score2 > Thresh_low2 exceed the set high threshold, and the complete wake-up modeling score exceeds the set low threshold.
[0107] The phoneme modeling scores of Score1 > Thresh_low1 and Score2 > Thresh_high2 exceed the set low threshold, and the complete wake word modeling score exceeds the set high threshold.
[0108] Figure 4This is a block diagram of a voice wake-up device provided in one embodiment of this application. The device includes at least the following modules: a data acquisition module 410, a first determination module 420, a second determination module 430, and a voice wake-up module 440.
[0109] The data acquisition module 410 is used to acquire voice data collected by the target device.
[0110] The first determining module 420 is used to determine the phoneme confidence level of each phoneme in the speech data and the first confidence level of the wake word.
[0111] The second determining module 430 is used to determine the second confidence level of the wake word based on the phoneme confidence level.
[0112] The voice wake-up module 440 is used to wake up the target device based on the voice data when the first confidence level of the wake-up word meets the first wake-up condition and the second confidence level of the wake-up word meets the second wake-up condition.
[0113] For relevant details, please refer to the above embodiments.
[0114] It should be noted that the voice wake-up device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the voice wake-up device can be divided into different functional modules to complete all or part of the functions described above. In addition, the voice wake-up device and the voice wake-up method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0115] This embodiment provides an electronic device, such as... Figure 5 As shown, the electronic device can be Figure 1 The target device in the context of the electronic device. This electronic device includes at least a processor 501 and a memory 502.
[0116] Processor 501 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 501 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 501 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0117] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 502 are used to store at least one instruction, which is executed by the processor 501 to implement the voice wake-up method provided in the method embodiments of this application.
[0118] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 501, memory 502, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuitry, a touch display screen, audio circuitry, and a power supply.
[0119] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.
[0120] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the voice wake-up method of the above method embodiments.
[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0122] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A voice wake-up method, characterized in that, The method includes: Acquire voice data collected by the target device; Determine the phoneme confidence level and the first confidence level of the wake word for each phoneme in the speech data; Determining the phoneme confidence score and the first confidence score of the wake word for each phoneme in the speech data includes: inputting the speech data into a pre-trained fusion wake-up model to obtain the phoneme confidence score and the first confidence score of the wake word; the fusion wake-up model includes a feature extraction network and a first output layer and a second output layer connected to the feature extraction network; the feature extraction network is used to extract phoneme features and wake word features from the speech data; the first output layer is used to output the phoneme confidence score; the second output layer is used to output the first confidence score of the wake word; The training process of the fusion wake-up model includes: acquiring training data, which includes sample speech data, phoneme labels for each phoneme corresponding to the sample speech data, and wake-up word labels for each wake-up word corresponding to the sample speech data; inputting the sample speech data into a pre-created initial network model to obtain a first training result; inputting the first training result and the phoneme labels into a preset first loss function to obtain a first loss function value; using the first loss function value to iteratively train the pre-created initial network model to obtain a trained initial network model; keeping the feature extraction layer parameters of the initial network model unchanged, and updating the parameters of the second output layer; inputting the sample speech data into the trained initial network model to obtain a second training result; inputting the second training result and the wake-up word labels into a preset second loss function to obtain a second loss function value; using the second loss function value to iteratively train the trained initial network model to obtain a trained fusion wake-up model. The second confidence level of the wake word is determined based on the phoneme confidence level; When the first confidence level of the wake word satisfies the first wake-up condition and the second confidence level of the wake word satisfies the second wake-up condition, the target device is woken up based on the voice data; The method of waking up the target device based on the voice data when the first confidence level of the wake word satisfies the first wake-up condition and the second confidence level of the wake word satisfies the second wake-up condition includes: waking up the target device based on the voice data when the first confidence level is greater than a first threshold and the second confidence level is greater than a second threshold; or, waking up the target device based on the voice data when the first confidence level is greater than a third threshold and the second confidence level is greater than a fourth threshold; wherein the third threshold is greater than the first threshold and the fourth threshold is less than the second threshold; or, waking up the target device based on the voice data when the first confidence level is greater than a fifth threshold and the second confidence level is greater than a sixth threshold; wherein the fifth threshold is less than the first threshold and the sixth threshold is greater than the second threshold.
2. The method according to claim 1, characterized in that, The phoneme confidence level is obtained based on the first wake-up model; the wake-up word first confidence level is obtained based on the second wake-up model. The first threshold, the third threshold, and the fifth threshold are positive and negative test sets using wake words. The positive test set is also called the wake-up rate test set, and the negative test set is also called the false wake-up test set. They are obtained based on the output of the first wake-up model when testing the first wake-up model. The second threshold, the fourth threshold, and the sixth threshold are obtained based on the output results of the second wake-up model when the same positive and negative test sets of wake-up words are used to test the second wake-up model.
3. The method according to claim 1, characterized in that, Determining the phoneme confidence score and the first confidence score of the wake word for each phoneme in the speech data includes: The speech data is input into the first wake-up model to obtain the phoneme confidence of each phoneme in the speech data; The voice data is input into the second wake-up model to obtain the first confidence level of each wake-up word in the voice data.
4. The method according to claim 1, characterized in that, The determination of the second confidence level of the wake word based on the phoneme confidence level includes: The second confidence level of the wake word is obtained by weighted averaging of the confidence levels of the phonemes.
5. The method according to claim 1, characterized in that, Before acquiring the voice data collected by the target device, the process also includes: Acquire audio data collected by the target device; The audio data is input into the speech recognition model to obtain speech data.
6. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program that is loaded and executed by the processor to implement the voice wake-up method as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement the voice wake-up method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Processing method, device and system for voice to be tested
CN103810996A
Method and device for comprehensively evaluating voice and electronic equipment
CN112951276A