Voice wake-up method and device, terminal and earphone
By working together with the headset and the terminal, and employing a multi-level verification mechanism that includes wake-up word recognition and voiceprint verification, the problem of accidental wake-up of voice interaction functions has been solved, achieving more efficient wake-up accuracy and reduced power consumption.
Patent Information
- Application Number
- CN202511167875.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-04
AI Technical Summary
In existing technologies, voice interaction functions are easily accidentally activated, leading to resource waste and unnecessary power consumption increases.
By collecting audio through headphones and recognizing the wake word, combined with voiceprint verification on the terminal, multi-level verification is implemented to wake up the voice interaction function, ensuring that voice interaction is only activated when the wake word and voiceprint match.
It effectively reduces the probability of voice interaction functions being accidentally woken up, reduces unnecessary power consumption and resource waste, and improves the accuracy of wake-up.
Smart Images

Figure CN120895034A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of voice wake-up, and particularly relate to a voice wake-up method, device, terminal and earphone. BACKGROUND
[0002] With the continuous improvement of terminal performance, voice interaction function gradually becomes a standard function of the terminal.
[0003] In the related art, the voice interaction function can be woken up through voice. For example, a user speaks a specific wake-up word to wake up the voice interaction function of the terminal. SUMMARY
[0004] Embodiments of the present application provide a voice wake-up method, device, terminal and earphone. The technical solution is as follows:
[0005] In one aspect, the present application provides a voice wake-up method, which is used for a terminal, and the method comprises:
[0006] obtaining first audio sent by an earphone, the first audio being collected by the earphone through a first microphone and being sent when a first wake-up word recognition result of the first audio represents that the wake-up word is contained, the wake-up word being used to wake up a voice interaction function of the terminal;
[0007] performing voiceprint verification on the first audio to obtain a voiceprint verification result;
[0008] in a case where the voiceprint verification result represents that the first audio passes the voiceprint verification, waking up the voice interaction function.
[0009] In another aspect, the present application provides a voice wake-up method, which is used for an earphone, and the method comprises:
[0010] collecting audio through a first microphone;
[0011] performing wake-up word recognition on the collected first audio to obtain a first wake-up word recognition result, the first wake-up word recognition result being used to represent whether the first audio contains a wake-up word, the wake-up word being used to wake up a voice interaction function of a terminal;
[0012] in a case where the first wake-up word recognition result represents that the first audio contains the wake-up word, sending the first audio to the terminal, the terminal being used to wake up the voice interaction function in a case where the first audio passes voiceprint verification.
[0013] In another aspect, the present application provides a voice wake-up device, which comprises:
[0014] The acquisition module is configured to acquire first audio transmitted by the earphone, the first audio being collected by the earphone through a first microphone and transmitted when a first wake-up word recognition result of the first audio indicates that a wake-up word is contained, the wake-up word being used to wake up a voice interaction function of the terminal.
[0015] The voiceprint verification module is configured to perform voiceprint verification on the first audio to obtain a voiceprint verification result.
[0016] The wake-up module is configured to wake up the voice interaction function when the voiceprint verification result indicates that the first audio passes the voiceprint verification.
[0017] In another aspect, an embodiment of the present application provides a voice wake-up device, the device comprising:
[0018] The first collection module is configured to collect audio through a first microphone.
[0019] The wake-up word recognition module is configured to perform wake-up word recognition on the collected first audio to obtain a first wake-up word recognition result, the first wake-up word recognition result being used to indicate whether the first audio contains a wake-up word, the wake-up word being used to wake up a voice interaction function of a terminal.
[0020] The transmission module is configured to transmit the first audio to the terminal when the first wake-up word recognition result indicates that the first audio contains the wake-up word, the terminal being configured to wake up the voice interaction function when the first audio passes voiceprint verification.
[0021] In another aspect, an embodiment of the present application provides a terminal, the terminal comprising a processor and a memory, the memory storing at least one computer instruction, the at least one computer instruction being loaded and executed by the processor to implement the voice wake-up method according to the above aspect.
[0022] In another aspect, an embodiment of the present application provides an earphone, the earphone comprising a processor, a memory, and a microphone, the memory storing at least one computer instruction, the at least one computer instruction being loaded and executed by the processor to implement the voice wake-up method according to the above aspect.
[0023] In another aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing at least one computer instruction, the at least one computer instruction being used to be executed by a processor to implement the voice wake-up method according to the above aspect.
[0024] In another aspect, an embodiment of the present application provides a computer program product, the computer program product comprising computer instructions, the computer instructions being executed by a processor to implement the voice wake-up method according to the above aspect.
[0025] In the embodiments of the present application, when the voice interaction function of the terminal is woken up through the earphone worn, the earphone performs wake-up word recognition on the first audio collected by the first microphone, and sends the first audio to the terminal when the wake-up word is recognized, and the terminal further performs voiceprint verification on the first audio, so as to wake up the voice interaction function when the first audio passes the voiceprint verification, thereby realizing multi-level verification of voice wake-up. Since the voice interaction function is woken up only when the wake-up word recognition and the voiceprint verification are both successful, the probability of false wake-up of the voice interaction function can be reduced. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 A schematic diagram of an implementation environment provided by an example embodiment of the present application is shown;
[0027] Figure 2 A flowchart of a voice wake-up method provided by an example embodiment of the present application is shown;
[0028] Figure 3 A flowchart of a three-segment voice wake-up process provided by an example embodiment of the present application is shown;
[0029] Figure 4 An implementation schematic diagram of a three-segment voice wake-up process provided by an example embodiment of the present application is shown;
[0030] Figure 5 A flowchart of a four-segment voice wake-up process provided by an example embodiment of the present application is shown;
[0031] Figure 6 An implementation schematic diagram of a four-segment voice wake-up process provided by an example embodiment of the present application is shown;
[0032] Figure 7 A flowchart of a voiceprint verification optimization process provided by an example embodiment of the present application is shown;
[0033] Figure 8 A structural block diagram of a voice wake-up device provided by an example embodiment of the present application is shown;
[0034] Figure 9 A structural block diagram of a voice wake-up device provided by another example embodiment of the present application is shown;
[0035] Figure 10 A structural block diagram of an earphone provided by an example embodiment of the present application is shown;
[0036] Figure 11 A structural block diagram of a terminal provided by an example embodiment of the present application is shown. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0038] In the present application, "a plurality of" refers to two or more. The "and / or" describes the association between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0039] Please refer to Figure 1 which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application, which includes earphone 110 and terminal 120.
[0040] The earphone 110 is an electronic device with audio playback function. According to the wearing method, the earphone 110 can be a head-mounted, ear-pressed, ear-inserted, earplug-type, bone conduction earphone; according to the connection method, the earphone 110 can be a wired earphone (such as through a 3.5mm interface, USB interface connection), wireless earphone (such as Bluetooth earphone), and the specific type of the earphone 110 is not limited by the present application embodiment.
[0041] In the present application embodiment, the earphone 110 has a sound collecting unit in addition to the sound generating unit, that is, a microphone, which is used to collect ambient sound and wearer's voice.
[0042] In some embodiments, the earphone 110 is provided with a first microphone, which is an air conduction microphone, that is, it collects sound transmitted through the air. Optionally, the earphone 110 is provided with a sound collecting hole, and the first microphone collects audio through the sound collecting hole.
[0043] In some embodiments, the earphone 110 is also provided with a second microphone, which is a bone conduction microphone, that is, it collects vibration signals directly through contact with the skin or bones, rather than collecting sound waves transmitted through the air, which is usually realized by a piezoelectric sensor or an accelerometer.
[0044] In the working state, the earphone 110 establishes a data communication link with the terminal 120 through wired or wireless mode (related to the connection mode of the earphone), and then transmits data through the link.
[0045] The terminal 120 is an electronic device with a voice interaction function, which can be a smartphone, a tablet computer, a wearable device, a personal computer, etc. The voice interaction function refers to a function of interaction between a user and the terminal by voice, i.e., the user can give an instruction to the terminal by voice, or ask the terminal a question by voice, and correspondingly, the terminal can also reply in the form of voice.
[0046] In some embodiments, the voice interaction function of the terminal 120 is in a dormant state by default and needs to be woken up by a specific wake-up word. The wake-up word can be a preset wake-up word or a custom wake-up word, and before the first voice interaction, the user needs to read the wake-up word, and the terminal 120 stores the audio of the user reading the wake-up word.
[0047] In some embodiments, the terminal 120 is provided with a microphone, and the terminal collects audio through the microphone and determines whether to wake up the voice interaction function based on the collected audio. In other embodiments, when the earphone 110 is connected to the terminal 120, audio collection and determination of whether to wake up the voice interaction function can also be performed by means of the earphone 110.
[0048] With the scheme provided in the embodiments of the present application, when the user wears an earphone and needs to wake up the voice interaction function of the terminal by means of the earphone, the earphone and the terminal can cooperate with each other and determine whether to wake up the voice interaction service through a multi-level verification mechanism. Figure 1 As shown in FIG. 11, after the user reads the wake-up word, the earphone 110 first performs wake-up word recognition on the collected audio, and when the wake-up word is recognized, the earphone 110 sends the collected audio to the terminal 120, and the terminal 120 further performs voiceprint verification on the audio, i.e., verifies whether the wake-up word is read by the user who previously recorded the wake-up word voice, and wakes up the voice interaction function after the voiceprint verification is passed.
[0049] In the following embodiments, the voice wake-up method is used for the earphone 110 and the terminal 120 in Figure 1 .
[0050] Please refer to Figure 2 , which shows a flowchart of a voice wake-up method provided in an example embodiment of the present application. In this embodiment, the method is used for an earphone and a terminal, and the method can include the following steps:
[0051] Step 201: The earphone collects audio through a first microphone.
[0052] In some embodiments, the first microphone is an air conduction microphone, and the first audio collected by the first microphone includes sounds transmitted through air, including environmental sounds and environmental speech, which can be emitted by a user wearing the earphone and / or by other users in the environment.
[0053] In a possible implementation, when the function of waking up the terminal-side voice interaction service via the earphone is in an open state, the earphone collects audio via the first microphone when a connection is established with the terminal.
[0054] When the function of waking up the terminal-side voice interaction service via the earphone is in a closed state, the earphone does not collect audio via the first microphone even when a connection is established with the terminal.
[0055] Optionally, the user can set the on-off state of the function of waking up the terminal-side voice interaction service via the earphone through the terminal.
[0056] In some embodiments, to reduce power consumption caused by continuous audio collection, the first microphone collects audio according to preset collection parameters, which include a sampling rate and a bit width. The preset collection parameters are lower than the collection parameters of the first microphone in a normal call scenario, for example, a lower sampling rate and a smaller bit width.
[0057] For example, the first microphone collects audio according to a 16 KHz sampling rate and a 16 bit width. The embodiments of this application do not limit the specific collection parameters used by the first microphone.
[0058] In some embodiments, the earphone buffers a preset time length of first audio recently collected by the first microphone, so as to perform subsequent wake-up word recognition and transmission on the buffered first audio. For example, the earphone buffers 2s of first audio recently collected by the first microphone.
[0059] In step 202, the earphone performs wake-up word recognition on the collected first audio to obtain a first wake-up word recognition result, which is used to indicate whether the first audio contains a wake-up word, which is used to wake up the voice interaction function of the terminal.
[0060] In some embodiments, the earphone has a wake-up word recognition function, by which the earphone performs wake-up word recognition on the first audio collected in real time to detect whether the first audio contains a wake-up word.
[0061] Optionally, since the wake-up word is used to wake up the voice interaction function of the terminal, the terminal synchronizes the wake-up word to the earphone after completing the wake-up word setting on the terminal side.
[0062] In one possible implementation, the headphones preprocess the first audio (e.g., noise reduction and enhancement, frame segmentation and windowing, feature extraction, etc.), and then use a model to perform wake-word recognition on the preprocessed result to obtain the first wake-word recognition result. This model can be an HMM (Hidden Markov Model) + DNN (Deep Neural Network) (where the HMM models the phoneme sequence, and the DNN outputs phoneme probabilities; the combination of the two determines whether a wake-word exists), or it can be an end-to-end model, such as TC-ResNet or MHA-TCN (where the input is the preprocessed audio frame sequence or spectrogram, and the output is the probability that the audio segment contains a wake-word). This embodiment does not limit the specific model used for wake-word recognition.
[0063] If the first wake-word recognition result indicates that a wake-word is present, the headphones perform the following step 203.
[0064] If the first wake-word recognition result does not contain a wake-word, the headphones will not send the first audio to the terminal, and correspondingly, the voice interaction function on the terminal side will not be activated.
[0065] Step 203: If the first wake-up word recognition result indicates that the first audio contains a wake-up word, the earphone sends the first audio to the terminal.
[0066] In some embodiments, the headphones send a first audio message containing a wake word to the terminal via a data communication link with the terminal.
[0067] For example, Bluetooth headphones send the first audio signal to the terminal via a Bluetooth link.
[0068] Step 204: The terminal acquires the first audio sent by the earphone. The first audio is collected by the earphone through the first microphone and sent when the first wake-up word recognition result of the first audio contains a wake-up word. The wake-up word is used to wake up the terminal's voice interaction function.
[0069] Correspondingly, the terminal receives the first audio signal sent by the headphones through a data communication link with the headphones.
[0070] For example, the terminal receives the first audio signal sent by the headset through the Bluetooth link between the terminal and the headset.
[0071] Step 205: The terminal performs voiceprint verification on the first audio and obtains the voiceprint verification result.
[0072] Since the first audio contains the wake-up word, which can be read by the local user (i.e., the user with the voice wake-up permission, the terminal user, the earphone wearer) or can be read out by other users in the environment, in order to avoid the wake-up word read out by the user other than the local user from mistakenly waking up the voice interaction function of the terminal, when the first audio is identified to contain the wake-up word, the voice interaction function of the terminal is not directly woken up, but the terminal needs to further perform voiceprint verification on the first audio.
[0073] In the voiceprint verification, the voiceprint of the voice corresponding to the wake-up word in the first audio is verified to be matched with the voiceprint of the pre-recorded wake-up word.
[0074] In a possible implementation, the terminal performs feature extraction on the first audio to obtain voiceprint features representing the identity of the user, where the voiceprint features can include frequency spectrum features obtained through FFT analysis, MFCC features, voice energy features, and the like, and the specific content of the voiceprint features is not limited in the embodiment.
[0075] Further, the terminal performs feature matching on the extracted voiceprint features and the voiceprint features of the pre-recorded wake-up word, and determines a voiceprint verification result based on the feature matching result.
[0076] When the voiceprint verification result indicates that the first audio passes the voiceprint verification, it indicates that the wake-up word in the first audio is issued by the user with the voice wake-up permission. When the voiceprint verification result indicates that the first audio fails the voiceprint verification, it indicates that the wake-up word in the first audio is issued by the user without the voice wake-up permission.
[0077] In step 206, in the case where the voiceprint verification result indicates that the first audio passes the voiceprint verification, the terminal wakes up the voice interaction function.
[0078] Further, when the first audio passes the voiceprint verification, the terminal wakes up the voice interaction function. After the voice interaction function is woken up, the user can continue to perform audio acquisition through the first microphone and transmit the acquired audio to the terminal, so that the terminal identifies and responds to the audio.
[0079] In some embodiments, during the voice interaction, the terminal can perform noise reduction processing on the audio transmitted by the earphone based on the voiceprint, and then identify and respond to the audio after the noise reduction processing, so as to improve the identification and response accuracy.
[0080] The voiceprint noise reduction can filter the voices of the users other than the local user (the user of the pre-recorded wake-up word) in a noisy environment, so as to avoid the interference of the voices of the other users to cause voice recognition errors and then response errors.
[0081] In one possible implementation, the terminal can achieve audio noise reduction through voiceprint feature matching (calculating the similarity between the voiceprint features of the pre-recorded speech and the audio features of the audio transmitted through the headphones, and then filtering out audio features with low similarity), or by using the audio transmitted through the headphones and the voiceprint embedding of the pre-recorded speech as input to a noise reduction model (an AI model trained based on deep learning) to obtain the noise-reduced audio output by the noise reduction model. This application does not limit the specific method of voiceprint noise reduction.
[0082] In some embodiments, if the voiceprint verification result indicates that the first audio has failed the voiceprint verification, the terminal will not activate the voice interaction function.
[0083] In summary, in this embodiment, when the voice interaction function on the terminal side is activated via a worn headset, the headset performs wake-up word recognition on the first audio signal captured by the first microphone. Upon detecting a wake-up word, the headset sends the first audio signal to the terminal, which then performs voiceprint verification on the first audio signal. The voice interaction function is activated only when the first audio signal passes voiceprint verification, thus achieving multi-level verification for voice activation. Since the voice interaction function is only activated when both wake-up word recognition and voiceprint verification are successful simultaneously, the probability of accidental activation of the voice interaction function is reduced.
[0084] Notification of wake-up results
[0085] In some embodiments, in order to enable the user to perceive whether the voice interaction function has been activated, after verifying the voiceprint of the first audio, the terminal sends a wake-up result notification to the headset.
[0086] The wake-up result notification includes a wake-up success notification and a wake-up failure notification. Specifically, when the voiceprint verification result indicates successful voiceprint verification, the terminal sends a wake-up success notification to the headset; when the voiceprint verification result indicates failed voiceprint verification, the terminal sends a wake-up failure notification to the headset.
[0087] Correspondingly, the headset receives the wake-up result notification sent by the terminal and provides a prompt based on the wake-up result notification.
[0088] In one possible implementation, when a wake-up success notification is received, the earphone displays a wake-up success message; when a wake-up failure notification is received, the earphone displays a wake-up failure message, or does not display any message.
[0089] Optionally, successful wake-up notifications may include, but are not limited to, playing a wake-up success message and vibration; failed wake-up notifications may include, but are not limited to, playing a wake-up failure message and vibration. The vibration patterns for successful and failed wake-up notifications differ (e.g., the number of vibrations and their duration).
[0090] Three-stage voice wake-up scheme
[0091] In a possible implementation, after receiving the first audio sent by the earphone, the terminal can directly perform voiceprint verification on the first audio.
[0092] In another possible implementation, in order to further improve the accuracy of the wake-up word recognition, after receiving the first audio sent by the earphone, the terminal can first perform wake-up word recognition on the first audio to obtain a second wake-up word recognition result, and then determine whether voiceprint verification is needed based on the second wake-up word recognition result.
[0093] In some embodiments, in a case where the second wake-up word recognition result indicates that the first audio contains a wake-up word, the terminal performs voiceprint verification on the first audio to obtain a voiceprint verification result.
[0094] In a case where the second wake-up word recognition result indicates that the first audio does not contain a wake-up word, the terminal does not perform voiceprint verification on the first audio, and accordingly, the terminal does not wake up the voice interaction function.
[0095] Optionally, in a case where the second wake-up word recognition result indicates that the first audio does not contain a wake-up word, the terminal sends a wake-up result notification (i.e., a wake-up failure notification) to the earphone.
[0096] In a possible design, the wake-up word recognition schemes of the earphone side and the terminal side are different, and therefore the second wake-up word recognition result obtained by the terminal side can be different from the first wake-up word recognition result obtained by the earphone side.
[0097] In addition, the reliability of the wake-up word recognition scheme of the terminal side is higher than that of the wake-up word recognition scheme of the earphone side, and accordingly, in a case where the first wake-up word recognition result and the second wake-up word recognition result are different, the second wake-up word recognition result is used as the reference.
[0098] Please refer to Figure 3 which shows a flowchart of a three-stage voice wake-up process provided by an example embodiment of the present application. The process can include the following steps:
[0099] In step 301, the earphone performs audio collection through a first microphone.
[0100] The implementation of this step can refer to the above step 201, and this embodiment will not be described here again.
[0101] In step 302, the earphone performs wake-up word recognition on the collected first audio through a first recognition model to obtain a first wake-up word recognition result.
[0102] In some embodiments, the first recognition model is a deep learning model pre-trained and deployed on the earphone side, and an input of the first recognition model is the first audio, and an output of the first recognition model is a probability that the first audio contains the wake-up word.
[0103] Optionally, when the probability that the first audio contains the wake-up word is higher than a probability threshold, the earphone determines that the first audio contains the wake-up word; and when the probability that the first audio contains the wake-up word is lower than the probability threshold, the earphone determines that the first audio does not contain the wake-up word. For example, the probability threshold can be 50%, 70%, 80%, etc.
[0104] Limited by the power consumption and volume of the earphone, compared with the terminal, the earphone side cannot set a processor with higher computing power, and therefore the first recognition model usually has a smaller model parameter amount to run on a processor (such as an MCU) with lower computing power. Correspondingly, the accuracy of the recognition result output by the first recognition model is relatively low, and the probability of identifying a word similar to the pronunciation of the wake-up word as the wake-up word is relatively high.
[0105] In step 303, when the first wake-up word recognition result indicates that the first audio contains the wake-up word, the earphone sends the first audio to the terminal.
[0106] The implementation of this step can refer to step 203 described above, and this embodiment will not be described here in detail.
[0107] In step 304, the terminal acquires the first audio sent by the earphone.
[0108] The implementation of this step can refer to step 204 described above, and this embodiment will not be described here in detail.
[0109] In step 305, the terminal performs wake-up word recognition on the first audio by using a second recognition model to obtain a second wake-up word recognition result, and the recognition accuracy of the second recognition model is higher than that of the first recognition model.
[0110] In order to improve the accuracy of wake-up word recognition, the terminal side is deployed with a second recognition model different from the first recognition model on the earphone side, and after receiving the first audio sent by the earphone, the second recognition model is used to perform secondary wake-up word recognition on the first audio.
[0111] Compared with the earphone, the terminal usually sets a processor with higher computing power, and therefore can run a more complex recognition model to obtain a more accurate wake-up word recognition result.
[0112] In a possible design, the first recognition model and the second recognition model are different in at least one of the model type and the model structure, and the model parameter amount of the second recognition model is greater than that of the first recognition model.
[0113] The model types are different, that is, the model backbone architectures adopted by the first recognition model and the second recognition model are different. Optionally, the performance of the model backbone architecture adopted by the second recognition model is higher than the performance of the model backbone architecture adopted by the first recognition model.
[0114] For example, the first recognition model adopts a TDNN (Time-Delay Neural Network) architecture, and the second recognition model adopts a Transformer or Conformer architecture. The embodiments of the present application do not limit the specific model types of the first recognition model and the second recognition model.
[0115] The model structures are different, that is, the model level setting manners (including the level connection manners and the number of levels, etc.) of the first recognition model and the second recognition model are different. Optionally, in the case where the first recognition model and the second recognition model adopt the same model backbone architecture, the number of model levels of the second recognition model is greater than the number of model levels of the first recognition model.
[0116] For example, the first recognition model and the second recognition model both adopt a Transformer architecture, and the first recognition model is obtained by network layer simplification on the basis of the second recognition model. The embodiments of the present application do not limit the specific model structures of the first recognition model and the second recognition model.
[0117] Since the number of model parameters of the second recognition model is greater than the number of model parameters of the first recognition model, the accuracy of the wake-up word recognition using the second recognition model is higher than the accuracy of the wake-up word recognition using the first recognition model.
[0118] In another possible design, the first recognition model is obtained by quantization of the second recognition model. Model quantization refers to converting parameters such as weights and activation values stored and calculated in high-precision floating-point numbers (such as FP32 and FP16) into data types with low bit width (such as INT8 and INT4), which helps to reduce model size, reduce calculation overhead, and improve inference speed.
[0119] In step 306, in the case where the second wake-up word recognition result indicates that the first audio contains a wake-up word, the terminal performs voiceprint verification on the first audio to obtain a voiceprint verification result.
[0120] The implementation of this step can refer to step 205 described above, and this embodiment will not be repeated here.
[0121] In step 307, in the case where the voiceprint verification result indicates that the first audio passes the voiceprint verification, the terminal wakes up the voice interaction function.
[0122] The implementation of this step can refer to step 206 described above, and will not be repeated here.
[0123] In an illustrative example, as shown in Figure 4 When the voice assistant is woken up using a three-stage voice wake-up scheme, the microphone of the earphone collects audio and uses the earphone-side small model to perform wake-up word recognition on the audio. If the wake-up word recognition result output by the earphone-side small model indicates that there is a wake-up word, the audio is transmitted to the mobile phone. If the wake-up word recognition result output by the earphone-side small model indicates that there is no wake-up word, the voice assistant wake-up process is stopped.
[0124] After the mobile phone receives the audio, the mobile phone-side large model is used to perform wake-up word recognition on the audio. If the wake-up word recognition result output by the mobile phone-side large model indicates that there is no wake-up word, the voice assistant wake-up process is stopped. If the wake-up word recognition result output by the mobile phone-side large model indicates that there is a wake-up word, the audio is further subjected to voiceprint verification.
[0125] When the audio passes the voiceprint verification, the mobile phone wakes up the voice assistant, and one-step voice interaction is performed between the voice assistant and the user.
[0126] In this embodiment, through a two-stage wake-up word recognition mechanism, the first recognition model deployed on the earphone side is used to filter audio that does not obviously contain a wake-up word, thereby reducing the data processing amount on the terminal side. For audio that passes the wake-up word recognition on the earphone side, the terminal further performs wake-up word recognition on the audio using a second recognition model with higher recognition accuracy, thereby filtering audio that contains a sound similar to the pronunciation of the wake-up word, avoiding unnecessary voiceprint recognition, and further reducing the computational amount on the terminal side.
[0127] Dynamic switching between the two-stage voice wake-up scheme and the three-stage voice wake-up scheme
[0128] In actual application, it is found that the recognition accuracy difference between the first recognition model and the second recognition model is small in a quiet environment, but the recognition accuracy difference between the first recognition model and the second recognition model is large in a noisy environment. Moreover, in the case where the confidence of the first wake-up word recognition result output by the first recognition model is high, the probability that the second wake-up word recognition result output by the second recognition model is different from the first wake-up word recognition result is higher than that in the case where the confidence of the first wake-up word recognition result output by the first recognition model is low.
[0129] Therefore, in order to reduce the wake-up delay of the voice assistant, the two-stage voice wake-up scheme and the three-stage voice wake-up scheme described above support dynamic switching.
[0130] In a possible implementation, after the terminal obtains the first audio, the terminal can dynamically determine whether the first audio needs to be subjected to secondary wake-up word recognition according to the confidence of the first wake-up word recognition result and / or the current sound collecting environment.
[0131] Optionally, the confidence of the first wake-up word recognition result is provided by the earphone.
[0132] In a possible implementation, the confidence of the first wake-up word recognition result is a probability that the first audio contains a wake-up word output by the first recognition model. For example, when the probability that the first audio contains a wake-up word output by the first recognition model is 0.95 and the probability that the first audio does not contain a wake-up word is 0.05, the first wake-up word recognition result indicates that the first audio contains a wake-up word, and the confidence is 0.95.
[0133] Optionally, the sound collecting environment is obtained by the terminal identifying the first audio, or is obtained by the earphone identifying the first audio and provided to the terminal.
[0134] In a possible implementation, the earphone / terminal can determine the current sound collecting environment based on time domain features (such as amplitude fluctuation features, signal-to-noise ratio), frequency domain features (such as spectral energy distribution features) and statistical features (such as zero-crossing rate, MFCC and the like) of the first audio. It should be noted that, compared with the terminal performing secondary wake-up word recognition, the calculation amount is lower when determining the sound collecting environment.
[0135] In some embodiments, in addition to sending the first audio to the terminal, the earphone also sends the confidence of the first wake-up word recognition result and / or environment information representing the sound collecting environment to the terminal, so that the terminal determines whether to perform wake-up word recognition on the first audio based on the confidence of the first wake-up word recognition result and / or the sound collecting environment.
[0136] Regarding the switching strategy of the voice wake-up scheme, in a possible implementation, when the confidence of the first wake-up word recognition result is lower than a confidence threshold and / or the sound collecting environment belongs to a first environment, the terminal performs wake-up word recognition on the first audio to obtain a second wake-up word recognition result.
[0137] When the confidence of the first wake-up word recognition result is higher than the confidence threshold and / or the sound collecting environment belongs to a second environment, the first audio is subjected to voiceprint verification to obtain a voiceprint verification result, and the first environment has a higher noise level than the second environment.
[0138] In other words, in a case that the sound environment is quiet and / or the confidence of the first wake-up word recognition result is high, the terminal can not perform the secondary wake-up word recognition because the accuracy of the first wake-up word recognition result is high; in a case that the sound environment is noisy and / or the confidence of the first wake-up word recognition result is low, the terminal needs to perform the secondary wake-up word recognition because the accuracy of the first wake-up word recognition result is low.
[0139] In an illustrative example, when the confidence threshold is set to 0.95, in a case that the confidence of the first wake-up word recognition result is 0.99 and the current sound environment is quiet, the terminal does not need to perform the secondary wake-up word recognition but directly performs the voiceprint verification; in a case that the confidence of the first wake-up word recognition result is 0.85 and the current sound environment is noisy, the terminal needs to perform the secondary wake-up word recognition.
[0140] In this embodiment, the terminal determines whether to perform the secondary wake-up word recognition based on at least one of the confidence of the earphone-side wake-up word recognition result and the current sound environment, which can avoid the waste of processing resources caused by the secondary wake-up word recognition in a case that the earphone-side wake-up word recognition result is relatively accurate, and help reduce the delay of voice wake-up.
[0141] In some embodiments, because the setting of the wake-up word has a great impact on the false trigger rate, the terminal can determine whether to support the dynamic switching between the two-stage voice wake-up scheme and the three-stage voice wake-up scheme based on the wake-up word.
[0142] Optionally, the terminal can determine the false trigger rate of the wake-up word based on at least one of the syllable structure, the semantic feature and the acoustic feature of the wake-up word, and then determine whether to support the dynamic switching between the two-stage voice wake-up scheme and the three-stage voice wake-up scheme based on the false trigger rate.
[0143] In a case that the false trigger rate is lower than a threshold, the terminal determines to support the dynamic switching between the two-stage voice wake-up scheme and the three-stage voice wake-up scheme; in a case that the false trigger rate is higher than the threshold, the terminal determines not to support the dynamic switching between the two-stage voice wake-up scheme and the three-stage voice wake-up scheme, i.e., the terminal needs to perform the secondary wake-up word recognition.
[0144] Optionally, the syllable structure includes at least one of the syllable length and the tone combination of the wake-up word, and the false trigger rate is negatively related to the syllable length (for example, the false trigger rate of a 4-syllable wake-up word is lower than that of a 2-syllable wake-up word), and the false trigger rate is negatively related to the complexity of the tone combination (for example, the false trigger rate of a wake-up word with mixed tones is lower than that of a wake-up word with a single tone).
[0145] Optionally, the semantic feature is used to represent at least one of semantic uniqueness and semantic naturalness of the wake-up word, and the false trigger rate is negatively correlated with the semantic uniqueness of the wake-up word (for example, the false trigger rate of a low-frequency word is lower than that of a high-frequency word), and the false trigger rate is negatively correlated with the semantic naturalness of the wake-up word (for example, the false trigger rate of a natural language wake-up word is higher than that of a special wake-up word).
[0146] Optionally, the acoustic feature includes a consonant feature, and the false trigger rate is negatively correlated with the number of consonants represented by the consonant feature (for example, the false trigger rate of a wake-up word containing consonants such as plosive or fricative is lower than that of a pure vowel wake-up word).
[0147] In some embodiments, the terminal determines the false trigger rate of the wake-up word based on the false trigger rate of the wake-up word in each dimension (at least one of syllable structure, semantic feature, and acoustic feature).
[0148] In this embodiment, the terminal estimates the false trigger rate of the wake-up word based on the multi-dimensional features of the wake-up word, and then starts the dynamic switching of the two-stage and three-stage schemes in the case of low false trigger rate, which helps to reduce power consumption; and in the case of high false trigger rate, the three-stage scheme is used to ensure the accuracy of wake-up word recognition.
[0149] Four-stage voice wake-up scheme
[0150] Since the user wears the earphone, the audio collected by the earphone may contain the local user's voice and the external user's voice, and the wake-up word is usually read by the user wearing the earphone. Therefore, in order to shorten the duration of audio collection by the first microphone and reduce the calculation amount caused by wake-up word recognition, in one possible implementation, a local user sound emission recognition mechanism can be additionally added before audio collection by the first microphone.
[0151] As for the way to recognize the local user's sound emission, in one possible implementation, a second microphone is arranged at the earphone, and the second microphone collects audio in a bone conduction manner, that is, the second microphone is a bone conduction microphone. The earphone recognizes whether there is local user's sound emission based on the audio collected by the second microphone.
[0152] Please refer to Figure 5 which shows a flowchart of a four-stage voice wake-up process provided by one exemplary embodiment of the present application. The process can include the following steps:
[0153] Step 501, the earphone collects audio through the second microphone.
[0154] In a possible implementation, the second microphone continuously collects audio. Different from the first microphone, the second microphone converts a vibration signal conducted through the skin or the bone into an audio signal.
[0155] Therefore, in order to improve the accuracy of the voice recognition of the local user, in some embodiments, the earphone collects audio through the second microphone when it is determined that the earphone wearing state meets the requirement.
[0156] In step 502, the earphone collects audio through the first microphone when it is determined that the local user speaks based on the second audio collected by the second microphone.
[0157] In a possible implementation, since the second microphone only collects a vibration signal conducted through the skin or the bone, and the frequency of the vibration signal generated when the user speaks is usually within a preset frequency range, the earphone determines that the local user speaks, i.e., the user wearing the earphone speaks, when the second audio is collected by the second microphone and the frequency range of the second audio is within the preset frequency range, and then triggers the first microphone to collect audio.
[0158] For example, when the frequency range of the second audio collected by the second microphone is within the frequency range of 500-1500 Hz (only for illustrative purposes), the earphone determines that the local user speaks.
[0159] In some embodiments, after the local user is determined to speak, the first microphone can continuously collect audio or stop collecting audio to reduce power consumption.
[0160] In step 503, the earphone performs wake-up word recognition on the collected first audio by using the first recognition model to obtain a first wake-up word recognition result.
[0161] The implementation of this step can refer to the above-described step 302, and the present embodiment will not be described here in detail.
[0162] In step 504, when the first wake-up word recognition result indicates that the first audio contains a wake-up word, the earphone sends the first audio to the terminal.
[0163] The implementation of this step can refer to the above-described step 303, and the present embodiment will not be described here in detail.
[0164] In step 505, the terminal obtains the first audio sent by the earphone.
[0165] The implementation of this step can refer to the above-described step 304, and the present embodiment will not be described here in detail.
[0166] Step 506: The terminal uses the second recognition model to identify the wake word of the first audio and obtains the second wake word recognition result. The recognition accuracy of the second recognition model is higher than that of the first recognition model.
[0167] The implementation method of this step can refer to step 305 above, and will not be repeated here in this embodiment.
[0168] Step 507: If the second wake-up word recognition result indicates that the first audio contains a wake-up word, the terminal performs voiceprint verification on the first audio to obtain the voiceprint verification result.
[0169] The implementation method of this step can refer to step 306 above, and will not be repeated here in this embodiment.
[0170] Step 508: If the voiceprint verification result indicates that the first audio has passed the voiceprint verification, the terminal activates the voice interaction function.
[0171] The implementation method of this step can refer to step 307 above, and will not be repeated here.
[0172] In an illustrative example, such as Figure 6 As shown, when activating the voice assistant using a four-segment voice wake-up scheme, the earphones first collect audio through a bone conduction microphone. When the user's voice is recognized based on the audio collected by the bone conduction microphone, the earphones collect audio through an air conduction microphone and use a small model on the earphones to recognize the wake word.
[0173] If the wake-up word recognition result output by the earpiece mini-model indicates the presence of a wake-up word, the audio is transmitted to the phone. If the wake-up word recognition result output by the earpiece mini-model indicates the absence of a wake-up word, the voice assistant wake-up process is stopped.
[0174] After receiving the audio, the phone uses a large-scale model on the phone to identify the wake word. If the wake word recognition result output by the large-scale model on the phone indicates that no wake word exists, the voice assistant wake-up process stops. If the wake word recognition result output by the large-scale model on the phone indicates that a wake word exists, the audio is further verified using voiceprint verification.
[0175] When the audio passes the voiceprint verification, the phone wakes up the voice assistant and conducts a one-step voice interaction with the user.
[0176] In this embodiment, the earphone side uses a bone conduction microphone to perform pre-recognition of the user's voice. When the user's voice is recognized, the air conduction microphone is used to collect audio, which helps to shorten the continuous audio collection time of the air conduction microphone and thus helps to reduce the computational load of wake word recognition.
[0177] It should be noted that when the four-segment voice wake-up scheme is used, the secondary wake-up word recognition on the terminal side can also not be performed. The process of determining whether to perform secondary wake-up word recognition can refer to the dynamic switching scheme of the two-segment voice wake-up scheme and the three-segment voice wake-up scheme, which will not be repeated here.
[0178] Optimization of voiceprint verification process
[0179] Since the audio collected by the first microphone includes environmental sound, a noisy listening environment can affect the accuracy of voiceprint verification. Since the bone conduction microphone only collects vibration signals transmitted through the skin or bones and is not affected by environmental sound, the terminal can use the audio collected by the bone conduction microphone to optimize the first audio collected by the first microphone to improve the accuracy of subsequent voiceprint verification.
[0180] Please refer to Figure 7 , Figure 7 is a flowchart of a voiceprint verification process according to an example embodiment of the present application. The process can include the following steps:
[0181] Step 701, the earphone collects audio through the second microphone.
[0182] Step 702, the earphone sends the second audio collected by the second microphone to the terminal.
[0183] In one possible implementation, in the case where the first audio is identified to contain a wake-up word, the earphone sends the second audio collected by the second microphone to the terminal.
[0184] Since a noisy listening environment can affect subsequent voiceprint verification, and a quiet listening environment has less impact on voiceprint verification, in one possible implementation, in the case where the current listening environment belongs to the first environment, the earphone sends the second audio to the terminal; in the case where the current listening environment belongs to the second environment, the earphone does not send the second audio to the terminal.
[0185] Wherein, the related content of identifying the current listening environment, the first environment and the second environment can refer to the above embodiments, which will not be repeated here.
[0186] Step 703, in the case where the second audio sent by the earphone is obtained, the terminal optimizes the first audio based on the second audio to obtain the optimized first audio.
[0187] In some embodiments, in the audio optimization process, the terminal separates the pure local user speech and environmental noise in the first audio by taking the second audio as a reference, and the separated pure local user speech is the optimized first audio.
[0188] In some embodiments, in order to guarantee the audio optimization effect, the terminal performs time alignment processing on the first audio and the second audio before the audio optimization, so as to guarantee that the second audio collected at the same collection time is used to perform the audio optimization on the first audio.
[0189] In a possible implementation, the terminal performs frequency domain mask processing on the first audio by using the second audio to obtain the optimized first audio.
[0190] Since the second audio represents the time interval in which the local user actually speaks, the terminal can generate a mask based on the audio features of the second audio in the frequency domain, and then perform mask processing on the first audio by using the mask to suppress the environmental noise in the first audio.
[0191] Optionally, after the terminal performs signal preprocessing and frequency domain conversion on the second audio, a binary mask is generated based on the frequency point energy of the first audio and the second audio at the same time, or a soft mask is generated based on the signal-to-noise ratio of the first audio and the second audio, and then the mask is applied to the amplitude spectrum of the first audio (that is, the amplitude spectrum in the amplitude range represented by the mask is retained), and the processed amplitude spectrum is subjected to frequency spectrum reconstruction and inverse STFT (short-time Fourier transform) processing to obtain the optimized first audio.
[0192] In another possible implementation, the terminal is provided with a noise filtering model, which is pre-trained based on sample airborne conduction audio (corresponding to the first audio), sample bone conduction audio (corresponding to the second audio), and sample optimized audio (corresponding to the optimized first audio). In the application process, the terminal inputs the first audio and the second audio into the noise filtering model to obtain the optimized first audio output by the noise filtering model.
[0193] Of course, in addition to the above-mentioned manner, the terminal can also obtain the optimized first audio by performing adaptive filtering processing and other manners on the first audio by using the second audio, and the embodiments of the present application are not limited thereto.
[0194] In step 704, the terminal performs voiceprint verification on the optimized first audio to obtain a voiceprint verification result.
[0195] The process of performing voiceprint verification on the optimized first audio can refer to the process of performing voiceprint verification on the first audio, and details are not repeated herein.
[0196] It should be noted that the optimization scheme of the voiceprint verification process can be combined with each of the above-mentioned embodiments to obtain a new embodiment, for example, the optimization scheme of the voiceprint verification process can be applied to the two-stage, three-stage or four-stage voice wake-up scheme, and details are not repeated herein.
[0197] In this embodiment, based on the sound collection characteristics of the bone conduction microphone and the air conduction microphone, the terminal optimizes the first audio by using the second audio collected at the same time as the first audio, which can reduce the negative influence of the ambient sound on the voiceprint verification of the first audio, and help to improve the accuracy of voiceprint verification.
[0198] It should be noted that, in each of the above embodiments, the steps taken by the earphone as the execution subject can be implemented alone as an earphone-side voice wake-up method, and the steps taken by the terminal as the execution subject can be implemented alone as a terminal-side voice wake-up method, and the present application embodiments will not be repeated.
[0199] Please refer to Figure 8 which shows a structural block diagram of a voice wake-up device provided by an example embodiment of the present application. The device includes:
[0200] The acquisition module 801 is configured to acquire first audio sent by an earphone, the first audio being collected by the earphone through a first microphone and sent when a first wake-up word recognition result of the first audio indicates that the wake-up word is contained, the wake-up word being used to wake up a voice interaction function of the terminal;
[0201] The voiceprint verification module 802 is configured to perform voiceprint verification on the first audio to obtain a voiceprint verification result.
[0202] The wake-up module 803 is configured to wake up the voice interaction function in a case where the voiceprint verification result indicates that the first audio passes the voiceprint verification.
[0203] Optionally, the device further includes a wake-up word identification module configured to:
[0204] perform wake-up word identification on the first audio to obtain a second wake-up word identification result.
[0205] The voiceprint verification module 802 is configured to:
[0206] perform voiceprint verification on the first audio to obtain the voiceprint verification result in a case where the second wake-up word identification result indicates that the first audio contains the wake-up word.
[0207] Optionally, the first wake-up word identification result is obtained by the earphone performing wake-up word identification on the first audio through a first identification model.
[0208] The wake-up word identification module is configured to:
[0209] perform wake-up word identification on the first audio through a second identification model to obtain the second wake-up word identification result, the identification accuracy of the second identification model being higher than that of the first identification model.
[0210] Optionally, the first recognition model and the second recognition model differ in at least one of model type and model structure, and the second recognition model has a larger number of model parameters than the first recognition model.
[0211] and / or,
[0212] The first recognition model is quantized by the second recognition model.
[0213] Optionally, the wake-up word recognition module is configured to:
[0214] In a case where the confidence of the first wake-up word recognition result is lower than a confidence threshold and / or the sound collection environment belongs to a first environment, perform wake-up word recognition on the first audio to obtain a second wake-up word recognition result.
[0215] Optionally, the voiceprint verification module 802 is configured to:
[0216] In a case where the confidence of the first wake-up word recognition result is higher than the confidence threshold and / or the sound collection environment belongs to a second environment, perform voiceprint verification on the first audio to obtain a voiceprint verification result, the first environment having a higher level of noise than the second environment.
[0217] Optionally, the wake-up module 803 is configured to:
[0218] In a case where the second wake-up word recognition result indicates that the first audio does not contain the wake-up word, do not wake up the voice interaction function.
[0219] Optionally, the earphone has a second microphone, and the second microphone is configured to collect audio in a bone conduction manner.
[0220] The voiceprint verification module 802 is configured to:
[0221] In a case where the second audio sent by the earphone is obtained, perform audio optimization on the first audio based on the second audio to obtain optimized first audio, the second audio being collected by the second microphone.
[0222] Perform voiceprint verification on the optimized first audio to obtain the voiceprint verification result.
[0223] Optionally, the voiceprint verification module 802 is configured to:
[0224] Perform frequency domain mask processing on the first audio using the second audio to obtain the optimized first audio.
[0225] input the first audio and the second audio into a noise filtering model to obtain the first audio filtered by the noise filtering model.
[0226] Optionally, the wake-up module 803 is configured to:
[0227] In a case where the voiceprint verification result indicates that the first audio fails to pass voiceprint verification, the voice interaction function is not woken up.
[0228] Optionally, the apparatus further comprises a notification module configured to:
[0229] The wake-up result notification is sent to the earphone, so that the earphone gives a prompt based on the wake-up result notification.
[0230] Please refer to Figure 9 which shows a structural block diagram of a voice wake-up apparatus provided by another example embodiment of the present application. The apparatus comprises:
[0231] A first collection module 901 is configured to collect audio through a first microphone.
[0232] A wake-up word recognition module 902 is configured to recognize a wake-up word from the collected first audio, to obtain a first wake-up word recognition result, the first wake-up word recognition result being used to indicate whether the first audio contains a wake-up word, the wake-up word being used to wake up a voice interaction function of a terminal.
[0233] A sending module 903 is configured to send the first audio to the terminal in a case where the first wake-up word recognition result indicates that the first audio contains the wake-up word, the terminal being configured to wake up the voice interaction function in a case where the first audio passes voiceprint verification.
[0234] Optionally, the wake-up word recognition module 902 is configured to:
[0235] The first audio is recognized for a wake-up word through a first recognition model, to obtain the first wake-up word recognition result, the first recognition model being different from a second recognition model used by the terminal to recognize a wake-up word from the first audio.
[0236] Optionally, the first recognition model and the second recognition model differ in at least one of a model type and a model structure, and a model parameter quantity of the second recognition model is greater than a model parameter quantity of the first recognition model.
[0237] and / or,
[0238] The first recognition model is obtained by quantization of the second recognition model.
[0239] Optionally, the sending module 903 is further configured to:
[0240] send, to the terminal, the confidence of the first wake-up word recognition result and / or the environment information representing the sound collection environment, so that the terminal determines whether to perform wake-up word recognition on the first audio based on the confidence of the first wake-up word recognition result and / or the sound collection environment.
[0241] Optionally, the earphone has a second microphone, and the second microphone is configured to collect audio in a bone conduction manner.
[0242] The apparatus further includes a second collecting module configured to:
[0243] collect audio through the second microphone.
[0244] The first collecting module 901 is configured to:
[0245] collect audio through the first microphone when the second audio collected by the second microphone is identified to represent the sound of the local user.
[0246] Optionally, the sending module 903 is further configured to:
[0247] send, to the terminal, the second audio, so that the terminal optimizes the first audio based on the second audio and performs voiceprint verification on the optimized first audio.
[0248] Optionally, the apparatus further includes a prompting module configured to:
[0249] receive the wake-up result notification sent by the terminal.
[0250] prompt based on the wake-up result notification.
[0251] In the embodiments of the present application, when the earphone is used to wake up the voice interaction function of the terminal, the earphone performs wake-up word recognition on the first audio collected by the first microphone, and when the first audio is identified to contain a wake-up word, the earphone sends the first audio to the terminal, and the terminal further performs voiceprint verification on the first audio, so that the voice interaction function is woken up when the first audio passes the voiceprint verification, and multi-level verification of voice wake-up is achieved. Since the voice interaction function is woken up only when the wake-up word recognition and the voiceprint verification are both successful, the probability of false wake-up of the voice interaction function can be reduced.
[0252] It should be noted that the apparatus provided in the above examples is only exemplarily described based on the division of the above functional modules. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above examples belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be described here.
[0253] Referring to Figure 10 , Figure 10 is a structural block diagram of an earphone provided in an example embodiment of the present application. The earphone includes a processor 1010, a memory 1020, a sound producing unit 1030, and a microphone 1040.
[0254] Optionally, the processor 1010 is connected to various parts in the earphone by various interfaces and lines, and performs various functions of the earphone and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1020, and calling data stored in the memory 1020.
[0255] Optionally, the processor 1010 can be a low-power processor such as a microcontroller unit (MCU).
[0256] The memory 1020 can include a random access memory (RAM) and can also include a read-only memory (ROM). Optionally, the memory 1020 includes a non-transitory computer-readable storage medium. The memory 1020 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 1020 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area can store data created according to the use of the earphone (such as audio data or wake-up word data), etc.
[0257] The sound producing unit 1030 is a unit for producing sound, which can be a moving coil unit, a moving iron unit, a planar magnetic unit, an electrostatic unit or a bone conduction unit.
[0258] The microphone 1040 is a unit for collecting sound. In the embodiments of the present application, the microphone 1040 can include an air conduction microphone and a bone conduction microphone.
[0259] In addition, those skilled in the art can understand that the structure of the earphone shown in the above-mentioned drawings does not constitute a limitation on the earphone, and the earphone can include more (such as a communication component for communicating with an external device, a sensor unit, a touch unit, and the like) or fewer components than those shown in the drawings, or some components are combined, or different component arrangements.
[0260] Referring to Figure 11 , Figure 11 is a structural block diagram of a terminal provided by an exemplary embodiment of the present application. The terminal can include one or more of the following components: a processor 1110 and a memory 1120.
[0261] Optionally, the processor 1110 connects various parts in the entire electronic device by using various interfaces and lines, and performs various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1120, and calling data stored in the memory 1120. Optionally, the processor 1110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA).
[0262] The processor 1110 can be integrated with one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), a neural-network processing unit (NPU), and a baseband chip. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; the NPU is used to implement artificial intelligence (AI) functions; and the baseband chip is used to process wireless communication. It can be understood that the above-mentioned baseband chip can also not be integrated into the processor 1110, but be implemented by a separate chip.
[0263] The memory 1120 can include a random access memory (RAM) and also include a read-only memory (ROM). Optionally, the memory 1120 includes a non-transitory computer-readable storage medium. The memory 1120 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 1120 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for at least one function, instructions for implementing each of the above-described method embodiments, and the like; and the data storage area can store data created according to the use of the electronic device, and the like.
[0264] In addition, those skilled in the art can understand that the structure of the terminal shown in the above-described drawings does not constitute a limitation on the terminal, and the terminal can include more (such as a communication component for communicating with an external device, such as Bluetooth) or fewer components than those shown, or combine certain components, or different component arrangements.
[0265] The embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores at least one computer instruction, the at least one computer instruction is used to be executed by a processor to realize the voice wake-up method as described in the above embodiment.
[0266] In another aspect, the embodiment of the present application provides a computer program product, the computer program product includes computer instructions, and a processor executes the computer instructions to realize the voice wake-up method as described in the above embodiment.
[0267] Those skilled in the art should realize that, in one or more examples described above, the functions described in the embodiments of the present application can be realized by hardware, software, firmware or any combination thereof. When realized by software, these functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on a computer readable medium. The computer readable medium includes a computer storage medium and a communication medium, where the communication medium includes any medium that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0268] The above description is only optional embodiments of the present application, and does not limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A voice wake-up method, characterized in that, The method is used in a terminal, and the method includes: The system acquires a first audio signal sent by the earphone, which is collected by the earphone through a first microphone and is sent when the first wake-up word recognition result of the first audio signal contains a wake-up word. The wake-up word is used to activate the voice interaction function of the terminal. The first audio file is subjected to voiceprint verification to obtain the voiceprint verification result. If the voiceprint verification result indicates that the first audio has passed voiceprint verification, the voice interaction function is activated.
2. The method according to claim 1, characterized in that, The method further includes: Perform wake-up word recognition on the first audio to obtain the second wake-up word recognition result; The step of performing voiceprint verification on the first audio to obtain the voiceprint verification result includes: If the second wake-word recognition result indicates that the first audio contains the wake-word, then the first audio is subjected to voiceprint verification to obtain the voiceprint verification result.
3. The method according to claim 2, characterized in that, The first wake-word recognition result is obtained by the earphones through the first recognition model to perform wake-word recognition on the first audio. The step of performing wake-up word recognition on the first audio to obtain the second wake-up word recognition result includes: The second recognition model is used to identify the wake word of the first audio, and the recognition result of the second wake word is obtained. The recognition accuracy of the second recognition model is higher than that of the first recognition model.
4. The method according to claim 3, characterized in that, The first recognition model and the second recognition model differ in at least one dimension of model type and model structure, and the number of model parameters of the second recognition model is greater than the number of model parameters of the first recognition model. And / or, The first recognition model is obtained by quantizing the second recognition model.
5. The method according to claim 2, characterized in that, The step of performing wake-up word recognition on the first audio to obtain the second wake-up word recognition result includes: If the confidence level of the first wake-up word recognition result is lower than the confidence level threshold, and / or the audio reception environment belongs to the first environment, the first audio is subjected to wake-up word recognition to obtain the second wake-up word recognition result.
6. The method according to claim 5, characterized in that, The step of performing voiceprint verification on the first audio to obtain the voiceprint verification result further includes: If the confidence level of the first wake-word recognition result is higher than the confidence level threshold, and / or the sound reception environment belongs to the second environment, the first audio is subjected to voiceprint verification to obtain the voiceprint verification result, wherein the noise level of the first environment is higher than the noise level of the second environment.
7. The method according to claim 2, characterized in that, The method further includes: If the second wake-up word recognition result indicates that the first audio does not contain the wake-up word, the voice interaction function will not be activated.
8. The method according to any one of claims 1 to 7, characterized in that, The earphone has a second microphone, which uses bone conduction to collect audio. The method further includes: Upon acquiring the second audio transmitted by the earphone, the first audio is optimized based on the second audio to obtain the optimized first audio, and the second audio is acquired by the second microphone; The step of performing voiceprint verification on the first audio to obtain the voiceprint verification result includes: The optimized first audio is subjected to voiceprint verification to obtain the voiceprint verification result.
9. The method according to claim 8, characterized in that, The step of optimizing the first audio based on the second audio to obtain the optimized first audio includes at least one of the following: The first audio is subjected to frequency domain masking processing using the second audio to obtain the optimized first audio. The first audio and the second audio are input into a noise filtering model to obtain the optimized first audio output by the noise filtering model.
10. The method according to any one of claims 1 to 7, characterized in that, The method further includes: If the voiceprint verification result indicates that the first audio has failed the voiceprint verification, the voice interaction function will not be activated.
11. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Send a wake-up result notification to the earphone so that the earphone can provide a prompt based on the wake-up result notification.
12. A voice wake-up method, characterized in that, The method is used for headphones, and the method includes: Audio is captured using the first microphone; The first audio is subjected to wake word recognition to obtain a first wake word recognition result. The first wake word recognition result is used to characterize whether the first audio contains a wake word. The wake word is used to wake up the voice interaction function of the terminal. If the first wake-up word recognition result indicates that the first audio contains the wake-up word, the first audio is sent to the terminal, and the terminal is used to wake up the voice interaction function if the first audio passes the voiceprint verification.
13. The method according to claim 12, characterized in that, The step of performing wake-up word recognition on the acquired first audio to obtain the first wake-up word recognition result includes: The first recognition model is used to identify the wake word of the first audio, and the first wake word recognition result is obtained. The first recognition model is different from the second recognition model used by the terminal when it identifies the wake word of the first audio.
14. The method according to claim 13, characterized in that, The first recognition model and the second recognition model differ in at least one dimension of model type and model structure, and the number of model parameters of the second recognition model is greater than the number of model parameters of the first recognition model. And / or, The first recognition model is obtained by quantizing the second recognition model.
15. The method according to claim 13, characterized in that, The method further includes: The confidence level of the first wake-up word recognition result is sent to the terminal, and / or environmental information of the receiving environment is represented, so that the terminal determines whether to perform wake-up word recognition on the first audio based on the confidence level of the first wake-up word recognition result, and / or the receiving environment.
16. The method according to any one of claims 12 to 15, characterized in that, The earphone has a second microphone, which uses bone conduction to collect audio. The method further includes: Audio is captured via the second microphone; The audio acquisition via the first microphone includes: When the user's voice is identified based on the second audio collected by the second microphone, audio is collected through the first microphone.
17. The method according to claim 16, characterized in that, The method further includes: The second audio is sent to the terminal so that the terminal can optimize the first audio based on the second audio and perform voiceprint verification on the optimized first audio.
18. The method according to any one of claims 12 to 15, characterized in that, The method further includes: Receive the wake-up result notification sent by the terminal; A notification will be sent based on the wake-up result.
19. A voice wake-up device, characterized in that, The device includes: The acquisition module is used to acquire a first audio signal sent by the earphone. The first audio signal is collected by the earphone through a first microphone and sent when the first wake-up word recognition result of the first audio signal contains a wake-up word. The wake-up word is used to wake up the voice interaction function of the terminal. The voiceprint verification module is used to perform voiceprint verification on the first audio and obtain the voiceprint verification result. A wake-up module is used to wake up the voice interaction function when the voiceprint verification result indicates that the first audio has passed voiceprint verification.
20. A voice wake-up device, characterized in that, The device includes: The first acquisition module is used to acquire audio through the first microphone; The wake-up word recognition module is used to recognize the wake-up word in the first audio and obtain the first wake-up word recognition result. The first wake-up word recognition result is used to characterize whether the first audio contains a wake-up word. The wake-up word is used to wake up the voice interaction function of the terminal. The sending module is used to send the first audio to the terminal when the first wake-up word recognition result indicates that the first audio contains the wake-up word, and the terminal is used to wake up the voice interaction function when the first audio passes the voiceprint verification.
21. A terminal, characterized in that, The terminal includes a processor and a memory, the memory storing at least one computer instruction, which is loaded and executed by the processor to implement the voice wake-up method as described in any one of claims 1 to 11.
22. An earphone, characterized in that, The headset includes a processor, a memory, a sound-generating unit, and a microphone. The memory stores at least one computer instruction, which is loaded and executed by the processor to implement the voice wake-up method as described in any one of claims 12 to 18.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer instruction, which is executed by a processor to implement the voice wake-up method as described in any one of claims 1 to 18.
24. A computer program product, characterized in that, The computer program product includes computer instructions, which, when executed by the processor, implement the voice wake-up method as described in any one of claims 1 to 18.
Citation Information
Cited By
Speech recognition method and system of Bluetooth headset
CN122024730A