Voice wake-up system and wake-up method thereof, and wearable device
The bone conduction microphone detects voice activity and combines the sonic microphone to recognize wake-up words, which solves the problems of high power consumption and high false wake-up rate when wake-up of existing wearable devices, and realizes an efficient and low-power voice wake-up system.
Patent Information
- Application Number
- CN202411975195.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-09
AI Technical Summary
Existing wearable devices need to continuously detect voice when wake up, resulting in large power consumption and it is difficult to avoid false wake-up.
The bone conduction microphone is used to detect voice activity to realize first-level voice wake-up, and after detecting voice activity, voice recognition is performed through the sonic microphone to detect wake-up words, realizing second-level voice wake-up.
While ensuring wake-up sensitivity, it effectively avoids false wake-up, reducing device power consumption.
Smart Images

Figure CN119964603A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and more specifically, to a voice wake-up system and a wake-up method thereof, and a wearable device. Background Art
[0002] For most existing wearable devices, multiple acoustic microphones are usually sampled for sound collection. For example, smart glasses can be equipped with multiple acoustic microphones on the frame or temples. Currently, most of these types of wearable devices need to continuously detect and maintain the voice wake-up function when waking up, and need to frequently switch between standby and working modes, which results in high power consumption of the device. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides a voice wake-up system and a wake-up method thereof, and a wearable device, which utilizes a bone conduction microphone to detect voice activity to achieve primary voice wake-up, and performs voice recognition on the collected sound signal only after the voice activity is detected, so as to perform secondary voice wake-up of the wake-up word. This enables the embodiment of the present invention to effectively avoid false wake-ups while ensuring the wake-up sensitivity, thereby reducing the power consumption of the device.
[0004] In a first aspect, an embodiment of the present invention provides a voice wake-up system, the voice wake-up system comprising:
[0005] at least one acoustic microphone;
[0006] at least one bone conduction microphone; and
[0007] The controller is used to detect a voice activity signal according to the bone conduction microphone, perform voice recognition on the collected sound signal in response to detecting the voice activity signal, and wake up the corresponding device in response to detecting a wake-up word.
[0008] Furthermore, the at least one acoustic wave microphone is distributed on both sides of the wearable device, and the bone conduction microphone is arranged in a predetermined area of the wearable device, wherein at least a part of the predetermined area contacts the wearer after the wearable device is worn.
[0009] Furthermore, the controller is used to perform speech recognition on the sound signal collected by the bone conduction microphone based on a pre-trained first speech recognition model in response to detecting a voice activity signal, so as to perform wake-up word detection.
[0010] Furthermore, the controller is used to control the activation of at least one acoustic wave microphone in response to detecting a voice activity signal, and perform speech recognition on the sound signal collected by the acoustic wave microphone based on a pre-trained second speech recognition model to perform wake-up word detection.
[0011] Furthermore, the controller is used to control the activation of at least one acoustic wave microphone in response to detecting a voice activity signal, and to fuse multiple sound signals collected by the bone conduction microphone and each of the acoustic wave microphones based on a pre-trained third speech recognition model to perform wake-up word detection.
[0012] Furthermore, the controller is used to control the activation of at least one acoustic wave microphone in response to detecting a voice activity signal, and perform speech recognition on the sound signal collected by the bone conduction microphone based on a pre-trained first speech recognition model; in response to the first speech recognition model detecting a wake-up word, perform speech recognition on the sound signal collected by each of the acoustic wave microphones based on a pre-trained second speech recognition model to verify the wake-up word detection.
[0013] Furthermore, the controller is used to perform speech recognition on the sound signal collected by the acoustic microphone based on a pre-trained second speech recognition model in response to the first speech recognition model detecting the wake-up word and the detection confidence is less than a predetermined value to verify the wake-up word detection.
[0014] Furthermore, the voice wake-up system further includes:
[0015] The master control is used to perform audio processing on the received sound signal after the controlled wake-up;
[0016] The controller is used to wake up the main control in response to detecting a wake-up word.
[0017] Furthermore, the main control is deployed with a fourth speech recognition model, and the main control is used to verify the wake-up word detection result of the controller according to the fourth speech recognition model.
[0018] In a second aspect, an embodiment of the present invention provides a wearable device, the wearable device comprising:
[0019] The device body; and
[0020] Voice wake-up system as described above.
[0021] Furthermore, the wearable device is smart glasses, the acoustic microphones in the voice wake-up system are distributed on the temples and / or frames, and the bone conduction microphones in the voice wake-up system are arranged in the area where the nose pads are located.
[0022] Furthermore, the acoustic wave microphones are symmetrically distributed on the temples and / or frames on both sides.
[0023] In a third aspect, an embodiment of the present invention provides a wake-up method for a voice wake-up system, wherein the voice wake-up system includes at least one acoustic wave microphone and at least one bone conduction microphone, and the method includes:
[0024] Activating the bone conduction microphone to detect a voice activity signal;
[0025] In response to the bone conduction microphone detecting a voice activity signal, collecting a sound signal from a target microphone, wherein the target microphone includes at least one of the acoustic wave microphone and / or the bone conduction microphone;
[0026] Performing speech recognition on the sound signal to detect a wake-up word;
[0027] In response to detecting the wake-up word, waking up the corresponding device.
[0028] The voice wake-up system of the embodiment of the present invention includes at least one acoustic microphone, at least one bone conduction microphone and a controller, wherein the controller is used to detect a voice activity signal according to the bone conduction microphone, perform voice recognition on the collected sound signal in response to the detection of the voice activity signal, and wake up the corresponding device in response to the detection of the wake-up word. Therefore, the embodiment of the present invention can use the bone conduction microphone voice activity detection to achieve the first-level voice wake-up, and perform voice recognition on the collected sound signal after the voice activity is detected to perform the second-level voice wake-up of the wake-up word, which enables the embodiment of the present invention to effectively avoid false wake-ups while ensuring the wake-up sensitivity, and reduce the power consumption of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0030] Figure 1 is a schematic diagram of a voice wake-up system according to an embodiment of the present invention;
[0031] Figure 2 is a schematic diagram of a wearable device according to an embodiment of the present invention;
[0032] Figure 3 is a schematic diagram of another wearable device according to an embodiment of the present invention;
[0033] Figure 4 is a flow chart of a wake-up method of a voice wake-up system according to an embodiment of the present invention;
[0034] Figure 5 is a flow chart of another wake-up method of a voice wake-up system according to an embodiment of the present invention;
[0035] Figure 6 is a schematic diagram of a wake-up process of a voice wake-up system according to an embodiment of the present invention;
[0036] Figure 7is a flow chart of another wake-up method of a voice wake-up system according to an embodiment of the present invention;
[0037] Figure 8 is a schematic diagram of a wake-up process of another voice wake-up system according to an embodiment of the present invention;
[0038] Fig. 9 is a flow chart of another wake-up method of a voice wake-up system according to an embodiment of the present invention;
[0039] Fig.10 is a schematic diagram of a wake-up process of another voice wake-up system according to an embodiment of the present invention;
[0040] Fig.11 is a flow chart of another wake-up method of a voice wake-up system according to an embodiment of the present invention;
[0041] Fig.12 is a schematic diagram of a wake-up process of another voice wake-up system according to an embodiment of the present invention;
[0042] Fig.13 is a flow chart of another wake-up method of a voice wake-up system according to an embodiment of the present invention;
[0043] Fig.14 4 is a schematic diagram of a wake-up process of another voice wake-up system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.
[0045] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.
[0046] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".
[0047] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.
[0048] The solutions described in this specification and in the examples, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0049] Figure 1 is a schematic diagram of a voice wake-up system according to an embodiment of the present invention. The voice wake-up system according to an embodiment of the present invention includes at least one acoustic wave microphone, at least one bone conduction microphone, and a controller. In this embodiment, a voice wake-up system including four acoustic wave microphones and one bone conduction microphone is taken as an example. It should be understood that this embodiment does not limit the number of acoustic wave microphones and bone conduction microphones in the voice wake-up system, and the number of each type of microphone can be set according to the specific application scenario and / or the appearance and wearing method of the wearable device.
[0050] like Figure 1 As shown, the voice wake-up system of this embodiment includes acoustic wave microphones Mic11-Mic14, a bone conduction microphone 15 and a controller 16. The acoustic wave microphones Mic11-Mic14 capture sound signals after being controlled to start. The bone conduction microphone 15 senses bone vibration to capture sound after being controlled to start.
[0051] A bone conduction microphone is a device that uses bone conduction technology to pick up and transmit sound. It transmits sound vibrations through the skull of the person wearing the device, thereby capturing sound signals. Since the bone conduction microphone receives the voiceprint vibration signal when the wearer speaks, but cannot receive the sounds of the surrounding environment and other speakers, the bone conduction microphone can more accurately detect whether the wearer is speaking, that is, whether there is voice activity, avoiding interference from the surrounding environment or the voices of other speakers, effectively avoiding false wake-ups while ensuring wake-up sensitivity, and reducing device power consumption.
[0052] Further, the voice wake-up system activates the bone conduction microphone 15 to detect voice activity information, and the controller 16 collects the sound signal of the target microphone in response to the bone conduction microphone detecting the voice activity signal, performs voice recognition on the collected sound signal to detect the wake-up word, and wakes up the corresponding device in response to detecting the wake-up word. The target microphone includes at least one acoustic wave microphone and / or a bone conduction microphone.
[0053] Therefore, this embodiment can use the bone conduction microphone voice activity detection to achieve primary voice wake-up, and perform voice recognition on the collected sound signal after detecting the voice activity to perform secondary voice wake-up of the wake-up word. This allows the embodiment of the present invention to effectively avoid false wake-ups while ensuring wake-up sensitivity, thereby reducing device power consumption.
[0054] In an optional implementation, the secondary voice wake-up of this embodiment is implemented based on the bone conduction microphone 15, that is, the above-mentioned target microphone includes a bone conduction microphone. Further, a secondary voice wake-up module is deployed in the controller 16. Among them, the secondary voice wake-up module performs voice recognition on the sound signal collected by the bone conduction microphone 15 according to the pre-trained first voice recognition model to detect the wake-up word. Among them, the first voice recognition model is acquired based on the training of the sound sample set collected by the bone conduction microphone. Among them, there are at least some sound samples carrying the wake-up word in the sound sample set. It should be understood that the first voice recognition model of this embodiment can be any network structure that can realize speech recognition (Automatic Speech Recognition, ASR), such as convolutional neural network (CNN), recurrent neural network (RNN) and its variant long short-term memory network (LSTM) and gated recurrent unit (GRU), Transformer model, end-to-end model, or hybrid model, etc., or other lightweight network models that can realize speech recognition, which will not be illustrated one by one in this embodiment.
[0055] In another optional implementation, the secondary voice wake-up of this embodiment is implemented based on at least one sonic microphone Mic, that is, the above-mentioned target microphone includes at least one sonic microphone Mic. Furthermore, a secondary voice wake-up module is deployed in the controller 16. Among them, the controller 16 is used to control the start-up of at least one sonic microphone to continuously collect sound signals in response to the detection of a voice activity signal, and perform voice recognition on the sound signal collected by the sonic microphone according to a pre-trained second voice recognition model through the secondary voice wake-up module to perform wake-up word detection. In this embodiment, since the sonic microphone can detect a wider range of sound frequencies, the sound signal collected by the sonic microphone is used for voice recognition during the secondary wake-up check, which can further improve the accuracy of wake-up word recognition, thereby ensuring the sensitivity of voice wake-up.
[0056] Among them, the second speech recognition model is trained and acquired based on a sound sample set collected by an acoustic microphone. Among them, there are at least some sound samples carrying wake-up words in the sound sample set. It should be understood that the second speech recognition model of this embodiment can be any network structure that can realize speech recognition, such as convolutional neural networks, recurrent neural networks and their variants, long short-term memory networks and gated recurrent units, Transformer models, end-to-end models, or hybrid models, etc., or other lightweight network models that can realize speech recognition, which will not be explained one by one in this embodiment.
[0057] In another optional implementation, the secondary voice wake-up of this embodiment is implemented based on at least one sonic microphone Mic and a bone conduction microphone 15, that is, the above-mentioned target microphone includes at least one sonic microphone Mic and a bone conduction microphone 15. Further, a secondary voice wake-up module is deployed in the controller 16. Among them, the controller 16 is used to control the start of at least one sonic microphone to continuously collect sound signals in response to the detection of a voice activity signal, and obtain the sound signal continuously captured by the bone conduction microphone 15, and then the secondary voice wake-up module performs fusion processing on the multiple sound signals collected by the sonic microphone Mic and the bone conduction microphone 15 according to the third voice recognition model based on pre-training to perform wake-up word detection. Therefore, this embodiment can combine the sonic microphone and the bone conduction microphone to detect the wake-up word in the secondary voice wake-up, further improve the accuracy of the wake-up word detection, avoid the situation of false wake-up, and further reduce the power consumption of the device. The third voice recognition model is obtained by training a set of sound sample pairs collected by the sonic microphone and the bone conduction microphone. Among them, the sound sample pair includes the synchronously recorded sound signals of the same sound by the sonic microphone and the bone conduction microphone.
[0058] Furthermore, the third speech recognition model of this embodiment can be an end-to-end model, and the multiple sound signals collected by the acoustic microphone Mic and the bone conduction microphone 15 are input into the third speech recognition model for fusion processing to obtain the wake-up word detection result. In other optional implementations, the third speech recognition model can also be a non-end-to-end model, and the controller 16 can first pre-process the sound signals collected by the acoustic microphone and the sound signals collected by the bone conduction microphone, and input the pre-processed sound features into the third speech recognition model for processing to obtain the wake-up word detection result. It should be understood that this embodiment does not limit the network structure of the third speech recognition model, and it can also be the model structure of the above example or other types of model structures, which will not be illustrated one by one here.
[0059] Further, the controller 16 can perform signal processing such as power amplification, filtering, and / or noise reduction on the sound signals synchronously collected by the acoustic microphone Mic and the bone conduction microphone 15, and perform feature extraction on the multiple sound signals after signal processing, such as extracting the MFCC (Mel Cepstral Coefficient) or speech spectrum of the sound signal after signal processing corresponding to the acoustic microphone to obtain the sound feature F1, and extracting the time domain features (such as amplitude envelope), frequency domain features (such as spectrum characteristics), and / or other high-level features (such as wavelet transform coefficients) of the sound signal after signal processing corresponding to the bone conduction microphone 15 to obtain the bone conduction feature F2. Further, the controller 16 performs signal alignment and normalization on the sound feature F1 and the bone conduction feature F2. Among them, signal alignment is used to solve the delay between signals, and time alignment can be performed by cross-correlation or other methods. Normalization is used to adjust the sound feature F1 and the bone conduction feature F2 to the same scale for subsequent processing. For example, through normalization or standardization processing, the sound feature F1 and the bone conduction feature F2 have a similar distribution range. Furthermore, the controller 16 may input the sound feature F1 and the bone conduction feature F2 after signal alignment and normalization into the third speech recognition model for fusion processing, or may splice the sound feature F1 and the bone conduction feature F2 after signal alignment and normalization and input them into the third speech recognition model for fusion processing to obtain the wake-up word detection result. It should be understood that the above-mentioned sound signal preprocessing process is only exemplary, and this embodiment does not limit the specific preprocessing method in the specific implementation, which can realize the fusion recognition of different types of sound signals.
[0060] In another optional implementation, the secondary voice wake-up of this embodiment is implemented based on at least one sonic microphone Mic and a bone conduction microphone 15, that is, the above-mentioned target microphone includes at least one sonic microphone Mic and a bone conduction microphone 15. Further, a secondary voice wake-up module is deployed in the controller 16. The controller 16 is used to control the start of at least one sonic microphone in response to the detection of a voice activity signal, and perform voice recognition on the sound signal collected by the bone conduction microphone 25 based on a pre-trained first voice recognition model, and in response to the first voice recognition model detecting the wake-up word, perform voice recognition on the sound signal collected by the sonic microphone Mic based on a pre-trained second voice recognition model to verify the detection of the wake-up word. As a result, this embodiment performs secondary voice wake-up through the sound signal collected by the bone conduction microphone, and verifies the detection of the wake-up word through the sound signal collected by the sonic microphone after the wake-up word is detected, which further improves the accuracy of the detection of the wake-up word, thereby avoiding the situation of false wake-up and reducing the power consumption of the device. It should be understood that the first voice recognition model and the second voice recognition model of this embodiment are similar to the above-mentioned implementation, and will not be described in detail here.
[0061] Further, the speech recognition model of this embodiment has a corresponding confidence in the output wake-up word detection result. In response to the first speech recognition model detecting the wake-up word and the detection confidence being less than a predetermined value, the controller 16 of this embodiment performs speech recognition on the sound signal collected by the acoustic microphone based on the pre-trained second speech recognition model to verify the detection of the wake-up word. That is, when the confidence of the result of the first speech recognition model detecting the wake-up word is high, it indicates that the detection credibility of the first speech recognition model is high. Therefore, the controller 16 can wake up the corresponding device when the first speech recognition model detects the wake-up word and the confidence is greater than a predetermined value. When the confidence of the result of the first speech recognition model detecting the wake-up word is low, it indicates that the detection credibility of the first speech recognition model is low. Therefore, the controller 16 uses the second speech recognition model to perform speech recognition on the sound signal collected by the acoustic microphone Mic to verify the detection of the wake-up word. After the output of the second speech recognition model indicates that the wake-up word is detected, the corresponding device is woken up. Therefore, this embodiment determines whether to perform wake-up word verification based on the acoustic microphone by judging the confidence, which further reduces the amount of data calculation while avoiding false wake-up, thereby further reducing the power consumption of the device.
[0062] Furthermore, the voice wake-up system of this embodiment also includes a main control 17. Among them, the controller 16 is used to wake up the main control 17 in response to detecting the wake-up word. After the controlled wake-up, the main control 17 performs signal fusion processing on the received sound signal, obtains the processed audio signal, and recognizes the audio signal to achieve corresponding calls, voice interaction, audio processing and other functions. It should be understood that the functions that the main control 17 can achieve are configured based on the functions required by the wearable device, and this embodiment will not be described in detail.
[0063] In an optional implementation, the main control 17 is deployed with a fourth speech recognition model, which is used to verify the wake-up word detection result of the controller according to the fourth speech recognition model. In this embodiment, if the computing resources of the controller 16 are insufficient to deploy a complete voice wake-up module, this embodiment can deploy the wake-up word verification function in the main control 17 to verify the detection result of any implementation method of the wake-up word recognition in the above-mentioned controller 16, so as to further avoid the situation of false detection. Accordingly, the fourth speech recognition module can adopt speech recognition based on an acoustic wave microphone, or it can adopt speech recognition combining a bone conduction microphone and an acoustic wave microphone, which will not be described in detail here. Furthermore, this embodiment can also determine whether to start the wake-up word detection verification in the main control 17 by the confidence of the wake-up word detection result of the controller 16, so as to reduce the startup frequency of the main control 17 and further reduce the power consumption of the device.
[0064] In an optional implementation, the controller 16 of this embodiment may be a controller having a signal processing module, such as a DSP (Digital Signal Processing) or an MCU (Microcontroller Unit).
[0065] In an optional implementation, the voice wake-up system of this embodiment may also include a memory for storing instructions that can be executed by the controller 16 and / or the main control 17, and the instructions are executed by the controller 16 and / or the main control 17 to realize wake-up word recognition. Furthermore, the memory can be used as a non-volatile computer-readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The controller 16 and / or the main control 17 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions and modules stored in the memory, that is, realizing wake-up word recognition. Furthermore, one or more voice wake-up modules can be stored in the memory, and when executed by the controller 16 and / or the main control 17, the wake-up word recognition in any of the above method embodiments is performed.
[0066] The voice wake-up system of the embodiment of the present invention includes at least one acoustic microphone, at least one bone conduction microphone and a controller, wherein the controller is used to detect a voice activity signal according to the bone conduction microphone, perform voice recognition on the collected sound signal in response to the detection of the voice activity signal, and wake up the corresponding device in response to the detection of the wake-up word. Therefore, the embodiment of the present invention can use the bone conduction microphone voice activity detection to achieve the first-level voice wake-up, and perform voice recognition on the collected sound signal after the voice activity is detected to perform the second-level voice wake-up of the wake-up word, which enables the embodiment of the present invention to effectively avoid false wake-ups while ensuring the wake-up sensitivity, and reduce the power consumption of the device.
[0067] Figure 2 is a schematic diagram of a wearable device according to an embodiment of the present invention. Figure 3 is a schematic diagram of another wearable device according to an embodiment of the present invention. The wearable device according to the embodiment of the present invention includes a device body and a voice wake-up system in any of the above implementations. This embodiment is described in detail by taking the wearable device as smart glasses as an example. It should be understood that this embodiment is not limited to this, and it can be any wearable device that needs to pick up sound. Figure 2 As shown, the smart glasses of this embodiment include a glasses body and a voice wake-up system. The voice wake-up system includes acoustic microphones Mic21-Mic24, a bone conduction microphone 25, and a controller and / or a main control (not shown in the figure).
[0068] Furthermore, if Figure 2As shown, the acoustic wave microphones Mic21 and Mic22 are arranged in the temple 10 of the smart glasses, and the acoustic wave microphones Mic23 and Mic24 are arranged in the temple 20 of the smart glasses. Optionally, the arrangement positions of the acoustic wave microphones Mic21 and Mic22 are symmetrical with the arrangement positions of the acoustic wave microphones Mic23 and Mic24. However, it should be understood that the acoustic wave microphones arranged in the temple 10 and the temple 20 may also be arranged asymmetrically. In other optional implementations, some of the acoustic wave microphones may also be arranged on the frame, such as Figure 3 As shown, for example, the acoustic wave microphone Mic22 is arranged on the temple 10, the acoustic wave microphone Mic21 is arranged on the frame on the same side of the temple 10, the acoustic wave microphone Mic24 is arranged on the temple 20, and the acoustic wave microphone Mic23 is arranged on the frame on the same side of the temple 20. It should be understood that this embodiment does not limit the specific positions of the arrangements.
[0069] Further, the bone conduction microphone 25 is arranged in a predetermined area of the smart glasses, that is, in the area 30 where the nose pads are located. It should be understood that the present embodiment can determine the size of the area where the nose pads are located based on experience or testing to ensure the effect of capturing the sound signal when the wearer is speaking. Further, the bone conduction microphone 25 is arranged in the nose pads, and it can also be arranged in the frame in the area 30 where the nose pads are located, such as on the connecting frame of the frames on both sides. In other optional implementations, the bone conduction microphone 25 can also be arranged in the contact area between the temples and the wearer. It should be understood that the present embodiment does not limit the specific setting position of the bone conduction microphone 25, and it can capture the bone vibration when the wearer is speaking.
[0070] Furthermore, the controller in the voice wake-up system of this embodiment is used to detect the voice activity signal according to the bone conduction microphone 25, perform voice recognition on the collected sound signal in response to the detection of the voice activity signal, and wake up the device in response to the detection of the wake-up word. After the controlled wake-up, the main control in the voice wake-up system performs signal fusion processing on the received sound signal, obtains the processed audio signal, and recognizes the audio signal to realize the corresponding call, voice interaction, audio processing and other functions.
[0071] Furthermore, the specific implementation methods of the controller of this embodiment (or in combination with the main control) performing the one-time wake-up of the bone conduction microphone and the secondary wake-up word wake-up based on the bone conduction microphone and / or the acoustic microphone can refer to the above embodiments and will not be described in detail one by one here.
[0072] The voice wake-up system in the wearable device of the embodiment of the present invention includes at least one acoustic microphone, at least one bone conduction microphone and a controller, wherein the controller is used to detect a voice activity signal according to the bone conduction microphone, perform voice recognition on the collected sound signal in response to the detection of the voice activity signal, and wake up the corresponding device in response to the detection of the wake-up word. Therefore, the embodiment of the present invention can use the bone conduction microphone voice activity detection to achieve the first-level voice wake-up, and perform voice recognition on the collected sound signal after the voice activity is detected to perform the second-level voice wake-up of the wake-up word, which enables the embodiment of the present invention to effectively avoid false wake-ups while ensuring the wake-up sensitivity, and reduce the power consumption of the device.
[0073] Figure 4 1 is a flow chart of a wake-up method of a voice wake-up system according to an embodiment of the present invention. The voice wake-up system includes at least one acoustic microphone and at least one bone conduction microphone. Figure 4 As shown, the wake-up method of the voice wake-up system of the embodiment of the present invention includes the following steps:
[0074] Step S110, start the bone conduction microphone to detect voice activity signals. Since the bone conduction microphone receives the voiceprint vibration signal when the wearer speaks, but cannot receive the sounds of the surrounding environment and other speakers, the present embodiment can more accurately detect whether the wearer of the device is speaking, that is, whether there is voice activity, through the bone conduction microphone, avoiding interference from the surrounding environment or other speakers, effectively avoiding false wake-ups while ensuring wake-up sensitivity, and reducing device power consumption.
[0075] Step S120, in response to the bone conduction microphone detecting the voice activity signal, collecting the sound signal of the target microphone, wherein the target microphone includes at least one of the acoustic wave microphone and / or the bone conduction microphone.
[0076] Step S130: Perform speech recognition on the collected sound signal to detect the wake-up word.
[0077] Step S140, in response to detecting the wake-up word, waking up the corresponding device.
[0078] The embodiment of the present invention detects a voice activity signal based on a bone conduction microphone, performs voice recognition on the collected sound signal in response to detecting the voice activity signal, and wakes up the corresponding device in response to detecting the wake-up word. Therefore, the embodiment of the present invention can use the bone conduction microphone voice activity detection to achieve primary voice wake-up, and perform voice recognition on the collected sound signal after detecting the voice activity to perform secondary voice wake-up of the wake-up word, which enables the embodiment of the present invention to effectively avoid false wake-ups while ensuring the wake-up sensitivity, and reduce the power consumption of the device.
[0079] Figure 5 FIG. 1 is a flowchart of another wake-up method of a voice wake-up system according to an embodiment of the present invention. In an optional implementation, the secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on a bone conduction microphone. Figure 5 As shown, the wake-up method of the voice wake-up system of the embodiment of the present invention includes the following steps:
[0080] Step S210, starting the bone conduction microphone to detect voice activity signals.
[0081] Step S220: In response to the bone conduction microphone detecting a voice activity signal, collecting a sound signal of the bone conduction microphone.
[0082] Step S230, based on the first speech recognition model, speech recognition is performed on the collected sound signal to detect the wake-up word. The first speech recognition model is acquired based on the training of the sound sample set collected by the bone conduction microphone. At least some sound samples carrying the wake-up word exist in the sound sample set. It should be understood that the first speech recognition model of this embodiment can be any network structure that can realize speech recognition (Automatic Speech Recognition, ASR), such as convolutional neural network (CNN), recurrent neural network (RNN) and its variant long short-term memory network (LSTM) and gated recurrent unit (GRU), Transformer model, end-to-end model, or hybrid model, etc. It can also be other lightweight network models that can realize speech recognition, and this embodiment will not be illustrated one by one here.
[0083] Step S240, in response to detecting the wake-up word, waking up the corresponding device.
[0084] Figure 6 Schematic diagram of the wake-up process of a voice wake-up system according to an embodiment of the present invention. The secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on a bone conduction microphone. Figure 6As shown, after the bone conduction microphone 61 is started, it captures the bone vibration signal, obtains the bone conduction signal s1 and transmits it to the controller 62. The first-level wake-up module 621 deployed in the controller 62 detects the bone conduction signal s1 to determine whether it represents voice activity. It should be understood that the first-level wake-up module 621 can pre-store a sound model representing voice activity, and after receiving the bone conduction signal s1, it determines whether it conforms to the pre-stored sound model to determine whether voice activity is currently occurring. After the first-level wake-up module 621 detects the voice activity signal, it controls the first voice recognition model M1 in the second-level wake-up module 622 to perform voice recognition processing on the continuously received bone conduction signal s2 to determine whether the bone conduction signal s2 contains a wake-up word. If the first voice recognition model M1 recognizes the wake-up word, the second-level wake-up module 622 generates a wake-up signal and sends it to the main control 63 to wake up the main control 63 to perform corresponding voice processing operations to achieve voice interaction, calls, recording and other functions.
[0085] The embodiment of the present invention detects a voice activity signal based on a bone conduction microphone, performs voice recognition on the sound signal collected by the bone conduction microphone based on a first voice recognition model in response to detecting the voice activity signal, and wakes up the corresponding device in response to detecting the wake-up word. Therefore, the voice wake-up system of the embodiment of the present invention only activates the bone conduction microphone to implement the primary voice wake-up and the secondary voice wake-up of the wake-up word detection, so that other microphones do not need to remain in working state, which enables the embodiment of the present invention to further reduce the power consumption of the device while ensuring the wake-up sensitivity.
[0086] Figure 7 1 is a flowchart of another wake-up method of a voice wake-up system according to an embodiment of the present invention. The secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on at least one acoustic microphone. Figure 7 As shown, the wake-up method of the voice wake-up system of the embodiment of the present invention includes the following steps:
[0087] Step S310, starting the bone conduction microphone to detect voice activity signals.
[0088] Step S320, in response to the bone conduction microphone detecting the voice activity signal, controlling to start at least one acoustic wave microphone to collect the corresponding sound signal.
[0089] Step S330, based on the second speech recognition model, speech recognition is performed on the sound signal collected by the acoustic microphone to detect the wake-up word. The second speech recognition model is trained and acquired based on the sound sample set collected by the acoustic microphone. At least some sound samples carrying the wake-up word exist in the sound sample set. It should be understood that the second speech recognition model of this embodiment can be any network structure that can realize speech recognition, such as convolutional neural networks, recurrent neural networks and their variants, long short-term memory networks and gated recurrent units, Transformer models, end-to-end models, or hybrid models, etc. It can also be other lightweight network models that can realize speech recognition, and this embodiment will not illustrate them one by one.
[0090] Step S340, in response to detecting the wake-up word, waking up the corresponding device.
[0091] Figure 8 Schematic diagram of another wake-up process of a voice wake-up system according to an embodiment of the present invention. The secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on at least one acoustic microphone. Figure 8 As shown, after the bone conduction microphone 81 is started, it captures the bone vibration signal, obtains the bone conduction signal s3 and transmits it to the controller 83. The first-level wake-up module 831 deployed in the controller 83 detects the bone conduction signal s3 to determine whether it represents voice activity. It should be understood that the first-level wake-up module 831 can pre-store a sound model representing voice activity, and after receiving the bone conduction signal s3, determine whether it conforms to the pre-stored sound model to determine whether voice activity is currently occurring. After the first-level wake-up module 831 detects the voice activity signal, the controller 83 controls the start of at least one sound wave microphone 82 to record the sound wave signal s4. The second speech recognition model M2 in the second-level wake-up module 832 in the controller 83 performs speech recognition processing on the continuously received sound wave signal s4 to determine whether the sound wave signal s4 contains a wake-up word. If the second speech recognition model M2 recognizes the wake-up word, the second-level wake-up module 832 generates a wake-up signal and sends it to the main control 84 to wake up the main control 84 to perform corresponding voice processing operations to achieve voice interaction, calls, recording and other functions.
[0092] The embodiment of the present invention detects a voice activity signal based on a bone conduction microphone, controls the start-up of at least one sonic wave microphone to collect a corresponding sound signal in response to the detection of the voice activity signal, performs speech recognition on the sound signal collected by the sonic wave microphone based on a second speech recognition model to detect a wake-up word, and wakes up the corresponding device in response to the detection of the wake-up word. Thus, the voice wake-up system of the embodiment of the present invention implements voice activity detection (i.e., primary voice wake-up) through a bone conduction microphone, and implements secondary voice wake-up for wake-up word detection based on the sonic wave microphone. In this embodiment, since the sonic wave microphone can detect a wider range of sound frequencies, the sound signal collected by the sonic wave microphone is used for speech recognition during the secondary wake-up check, which can further improve the accuracy of wake-up word recognition, thereby reducing the power consumption of the device while ensuring the sensitivity of voice wake-up.
[0093] Fig. 9 2 is a flowchart of another wake-up method of a voice wake-up system according to an embodiment of the present invention. The secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on at least one acoustic microphone and a bone conduction microphone. Fig. 9 As shown, the wake-up method of the voice wake-up system of the embodiment of the present invention includes the following steps:
[0094] Step S410, starting the bone conduction microphone to detect a voice activity signal.
[0095] Step S420, in response to the bone conduction microphone detecting a voice activity signal, controlling to start at least one acoustic wave microphone to collect a corresponding sound signal.
[0096] Step S430, obtaining the sound signal collected by the bone conduction microphone.
[0097] Step S440, based on the third speech recognition model, speech recognition is performed on multiple sound signals collected by the acoustic microphone and the bone conduction microphone to detect the wake-up word. The third speech recognition model is obtained by training a set of sound sample pairs collected by the acoustic microphone and the bone conduction microphone. The sound sample pair includes a synchronously recorded sound signal of the same sound by the acoustic microphone and the bone conduction microphone.
[0098] Step S450, in response to detecting the wake-up word, waking up the corresponding device.
[0099] Furthermore, the third speech recognition model of this embodiment can be an end-to-end model, and multiple sound signals collected by the acoustic wave microphone and the bone conduction microphone are input into the third speech recognition model for fusion processing to obtain the wake-up word detection result. In other optional implementations, the third speech recognition model can also be a non-end-to-end model. The controller can first pre-process the sound signals collected by the acoustic wave microphone and the sound signals collected by the bone conduction microphone, and input the pre-processed sound features into the third speech recognition model for processing to obtain the wake-up word detection result. It should be understood that this embodiment does not limit the network structure of the third speech recognition model. It can also be the model structure of the above example or other types of model structures, which will not be explained one by one here.
[0100] Furthermore, the present embodiment can perform signal processing such as power amplification, filtering, and / or noise reduction on the sound signals synchronously collected by the acoustic microphone and the bone conduction microphone, and perform feature extraction on the multiple sound signals after signal processing to obtain the sound feature F1 and the bone conduction feature F2. Furthermore, the present embodiment performs signal alignment and normalization on the sound feature F1 and the bone conduction feature F2, and inputs the aligned and normalized sound feature F1 and the bone conduction feature F2 into the third speech recognition model for fusion processing, or splices the aligned and normalized sound feature F1 and the bone conduction feature F2 and inputs them into the third speech recognition model for fusion processing to obtain the wake-up word detection result. It should be understood that the above-mentioned sound signal preprocessing process is only exemplary, and the present embodiment does not limit the specific preprocessing method in the specific implementation, which can realize the fusion recognition of different types of sound signals.
[0101] Fig.10 Schematic diagram of another wake-up process of a voice wake-up system according to an embodiment of the present invention. The secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on at least one acoustic microphone and a bone conduction microphone. Fig.10As shown, after the bone conduction microphone 101 is activated, it captures the bone vibration signal, obtains the bone conduction signal s5 and transmits it to the controller 103. The first-level wake-up module 1031 deployed in the controller 103 detects the bone conduction signal s5 to determine whether it represents voice activity. It should be understood that the first-level wake-up module 1031 can pre-store a sound model representing voice activity, and after receiving the bone conduction signal s5, it determines whether it conforms to the pre-stored sound model to determine whether voice activity is currently occurring. After the first-level wake-up module 1031 detects the voice activity signal, the controller 103 controls the activation of at least one sound wave microphone 102 to record the sound wave signal s7. The third speech recognition model M3 in the secondary wake-up module 1032 in the controller 103 performs speech fusion processing on the continuously received bone conduction signal s6 and sound wave signal s7, so as to determine whether the sound emitted by the person wearing the device contains a wake-up word in combination with the bone conduction signal s6 and the sound wave signal s7. If the third speech recognition model M3 recognizes the wake-up word, the secondary wake-up module 1032 generates a wake-up signal and sends it to the main control 104 to wake up the main control 104 to perform corresponding speech processing operations to realize voice interaction, calls, recording and other functions. This embodiment is described by taking the third speech recognition model M3 as an end-to-end model as an example. It should be understood that the above-mentioned other models combined with the signal processing module can also be applied to this embodiment, which will not be described in detail here.
[0102] The embodiment of the present invention detects a voice activity signal based on a bone conduction microphone, and in response to detecting a voice activity signal, controls the start-up of at least one acoustic wave microphone to collect a corresponding sound signal, obtains the sound signal collected by the bone conduction microphone, performs speech recognition on multiple sound signals collected by the acoustic wave microphone and the bone conduction microphone based on a third speech recognition model to perform wake-up word detection, and wakes up the corresponding device in response to detecting the wake-up word. As a result, the voice wake-up system of the embodiment of the present invention implements voice activity detection (i.e., primary voice wake-up) through a bone conduction microphone, and combines the acoustic wave microphone and the bone conduction microphone in the secondary voice wake-up to perform wake-up word detection, further improving the accuracy of wake-up word detection, avoiding false wake-ups, and further reducing device power consumption.
[0103] Fig.11 FIG. 1 is a flowchart of another method for waking up a voice wake-up system according to an embodiment of the present invention. The second-level voice wake-up of this embodiment is implemented based on at least one acoustic microphone and a bone conduction microphone. Fig.11 As shown, the wake-up method of the voice wake-up system of the embodiment of the present invention includes the following steps:
[0104] Step S510, starting the bone conduction microphone to detect voice activity signals.
[0105] Step S520, in response to the bone conduction microphone detecting a voice activity signal, controlling to start at least one acoustic wave microphone to collect a corresponding sound signal.
[0106] Step S530: Based on the first speech recognition model, the sound signal collected by the bone conduction microphone is speech recognized to obtain the wake-up word detection result. It should be understood that the first speech recognition model of this embodiment is similar to the above implementation mode and will not be described in detail here.
[0107] Step S540, determining whether the wake-up word detection result obtained in the above step S530 detects the wake-up word. If the wake-up word is detected, executing step S550, if the wake-up word is not detected, re-executing step S510.
[0108] Step S550, in response to detecting the wake-up word, determine whether the detection confidence of the wake-up word recognized by the first speech recognition model is greater than a predetermined value, if the confidence is greater than the predetermined value, execute step S580, if the confidence is not greater than the predetermined value, execute step S560.
[0109] Step S560: Based on the second speech recognition model, the sound signal collected by the sound wave microphone is speech recognized to obtain the wake-up word detection result. It should be understood that the second speech recognition model of this embodiment is similar to the above implementation mode and will not be described in detail here.
[0110] Step S570, determining whether the wake-up word detection result obtained in the above step S560 detects the wake-up word. If the wake-up word is detected, executing step S580, if the wake-up word is not detected, re-executing step S510.
[0111] Step S580, in response to the first speech recognition model detecting the wake-up word and the detection confidence is greater than a predetermined value, or the second speech recognition model detecting the wake-up word, waking up the corresponding device.
[0112] Fig.12 Schematic diagram of another wake-up process of a voice wake-up system according to an embodiment of the present invention. The secondary voice wake-up (i.e., the wake-up word detection process) of this embodiment is implemented based on at least one acoustic microphone and a bone conduction microphone. Fig.12As shown, after the bone conduction microphone 121 is activated, it captures the bone vibration signal, obtains the bone conduction signal s8 and transmits it to the controller 123. The first-level wake-up module 1231 deployed in the controller 123 detects the bone conduction signal s8 to determine whether it represents voice activity. It should be understood that the first-level wake-up module 1231 can pre-store a sound model representing voice activity, and after receiving the bone conduction signal s8, it determines whether it conforms to the pre-stored sound model to determine whether voice activity is currently occurring. After the first-level wake-up module 1231 detects the voice activity signal, the controller 123 controls the activation of at least one sound wave microphone 122 to record the sound wave signal s10. After the first-level wake-up module 1231 detects a voice activity signal, the first voice recognition model M1 in the second-level wake-up module 1232 is controlled to perform voice recognition processing on the continuously received bone conduction signal s9 to determine whether the bone conduction signal s9 contains a wake-up word. If the first voice recognition model M1 recognizes the wake-up word and the corresponding confidence level is greater than a predetermined value, the second-level wake-up module 1232 generates a wake-up signal and sends it to the main control 124 to wake up the main control 124 to perform corresponding voice processing operations to achieve voice interaction, calls, recording and other functions.
[0113] If the first speech recognition model M1 recognizes the wake-up word and the corresponding confidence level is not greater than a predetermined value, the second speech recognition model M2 in the secondary wake-up module 1232 in the controller 123 performs speech recognition processing on the continuously received sound wave signal s10 to determine whether the sound wave signal s10 contains the wake-up word. If the second speech recognition model M2 recognizes the wake-up word, the secondary wake-up module 1232 generates a wake-up signal and sends it to the main control 124 to wake up the main control 84 to perform corresponding speech processing operations to achieve voice interaction, calls, recording and other functions.
[0114] This embodiment performs secondary voice wake-up through the sound signal collected by the bone conduction microphone, and verifies the wake-up word detection through the sound signal collected by the acoustic microphone after the wake-up word is detected and the confidence level is not greater than a predetermined value, further improving the accuracy of the wake-up word detection, thereby avoiding false wake-ups and reducing device power consumption. At the same time, this embodiment determines whether to perform wake-up word verification based on the acoustic microphone by judging the confidence level, further reducing the amount of data calculation while avoiding false wake-ups, thereby further reducing device power consumption.
[0115] In other optional implementations, this embodiment may not perform a confidence level judgment after the bone conduction microphone detects the wake-up word, and directly verify the wake-up word detection based on the acoustic wave microphone, which can further avoid false wake-ups.
[0116] Fig.13This is a flowchart of another wake-up method of a voice wake-up system according to an embodiment of the present invention. In this embodiment, if the computing resources of the controller in the voice wake-up system are insufficient to deploy a complete voice wake-up module, this embodiment can deploy the wake-up word verification function in the main control of the voice wake-up system to verify the detection results of any wake-up word recognition implementation method in the above controller, thereby further avoiding false detection. Fig.13 As shown, the wake-up method of the voice wake-up system of this embodiment includes the following steps:
[0117] Step S610, starting the bone conduction microphone to detect voice activity signals.
[0118] Step S620: In response to the bone conduction microphone detecting the voice activity signal, the sound signal of the target microphone is collected, wherein the target microphone includes at least one of the acoustic wave microphone and / or the bone conduction microphone.
[0119] Step S630: Perform speech recognition on the collected sound signal to detect the wake-up word. It should be understood that the wake-up word detection process executed in the controller of this embodiment can adopt any of the above implementations, which will not be described in detail here.
[0120] Step S640, in response to detecting the wake-up word, transmit the collected at least one sound signal to the main control, so as to perform speech recognition on the sound signal based on the fourth speech recognition model deployed in the main control, so as to verify the detection of the wake-up word. Accordingly, the fourth speech recognition module can adopt speech recognition based on an acoustic wave microphone, or speech recognition combining a bone conduction microphone and an acoustic wave microphone, which will not be described in detail here.
[0121] Step S640: In response to the fourth speech recognition model detecting the wake-up word, waking up the corresponding device.
[0122] In a further optional implementation, this embodiment can also determine whether to start the wake-up word detection check in the main control by the confidence level of the wake-up word detection result of the controller, so as to reduce the startup frequency of the main control and further reduce the power consumption of the device.
[0123] Fig.14 Schematic diagram of the wake-up process of another voice wake-up system of an embodiment of the present invention. In this embodiment, if the computing resources of the controller in the voice wake-up system are insufficient to deploy a complete voice wake-up module, this embodiment can deploy the wake-up word verification function in the main control of the voice wake-up system to verify the detection result of any wake-up word recognition implementation method in the above controller, thereby further avoiding false detection. Fig.14As shown, the controller 143 of this embodiment is deployed with a primary wake-up module 1431 and a secondary wake-up module 1432. After the primary wake-up module 1431 detects voice activity based on the bone conduction signal s11, the secondary wake-up module 1432 performs voice recognition on the bone conduction signal s11 and / or the sound wave signal s12 to determine whether the wake-up word is detected. It should be understood that the wake-up word detection process performed by the secondary wake-up module 1432 in the controller 143 of this embodiment can adopt any of the above implementation methods, which will not be described in detail here. If the secondary wake-up module 1432 detects the wake-up word, it generates a start signal to start the fourth voice recognition model M4 in the main control 144. The fourth voice recognition model M4 performs voice recognition on the bone conduction signal s11 and / or the sound wave signal s12 to determine whether the wake-up word is detected. If the wake-up word is detected, a wake-up signal is generated to wake up other modules in the main control 144, and then perform corresponding voice processing operations to achieve voice interaction, calls, recording and other functions. It should be understood that the fourth speech recognition model M4 can be deployed according to specific application requirements. According to the speech recognition method, it can be the same or a type of model as the above-mentioned first speech recognition model M1, second speech recognition model M2, or third speech recognition model M3, and will not be described in detail here.
[0124] The embodiment of the present invention detects a voice activity signal based on a bone conduction microphone, performs voice recognition on the collected sound signal in response to the detection of the voice activity signal, transmits at least one collected sound signal to the main control in response to the detection of the wake-up word, performs voice recognition on the sound signal based on the fourth voice recognition model deployed in the main control, verifies the detection of the wake-up word, and wakes up the corresponding device in response to the fourth voice recognition model detecting the wake-up word. Therefore, when the computing resources of the controller are insufficient to deploy a complete voice wake-up module, this embodiment can deploy the wake-up word verification function in the main control of the voice wake-up system to perform the wake-up word detection verification, so as to further avoid false wake-ups and reduce device power consumption.
[0125] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used for a computer to execute part or all of the above method embodiments.
[0126] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0127] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A voice wake-up system, characterized in that: The voice wake-up system comprises: at least one acoustic microphone; at least one bone conduction microphone; and The controller is used to detect a voice activity signal according to the bone conduction microphone, perform voice recognition on the collected sound signal in response to detecting the voice activity signal, and wake up the corresponding device in response to detecting a wake-up word.
2. The voice wake-up system according to claim 1, characterized in that: The at least one acoustic wave microphone is distributed on both sides of the wearable device, and the bone conduction microphone is arranged in a predetermined area of the wearable device, wherein at least a part of the predetermined area contacts the wearer after the wearable device is worn.
3. The voice wake-up system according to claim 1, characterized in that: The controller is used for performing speech recognition on the sound signal collected by the bone conduction microphone based on a pre-trained first speech recognition model in response to detecting a voice activity signal, so as to perform wake-up word detection.
4. The voice wake-up system according to claim 1, characterized in that: The controller is used to control the activation of at least one acoustic wave microphone in response to detecting a voice activity signal, and perform voice recognition on a sound signal collected by the acoustic wave microphone based on a pre-trained second voice recognition model to perform wake-up word detection.
5. The voice wake-up system according to claim 1, characterized in that: The controller is used to control the activation of at least one acoustic wave microphone in response to detecting a voice activity signal, and to perform fusion processing on multiple sound signals collected by the bone conduction microphone and each of the acoustic wave microphones based on a pre-trained third speech recognition model to perform wake-up word detection.
6. The voice wake-up system according to claim 1, characterized in that: The controller is used to control the activation of at least one acoustic wave microphone in response to detecting a voice activity signal, perform voice recognition on the sound signal collected by the bone conduction microphone based on a pre-trained first voice recognition model, and in response to the first voice recognition model detecting a wake-up word, perform voice recognition on the sound signal collected by each of the acoustic wave microphones based on a pre-trained second voice recognition model to verify the wake-up word detection.
7. The voice wake-up system according to claim 6, characterized in that: The controller is used to perform speech recognition on the sound signal collected by the acoustic microphone based on a pre-trained second speech recognition model in response to the first speech recognition model detecting the wake-up word and the detection confidence is less than a predetermined value to verify the wake-up word detection.
8. The voice wake-up system according to claim 1, characterized in that: The voice wake-up system also includes: The master control is used to perform audio processing on the received sound signal after the controlled wake-up; The controller is used to wake up the main control in response to detecting a wake-up word.
9. The voice wake-up system according to claim 8, characterized in that: The main control is deployed with a fourth speech recognition model, and the main control is used to verify the wake-up word detection result of the controller according to the fourth speech recognition model.
10. A wearable device, characterized in that: The wearable device comprises: The device body; and A voice wake-up system as claimed in any one of claims 1 to 9.
11. The wearable device according to claim 10, characterized in that: The wearable device is smart glasses, the sound wave microphones in the voice wake-up system are distributed on the temples and / or the frames, and the bone conduction microphones in the voice wake-up system are arranged in the area where the nose pads are located.
12. The wearable device according to claim 11, characterized in that: The acoustic microphones are symmetrically distributed on the temples and / or frames on both sides.
13. A wake-up method for a voice wake-up system, characterized in that: The voice wake-up system includes at least one acoustic wave microphone and at least one bone conduction microphone, and the method includes: Activating the bone conduction microphone to detect a voice activity signal; In response to the bone conduction microphone detecting a voice activity signal, collecting a sound signal from a target microphone, wherein the target microphone includes at least one of the acoustic wave microphone and / or the bone conduction microphone; Performing speech recognition on the sound signal to detect a wake-up word; In response to detecting the wake-up word, waking up the corresponding device.