Voice activity detection method, device and equipment

By combining air conduction and bone conduction sound data for speech activity detection, the problem of low robustness of voice interaction in complex acoustic environments of smart wearable devices is solved. This enables accurate differentiation of speaker identity and speech activity detection, thereby improving the voice interaction experience of the device.

CN121583293APending Publication Date: 2026-02-27SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511770587.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing smart wearable devices have low robustness in voice interaction under complex acoustic environments, especially in high-noise and multi-person dialogue environments, where it is difficult to distinguish the identity of the target speaker, leading to false wake-up and voice recognition failure.

Method used

A speech activity detection method combining air conduction and bone conduction sound data is adopted. Deep features are obtained through air conduction speech activity detection module and bone conduction speech activity detection module. The detection results of the two are combined to make joint decisions, thereby realizing the differentiation of speaker identity and speech activity detection.

Benefits of technology

It improves the robustness of voice activity detection, reduces the false negative and false positive rates of device wake-up, and enhances the accuracy of voice interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583293A_ABST
    Figure CN121583293A_ABST
Patent Text Reader

Abstract

The invention discloses a voice activity detection method and device, a face-to-face translation method and device, an equipment awakening method and device, a speaker recognition method and device, a voice activity detection model construction method and device and electronic equipment. The voice activity detection method comprises the following steps: acquiring air conduction sound data and bone conduction sound data; acquiring an air conduction voice activity detection result according to the air conduction voice data; acquiring a bone conduction voice activity detection result according to the bone conduction voice data; obtaining a joint voice activity detection result according to the air conduction voice activity detection result and the bone conduction voice activity detection result; obtaining a joint voice activity detection result according to the air conduction voice activity detection result and the bone conduction voice activity detection result; the joint voice activity detection result comprises no person speaking, equipment user speaking or non-equipment user speaking. By adopting the processing mode, the robustness of voice activity detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, specifically to speech activity detection methods and apparatus, face-to-face translation methods and apparatus, device wake-up methods and apparatus, speaker recognition methods and apparatus, speech activity detection model construction methods and apparatus, and electronic devices. Background Technology

[0002] With the widespread application of smart wearable devices in daily life, open-back headphones, Bluetooth headphones, sports headphones, smart glasses, and AR / VR smart hardware have generally integrated multimodal acoustic sensing systems combining microphone arrays and bone conduction microphones to meet the voice interaction needs in complex acoustic environments. These devices typically need to achieve stable and reliable voice input and interaction functions in high noise, strong winds, mobile conditions, and open spaces, placing higher demands on front-end voice acquisition and processing technologies.

[0003] In typical application scenarios, voice interaction has become one of the core functions of smart glasses. For example, in outdoor sports scenarios (such as running and cycling), environmental noise (such as wind noise and traffic noise) significantly interferes with the sound pickup quality of traditional air conduction microphones, leading to decreased call clarity or even speech recognition failure. Bone conduction microphones, on the other hand, directly acquire speech signals that are highly correlated with vocal cord activity by detecting the mechanical vibration of the skull when the user speaks. They have good resistance to environmental noise, especially showing a significant advantage under conditions of severe wind noise, effectively making up for the shortcomings of air conduction microphones.

[0004] More importantly, current smart devices commonly face the challenge of being accidentally woken up by non-wearers—that is, the device is mistakenly activated by the voices of others in the vicinity, seriously affecting user experience and privacy security. Especially in public places or multi-person conversation environments, relying solely on air conduction microphones is insufficient to effectively distinguish the identity of the target speaker. There is an urgent need to introduce more individual-specific perceptual modalities and discrimination mechanisms to improve the voice interaction experience of smart wearable devices in real-world usage environments. Summary of the Invention

[0005] This application provides a speech activity detection method to address the problem of low robustness in existing speech activity detection technologies. This application also provides a speech activity detection apparatus, a face-to-face translation method and apparatus, a device wake-up method and apparatus, a speaker recognition method and apparatus, a speech activity detection model construction method and apparatus, and an electronic device.

[0006] This application provides a method for detecting voice activity, including: Collect air conduction sound data and bone conduction sound data; Based on air conduction sound data, obtain air conduction speech activity detection results; and based on bone conduction sound data, obtain bone conduction speech activity detection results. Based on the air conduction speech activity detection results and the bone conduction speech activity detection results, a joint speech activity detection result is obtained; the joint speech activity detection result includes: no one speaking, the device user speaking, or a non-device user speaking.

[0007] Optionally, obtaining the air-conducted speech activity detection result based on the air-conducted sound data includes: The speech activity detection model includes an air-conducted speech activity detection module, which obtains the depth features of air-conducted sound based on the air-conducted sound data; and obtains the air-conducted speech activity detection results based on the depth features of the air-conducted sound. The step of obtaining bone conduction speech activity detection results based on bone conduction sound data includes: The speech activity detection model includes a bone conduction speech activity detection module, which obtains the depth features of bone conduction sound data based on the bone conduction sound data; and obtains the bone conduction speech activity detection results based on the depth features of the bone conduction sound.

[0008] Optionally, the air-conducting speech activity detection module included in the speech activity detection model obtains the depth features of the air-conducting sound based on the air-conducting sound data, including: The air-conducted speech activity detection module includes an air-conducted sound depth feature extraction submodule, which obtains the depth features of air-conducted sound based on the air-conducted sound data. The speech activity detection model includes a bone conduction speech activity detection module, which obtains depth features of bone conduction sound based on bone conduction sound data, including: The bone conduction speech activity detection module includes a bone conduction sound depth feature extraction submodule, which obtains the depth features of bone conduction sound based on the bone conduction sound data. Both the air-conducting sound depth feature extraction submodule and the bone-conducting sound depth feature extraction submodule adopt a neural network structure; the neural network size of the air-conducting sound depth feature extraction submodule is smaller than that of the bone-conducting sound depth feature extraction submodule.

[0009] Optionally, the air-conducted speech activity detection module included in the speech activity detection model obtains the air-conducted speech activity detection results based on the depth features of the air-conducted sound, including: The air conduction speech activity detection module includes a speech activity probability acquisition submodule, which acquires the air conduction speech activity probability based on the depth features of the air conduction sound. The smooth decision submodule included in the air-conducted speech activity detection module obtains the air-conducted speech activity detection result based on the air-conducted speech activity probability and the air-conducted speech activity probability threshold. The speech activity detection model includes a bone conduction speech activity detection module, which obtains bone conduction speech activity detection results based on the depth features of bone conduction sound. These results include: The bone conduction speech activity detection module includes a speech activity probability acquisition submodule, which acquires the bone conduction speech activity probability based on the depth features of the bone conduction sound. The smooth decision submodule included in the bone-guided speech activity detection module obtains the bone-guided speech activity detection results based on the bone-guided speech activity probability and the bone-guided speech activity probability threshold. The method further includes: adjusting the air conduction speech activity probability threshold and / or the bone conduction speech activity probability threshold.

[0010] Optionally, the adjustment of the air conduction speech activity probability threshold and / or bone conduction speech activity probability threshold may be performed in at least one of the following ways: If it is a device wake-up scenario, then reduce the bone conduction speech activity probability threshold and / or increase the air conduction speech activity probability threshold; In face-to-face translation scenarios, increase the bone conduction speech activity probability threshold and / or decrease the air conduction speech activity probability threshold.

[0011] Optionally, the acquisition of air conduction sound data includes: acquiring multiple channels of air conduction sound data through an air conduction microphone array; The method further includes: Based on multi-channel air-conducting sound data, the user's voice is enhanced. The step of obtaining the depth features of air-conducted sound based on air-conducted sound data includes: Depth features of air conduction sound are obtained based on enhanced air conduction sound data from equipment users and air conduction sound data from non-equipment users.

[0012] Optionally, the acquisition of air conduction sound data includes: acquiring multiple channels of air conduction sound data through an air conduction microphone array; The method further includes: Based on multi-channel air-conducted sound data, different degrees of voice enhancement are applied to the voice of equipment users and the voice of non-equipment users; The step of obtaining the depth features of air-conducted sound based on air-conducted sound data includes: Depth features of air conduction sound are obtained based on enhanced air conduction sound data from equipment users and non-equipment users.

[0013] Optionally, the step of enhancing the voice of the device user and the voice of non-device users to different degrees based on multi-channel air-conduction sound data includes: Through the first beamforming, voice enhancement is performed in the direction of the device user based on multi-channel air-guided sound data; By using second beamforming, voice enhancement is performed in the direction of non-equipment users based on multi-channel air-guided sound data; The angle targeted by the first beamforming towards the device user is different from the angle targeted by the second beamforming towards the non-device user.

[0014] Optionally, the objective function of beamforming includes a white noise gain constraint term and a null direction control term. By optimizing the weight vectors in multiple beam pointing directions, the output energy of beamforming can be reduced while maintaining the integrity of the desired signal in the target direction.

[0015] This application provides a method for constructing a speech activity detection model, including: Obtain the training dataset; the training data includes: air conduction sound data, bone conduction sound data, air conduction speech activity detection result annotation data, and bone conduction speech activity detection result annotation data; the air conduction speech activity detection result annotation data includes whether someone is speaking or not; the bone conduction speech activity detection result annotation data includes whether someone is speaking or not. A speech activity detection model is learned from the training dataset; the model includes an air-conducting speech activity detection module and a bone-conducting speech activity detection module; the air-conducting speech activity detection module is used to obtain the depth features of air-conducting sound data and obtain the air-conducting speech activity detection result based on the depth features of air-conducting sound; the bone-conducting speech activity detection module is used to obtain the depth features of bone-conducting sound data and obtain the bone-conducting speech activity detection result based on the depth features of bone-conducting sound.

[0016] Optionally, the air conduction sound data includes multiple air conduction sound data; The method further includes: Based on multi-channel air-conducted sound data, different degrees of voice enhancement are applied to the voice of equipment users and the voice of non-equipment users; The step of obtaining the depth features of air-conducted sound based on air-conducted sound data includes: Depth features of air conduction sound are obtained based on enhanced air conduction sound data from equipment users and non-equipment users.

[0017] This application provides a face-to-face translation method, including: Collect air conduction sound data and bone conduction sound data; Based on air conduction sound data, obtain air conduction speech activity detection results; and based on bone conduction sound data, obtain bone conduction speech activity detection results. If the air conduction speech activity detection result indicates that someone is speaking, and the bone conduction speech activity detection result indicates that no one is speaking, then the joint speech activity detection result is determined to be that the non-device user is speaking; based on the air conduction sound data, the translation data of the non-device user's speech content is obtained.

[0018] This application provides a device wake-up method, including: Collect air conduction sound data and bone conduction sound data; Based on air conduction sound data, obtain air conduction speech activity detection results; and based on bone conduction sound data, obtain bone conduction speech activity detection results. If both the air conduction speech activity detection result and the bone conduction speech activity detection result indicate that someone is speaking, then the combined speech activity detection result is determined to be the device user speaking; device wake-up recognition is performed based on the air conduction sound data.

[0019] This application provides a method for detecting voice activity, including: Collect air conduction sound data and bone conduction sound data; The feature processing module included in the speech activity detection model obtains deep fusion features of air conduction sound and bone conduction sound based on air conduction sound data and bone conduction sound data. The voice activity detection model includes a classification module, which obtains voice activity detection results based on the deep fusion features. The voice activity detection results include: no one speaking, device user speaking, or non-device user speaking.

[0020] Optionally, the feature processing module included in the speech activity detection model obtains deep fusion features of air conduction sound and bone conduction sound based on air conduction sound data and bone conduction sound data, including: The feature processing module includes an air conduction feature extraction submodule, which obtains the depth features of air conduction sound based on the air conduction sound data. The feature processing module includes a bone conduction feature extraction submodule, which obtains the depth features of bone conduction sound based on the bone conduction sound data. The feature processing module includes a feature fusion submodule, which obtains the depth fusion features of air-conducted sound and bone-conducted sound based on the depth features of air-conducted sound and bone-conducted sound.

[0021] This application provides a voice activity detection device, including: The sound data acquisition unit is used to acquire air conduction sound data and bone conduction sound data; The two detection units are used to obtain air conduction speech activity detection results based on air conduction sound data and bone conduction speech activity detection results based on bone conduction sound data. The joint decision-making unit is used to obtain a joint speech activity detection result based on the air-guided speech activity detection result and the bone-guided speech activity detection result; the joint speech activity detection result includes: no one speaking, the device user speaking, or a non-device user speaking.

[0022] This application provides an electronic device, including: Processor; and A memory for storing a program for implementing the method described in any of the preceding methods, wherein the device is powered on and the program of the method is executed by the processor.

[0023] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the various methods described above.

[0024] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the various methods described above.

[0025] Compared with the prior art, this application has the following advantages: The speech activity detection method provided in this application collects air-conducted sound data and bone-conducted sound data; obtains air-conducted speech activity detection results based on the air-conducted sound data; obtains bone-conducted speech activity detection results based on the bone-conducted sound data; obtains a combined speech activity detection result based on the air-conducted speech activity detection results and the bone-conducted speech activity detection results; the combined speech activity detection result includes: no one speaking, the device user speaking, or a non-device user speaking. This processing method allows for the differentiation between someone speaking and no one speaking based on air-conducted sound data, and when the air-conducted speech activity detection result indicates someone speaking, it determines whether the speaker is a non-device user or a device user based on the bone-conducted speech activity detection result, thereby achieving speech activity detection that can distinguish the speaker's identity. Compared to speech activity detection that relies solely on bone-conducted sound, this processing method allows for enhanced detection of device users based on air-conducted sound even when bone-conducted sound is weak (e.g., the device user speaks at a near-whisper volume with minimal skull vibration); therefore, it can effectively reduce the false negative rate for device users. Compared to speech activity detection that relies solely on air conduction sound, this processing method enables enhanced detection of device users based on bone conduction sound in complex acoustic environments. Therefore, it effectively reduces the occurrence of misidentifying non-device users as device users in complex acoustic environments. In summary, the speech activity detection method provided in this application can effectively improve the robustness of speech activity detection, thereby enhancing the voice interaction experience of the device in real-world usage environments.

[0026] The speech activity detection model construction method provided in this application embodiment obtains a training dataset. The training data includes: air-conducted sound data, bone-conducted sound data, and labeled air-conducted speech activity detection results. The labeled air-conducted speech activity detection results include whether someone is speaking or not. The labeled bone-conducted speech activity detection results include whether someone is speaking or not. A speech activity detection model is learned from the training dataset. The model includes an air-conducted speech activity detection module and a bone-conducted speech activity detection module. The air-conducted speech activity detection module is used to obtain the depth features of the air-conducted sound based on the air-conducted sound data and to obtain the air-conducted speech activity detection results based on the depth features of the air-conducted sound. The bone-conducted speech activity detection module is used to obtain the depth features of the bone-conducted sound based on the bone-conducted sound data and to obtain the bone-conducted speech activity detection results based on the depth features of the bone-conducted sound. This processing approach enables the fusion modeling of air conduction and bone conduction sound characteristics. The constructed end-to-end model includes speech activity detection based on bone conduction sound and speech activity detection based on air conduction sound. The two detection results are fused to make speaker identity decisions, achieving lightweight speaker identity discrimination. Therefore, it can effectively reduce the computational resource consumption of the model and is suitable for devices with limited computing resources.

[0027] The face-to-face translation method provided in this application collects air-conducted sound data and bone-conducted sound data; obtains air-conducted speech activity detection results based on the air-conducted sound data; and obtains bone-conducted speech activity detection results based on the bone-conducted sound data. If the air-conducted speech activity detection result indicates someone is speaking and the bone-conducted speech activity detection result indicates no one is speaking, then the combined speech activity detection result is determined to be that the speaker is not the device user. Translation data for the non-device user's speech content is obtained based on the air-conducted sound data. This processing method enables face-to-face translation when a device user is speaking face-to-face with another person. It distinguishes between someone speaking and no one speaking based on the air-conducted sound data. If the air-conducted speech activity detection result indicates someone is speaking, it determines whether the speaker is a non-device user or the device user based on the bone-conducted speech activity detection result. If the determination result is that the speaker is a non-device user, then the air-conducted sound is the non-device user's voice, and the speech is translated, thus achieving face-to-face translation. Compared to face-to-face translation that relies solely on air-conducted sound, the face-to-face translation method provided in this application allows for the determination of whether the speaker is a non-device user or a device user based on the bone-conducted sound activity detection result when the air-conducted speech activity detection result indicates that someone is speaking. This effectively reduces the occurrence of misidentifying non-device users as device users in complex acoustic environments, thereby improving the accuracy of face-to-face translation, i.e., improving the robustness of face-to-face translation, and ultimately enhancing the voice interaction experience of the device in real-world usage environments.

[0028] The device wake-up method provided in this application collects air-conduction sound data and bone-conduction sound data; obtains air-conduction speech activity detection results based on the air-conduction sound data; obtains bone-conduction speech activity detection results based on the bone-conduction sound data; if both the air-conduction and bone-conduction speech activity detection results indicate someone is speaking, the combined speech activity detection result is determined to be the device user speaking, and device wake-up recognition is performed based on the air-conduction sound data. This processing method allows for the differentiation between speaking and non-speaking based on air-conduction sound data when the device is in sleep mode. When the air-conduction speech activity detection result indicates someone is speaking, the bone-conduction speech activity detection result is used to determine whether the speaker is a non-device user or the device user. If the determination result is the device user, the air-conduction sound is the device user's voice, and wake-up word recognition is performed on this sound segment, thereby achieving device wake-up. Compared to device wake-up relying solely on bone conduction sound, the device wake-up method provided in this application allows for enhanced detection of the device user based on air conduction sound even when bone conduction sound is weak (e.g., when the device user speaks at a near-whisper volume with minimal skull vibration). Therefore, it effectively reduces the false detection rate of device wake-up, thereby improving the robustness of device wake-up and ultimately enhancing the voice interaction experience in real-world usage environments. Furthermore, compared to device wake-up relying solely on air conduction sound, the device wake-up method provided in this application allows for the determination of whether the speaker is a non-device user or a device user based on the bone conduction sound activity detection result when the air conduction voice activity detection result indicates someone is speaking. This effectively reduces the occurrence of misidentifying non-device users as device users in complex acoustic environments, thus lowering the false detection rate of device wake-up, improving the robustness of device wake-up, and ultimately enhancing the voice interaction experience in real-world usage environments.

[0029] The speech activity detection method provided in this application collects air-conducted sound data and bone-conducted sound data; through the feature processing module included in the speech activity detection model, it obtains deep fusion features of air-conducted sound and bone-conducted sound based on the air-conducted sound data and bone-conducted sound data; through the classification module included in the speech activity detection model, it obtains speech activity detection results based on the deep fusion features, and the speech activity detection results include: no one speaking, the device user speaking, or a non-device user speaking. This processing method integrates air-conducted and bone-conducted sound to achieve speech activity detection that can distinguish the speaker's identity. Compared to speech activity detection that relies solely on bone-conducted sound, this processing method allows for enhanced detection of the device user based on air-conducted sound even when bone-conducted sound is weak (e.g., the device user speaks at a near-whisper volume with slight skull vibration); therefore, it can effectively reduce the false negative rate for device users. Compared to speech activity detection that relies solely on air-conducted sound, this processing method allows for enhanced detection of the device user based on bone-conducted sound in complex acoustic environments; therefore, it can effectively reduce the occurrence of misidentifying non-device users as device users in complex acoustic environments. In summary, the voice activity detection method provided in this application can effectively improve the robustness of voice activity detection, thereby enhancing the voice interaction experience of the device in real-world usage environments. Attached Figure Description

[0030] Figure 1 This is a flowchart illustrating an embodiment of the voice activity detection method provided in this application; Figure 2 This is a schematic diagram of the model network structure of an embodiment of the speech activity detection method provided in this application; Figure 3 This is a schematic flowchart illustrating a specific embodiment of the voice activity detection method provided in this application; Figure 4 This is a schematic diagram of the model architecture of an embodiment of the speech activity detection method provided in this application; Figure 5 This is a schematic diagram of beamforming in an embodiment of the voice activity detection method provided in this application; Figure 6 This is a flowchart illustrating an embodiment of the device wake-up method provided in this application. Detailed Implementation

[0031] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0032] This application provides a method and apparatus for voice activity detection, a face-to-face translation method and apparatus, a device wake-up method and apparatus, a speaker recognition method and apparatus, a voice activity detection model construction method and apparatus, and an electronic device. The various solutions are described in detail below in each embodiment.

[0033] First Embodiment Please refer to Figure 1 This is a flowchart of the speech activity detection method of this application. In this embodiment, the method may include the following steps: Step S101: Collect air conduction sound data and bone conduction sound data.

[0034] The speech activity detection method provided in this application can be used in devices that integrate air conduction and bone conduction pickup devices, such as wearable devices like smart glasses and Bluetooth headsets. This method collects air conduction sound data through the air conduction pickup device and bone conduction sound data simultaneously through the bone conduction pickup device. The speech activity detection device used to implement this method can be a front-end speech acquisition and processing device of the device.

[0035] Air-conducted sound pickup devices can be single-channel air-conducted microphones, acquiring single-channel air-conducted sound data. Alternatively, they can be air-conducted microphone arrays, acquiring multiple channels of air-conducted sound data. Air-conducted speech (ACS) data refers to the speech signal received by the inner ear after sound travels through the air, passing through the outer and middle ear. ACS is susceptible to environmental noise, wind noise, and interference from other people's speech, but it can reflect the complete spectral characteristics of speech. For example, ACS can be acquired by a miniature microphone array placed on the frame of smart glasses.

[0036] Bone conduction microphones are also known as bone conduction devices. Bone-conducted speech (BCS) refers to the speech signal transmitted directly to the inner ear through the skull during speech, caused by the vibration of the vocal cords. BCS is collected by bone conduction sensors attached to bone regions of the head (such as the temporal region or behind the ear). It has strong speaker-specificity and is not sensitive to external environmental noise, but it suffers from attenuation in high-frequency components, resulting in lower speech intelligibility.

[0037] Step S103: Obtain air conduction speech activity detection results based on air conduction sound data; and obtain bone conduction speech activity detection results based on bone conduction sound data.

[0038] Voice Activity Detection (VAD) is used to determine whether valid speech activity exists in an audio signal. This step temporally aligns the information from both air-guided speech (ACS) and bone-guided speech (BCS) perceptual modalities and performs VAD detection in two paths. One path performs VAD detection based on air-guided sound data, and the detected air-guided speech activity result can be either no one is speaking or someone is speaking. The other path performs VAD detection based on bone-guided sound data, and the detected bone-guided speech activity result can also be either no one is speaking or someone is speaking.

[0039] In practice, the results of air conduction speech activity detection can be obtained based on air conduction sound data, and the results of bone conduction speech activity detection can be obtained based on bone conduction sound data. Both can be achieved using relatively mature existing technologies, which will not be elaborated here.

[0040] In one example, the alignment between air-conducted sound data and bone-conducted sound data can include amplifying the energy of the bone-conducted sound data. Bone-conducted and air-conducted sounds typically differ significantly in energy, and this large difference can hinder model learning. Amplifying the energy of the bone-conducted sound data can align this energy difference, thereby reducing the learning difficulty of the model. Specifically, the average power of the bone-conducted and air-conducted sounds can be calculated using an average RMS algorithm, and the bone-conducted sound can be amplified.

[0041] In one example, the alignment between air-conducted sound data and bone-conducted sound data may further include: aligning the bone-conducted sound data and air-conducted sound data in the time domain. Bone-conducted signals and air-conducted signals typically have slight phase differences; the delay difference between the two is calculated based on a cross-correlation algorithm, and then aligned in the time domain.

[0042] Step S105: Obtain a joint speech activity detection result based on the air conduction speech activity detection result and the bone conduction speech activity detection result; the joint speech activity detection result includes: no one speaking, the device user speaking, or a non-device user speaking.

[0043] Air-conducted sound and bone-conducted sound differ in acoustic characteristics; therefore, VAD detection results based on air-conducted sound are complementary to those based on bone-conducted sound. Air-conducted sound is susceptible to environmental noise, wind noise, and interference from other people's speech, but it reflects the complete speech spectrum characteristics. Therefore, air-conducted sound can more comprehensively identify whether someone is speaking. If someone is detected speaking based on air-conducted sound, the speaker could be either a device user or a non-device user. Bone-conducted sound has strong speaker-specificity, is insensitive to external environmental noise, and experiences attenuation in high-frequency components, resulting in lower speech intelligibility. If someone is detected speaking based on bone-conducted sound, the speaker is a device user; if the detection result is no one speaking, there are two possibilities: no one is speaking at all, or a non-device user is speaking. Thus, bone-conducted sound can perform VAD detection for device users, while air-conducted sound can perform VAD detection for all users, and the two detection results are complementary.

[0044] This step combines the air-conduction speech activity detection results obtained from air-conduction sound data with the bone-conduction speech activity detection results obtained from bone-conduction data. Leveraging the complementarity of these two detection results, a joint decision is made to obtain a combined speech activity detection result. This combined speech activity detection result can indicate no one is speaking, the device user is speaking, or a non-device user is speaking. This achieves discriminative VAD detection, meaning it can not only detect the presence of speech but also distinguish whether the speech originates from the device user (device wearer) or a non-device user (someone other than the device wearer, referred to as a non-wearer). Based on the discriminative VAD detection results, applications requiring speaker identification, such as false wake-up suppression and face-to-face translation, can be supported.

[0045] In specific implementation, step S105 can adopt at least one of the following methods: if the air conduction speech activity detection result is that someone is speaking and the bone conduction speech activity detection result is that someone is speaking, then the combined speech activity detection result is determined to be that the device user is speaking; if the air conduction speech activity detection result is that someone is speaking and the bone conduction speech activity detection result is that no one is speaking, then the combined speech activity detection result is determined to be that no one is speaking; if the air conduction speech activity detection result is that no one is speaking and the bone conduction speech activity detection result is that no one is speaking, then the combined speech activity detection result is determined to be that no one is speaking.

[0046] In one example, the air-guided speech activity detection based on air-guided sound data in step S103 can be implemented as follows: Using the air-guided speech activity detection module included in the speech activity detection model, the depth features of the air-guided sound are obtained based on the air-guided sound data; and the air-guided speech activity detection result is obtained based on the depth features of the air-guided sound. Similarly, the bone-guided speech activity detection based on bone-guided sound data in step S103 can be implemented as follows: Using the bone-guided speech activity detection module included in the speech activity detection model, the depth features of the bone-guided sound are obtained based on the bone-guided sound data; and the bone-guided speech activity detection result is obtained based on the depth features of the bone-guided sound. This processing method enables the fusion modeling of air-guided and bone-guided signal characteristics. The end-to-end speech activity detection model includes dual-path VADs (air-guided VAD and bone-guided VAD), achieving semantic-level speaker identification. Since dual-channel VADs rely on different types of sound data, and VAD models that rely on single-type sound data are smaller in scale, the end-to-end speech activity detection model is a lightweight VAD fusion and discrimination network, which can effectively reduce the computational load of the model and is suitable for devices with limited computing resources.

[0047] In one example, the method provided in this application embodiment may further include the following steps: obtaining acoustic features of air-conducted sound data based on air-conducted sound data; and obtaining acoustic features of bone-conducted sound data based on bone-conducted sound data. Correspondingly, obtaining depth features of air-conducted sound based on air-conducted sound data using the air-conducted speech activity detection module included in the speech activity detection model can be implemented as follows: obtaining depth features of air-conducted sound based on the acoustic features of air-conducted sound data using the air-conducted speech activity detection module included in the speech activity detection model; obtaining depth features of bone-conducted sound based on bone-conducted sound data using the bone-conducted speech activity detection module included in the speech activity detection model can be implemented as follows: obtaining depth features of bone-conducted sound based on the acoustic features of bone-conducted sound data using the bone-conducted speech activity detection module included in the speech activity detection model. The acoustic features can be existing acoustic features, such as log-Mel-spectrogram features. In specific implementations, the acoustic features of the sound data can be obtained through short-time Fourier transform (STFT). This processing method makes the input data of the end-to-end model acoustic features, thus effectively reducing the model complexity.

[0048] In one example, the method provided in this application embodiment may further include one or more of the following steps: 1) performing time delay estimation (TDS) on the air conduction sound data; 2) preprocessing the air conduction sound data, the preprocessing including but not limited to one or more of the following processes: removing power frequency, pre-emphasis, and windowing.

[0049] In one example, the air-conduction speech activity detection module included in the speech activity detection model obtains the depth features of air-conduction sound based on the air-conduction sound data. This can be achieved as follows: the air-conduction sound depth feature extraction submodule within the air-conduction speech activity detection module extracts the depth features of the air-conduction sound based on the air-conduction sound data. Similarly, the bone-conduction speech activity detection module included in the speech activity detection model obtains the depth features of bone-conduction sound based on the bone-conduction sound data. This can be achieved as follows: the bone-conduction sound depth feature extraction submodule within the bone-conduction speech activity detection module extracts the depth features of the bone-conduction sound based on the bone-conduction sound data. Both the air-conduction sound depth feature extraction submodule and the bone-conduction sound depth feature extraction submodule employ neural network structures; the neural network size of the air-conduction sound depth feature extraction submodule is smaller than that of the bone-conduction sound depth feature extraction submodule. Because bone-conduction sound data has a natural ability to suppress environmental noise (especially non-wearer speech and background interference), its signal-to-noise ratio is high. Therefore, the end-to-end model can significantly reduce complexity while maintaining high robustness.

[0050] In practical implementation, the air-guided sound depth feature extraction submodule can be implemented using a convolutional-recurrent network or other network architectures. The bone-guided sound depth feature extraction submodule can adopt a network architecture similar in structure to the air-guided sound depth feature extraction submodule but with a smaller scale. For example, the air-guided sound depth feature extraction submodule uses a convolutional-recurrent network architecture, first extracting local features along the frequency axis through a one-dimensional convolutional layer (1D Convolution) to capture the distribution pattern of speech in the frequency domain; then, a gated recurrent unit (GRU) is connected to model the temporal dynamic characteristics of the speech signal. The number of parameters in the air-guided sound depth feature extraction submodule is less than 58,000. The bone-guided sound depth feature extraction submodule can adopt a convolutional-recurrent network architecture similar in structure to the air-guided sound depth feature extraction submodule but with a smaller scale, including 1D convolutional layers and GRU layers, to learn the temporal-frequency features and temporal dependencies of bone-guided speech. The number of parameters in the bone-guided sound depth feature extraction submodule is less than 5,000, significantly reducing computational overhead while ensuring detection performance.

[0051] In one example, the acoustic feature dimension of bone conduction sound data is smaller than that of air conduction sound data; for instance, the acoustic feature dimension of bone conduction sound data is 32 dimensions, while that of air conduction sound data is 64 dimensions. This approach effectively reduces the size of the bone conduction sound depth feature extraction submodule, thus further reducing model complexity.

[0052] In one example, the air-conducted speech activity detection module included in the speech activity detection model obtains the air-conducted speech activity detection result based on the depth features of the air-conducted sound. This can be achieved as follows: the speech activity probability acquisition submodule included in the air-conducted speech activity detection module acquires the air-conducted speech activity probability based on the depth features of the air-conducted sound; the smoothing decision submodule included in the air-conducted speech activity detection module acquires the air-conducted speech activity detection result based on the air-conducted speech activity probability and the air-conducted speech activity probability threshold. Correspondingly, the bone-conducted speech activity detection module included in the speech activity detection model obtains the bone-conducted speech activity detection result based on the depth features of the bone-conducted sound. This can be achieved as follows: the speech activity probability acquisition submodule included in the bone-conducted speech activity detection module acquires the bone-conducted speech activity probability based on the depth features of the bone-conducted sound; the smoothing decision submodule included in the bone-conducted speech activity detection module acquires the bone-conducted speech activity result based on the bone-conducted speech activity probability and the bone-conducted speech activity probability threshold. Accordingly, the method provided in this application embodiment may further include the following steps: adjusting the air-conducted speech activity probability threshold and / or the bone-conducted speech activity probability threshold. In practice, based on the discrimination results of dual-channel VADs, thresholds and smoothing strategies can be set to obtain the wearer and non-wearer discrimination results for each frame. This processing method allows for control over the accuracy of speech activity detection by adjusting at least one of the air-conduction speech activity probability threshold and the bone-conduction speech activity probability threshold.

[0053] In one example, adjusting the air conduction speech activity probability threshold and / or bone conduction speech activity probability threshold can be achieved as follows: in a device wake-up scenario, decrease the bone conduction speech activity probability threshold and / or increase the air conduction speech activity probability threshold. This approach reduces the false negative rate for device users by decreasing the bone conduction speech activity probability threshold, and reduces the false positive rate for non-device users being identified as device users by increasing the air conduction speech activity probability threshold; therefore, it can effectively improve the detection accuracy in device wake-up scenarios.

[0054] Figure 2 A schematic diagram of the network structure of the speech activity detection model is shown, where VPU represents the collected bone conduction sound data and Mic represents the collected air conduction sound data. Figure 2As can be seen, the speech activity detection process includes the following steps: 1) Preprocessing bone conduction and air conduction sound data through a preprocessing module; 2) Extracting acoustic features (such as log-Mel spectrograms) from the sound data through an acoustic feature extraction module. The acoustic features of bone conduction sound data are 32-dimensional features, and the acoustic features of air conduction sound data are 64-dimensional features; 3) Inputting the 64-dimensional air conduction acoustic features into the air conduction speech activity detection module (air conduction VAD module), extracting local features along the frequency axis through four one-dimensional convolutional layers to capture the distribution pattern of speech in the frequency domain; then connecting two gated recurrent units (GRUs) to model the temporal dynamic characteristics of the speech signal, thereby extracting the depth features of bone conduction sound; finally, completing the classification decision through two fully connected layers: the hidden layer uses the ReLU activation function, and the output layer uses the Sigmoid activation function, generating a range of [0, The bone-guided speech activity probability value [1] enables fine-grained judgment between speech segments and non-speech segments; the model has approximately 58,000 parameters and is denoted as AIR-VAD. Another path inputs 32-dimensional bone-guided acoustic features into the bone-guided speech activity detection module (bone-guided VAD module). This module employs a convolutional-recurrent network architecture similar to AIR-VAD but smaller in scale, including two 1D convolutional layers and two GRU layers, to learn the time-frequency features and temporal dependencies of bone-guided speech. Finally, a lightweight fully connected layer outputs the VAD decision result. Because bone-guided signals naturally suppress environmental noise (especially non-wearer speech and background interference), their signal-to-noise ratio is high, thus the speech activity detection model can significantly reduce complexity while maintaining high robustness. The bone-guided VAD module (BC-VAD model) has only about 5,000 parameters, significantly reducing computational overhead while ensuring detection performance.

[0055] In one example, step S101 involves acquiring multiple air-conduction sound data via an air-conduction microphone array. Correspondingly, the method provided in this application embodiment may further include the following step: enhancing the device user's speech based on the multiple air-conduction sound data. Correspondingly, the air-conduction speech activity detection module included in step S103, which obtains the depth features of the air-conduction sound based on the air-conduction sound data, can be implemented as follows: The air-conduction speech activity detection module included in the speech activity detection model obtains the depth features of the air-conduction sound based on the enhanced air-conduction sound data of the device user and the air-conduction sound data of non-device users. This processing method enhances the device user's speech, enabling better identification of the wearer even in complex acoustic environments using a lightweight speech activity detection model. Therefore, it effectively improves the robustness of wearer speech activity detection while maintaining a small model size, making it suitable for device wake-up and other processing on resource-constrained end-user devices.

[0056] In specific implementation, the voice enhancement of the device user's speech based on multi-channel air-guided sound data can be achieved in the following way: Through beamforming, voice enhancement is performed in the direction of the device user based on the multi-channel air-guided sound data. Correspondingly, the air-guided voice activity detection module included in the voice activity detection model obtains the depth features of the air-guided sound based on the enhanced air-guided sound data in the direction of the device user and the air-guided sound data in the direction of non-device users. For example, a Minimum Variance Distortionless Response (MVDR) beamformer can be used. In specific implementation, other voice enhancement techniques can also be used.

[0057] Please refer to Figure 3This is a schematic diagram illustrating the specific flow of the speech activity detection method of this application. In one example, the method provided in this embodiment may further include the following step S301: enhancing the speech of the device user and the speech of the non-device user to different degrees based on multiple air-conducting sound data. This processing method ensures that the first difference is greater than the second difference. The first difference is the difference between the enhanced air-conducting sound data of the device user and the enhanced air-conducting sound data of the non-device user, and the second difference is the difference between the unenhanced air-conducting sound data of the device user and the unenhanced air-conducting sound data of the non-device user. Correspondingly, the air-conducting speech activity detection module included in the speech activity detection model in step S103, which obtains the depth features of the air-conducting sound based on the air-conducting sound data, can be implemented as follows: obtaining the depth features of the air-conducting sound based on the enhanced air-conducting sound data of the device user and the enhanced air-conducting sound data of the non-device user. This approach allows for speech enhancement for both device users and non-device users, with different enhancement intensities applied to amplify the differences between the two speech types (such as differences in spectrum and energy). This enables a lightweight speech activity detection model to better identify wearers and non-wearers even in complex acoustic environments. Therefore, it effectively improves the robustness of wearer speech activity detection while maintaining a relatively small model size, making it suitable for device wake-up, face-to-face translation, and other processing tasks on resource-constrained devices.

[0058] In practice, the first beamforming enhances speech in the direction of the device user based on multiple air-conducted sound data; the second beamforming enhances speech in the direction other than the device user based on the same multiple air-conducted sound data. The angle of the first beamforming targeting the device user direction is set to be greater than or less than the angle of the second beamforming targeting the non-device user direction. This ensures that the first difference between the enhanced air-conducted sound data in the device user direction and the enhanced air-conducted sound data in the non-device user direction is greater than the second difference between the unenhanced air-conducted sound data in the device user direction and the unenhanced air-conducted sound data in the non-device user direction. For example, the angle of the first beamforming targeting the device user direction can be set to a small angle relative to the user's mouth, and the angle of the second beamforming targeting the non-device user direction can be set to a 30-degree area in front of the user. Other speech enhancement techniques can also be used in practice.

[0059] Figure 4 This diagram illustrates the model architecture for voice activity detection. Figure 4As can be seen, the voice activity detection process includes the following steps: 1) Acquiring 5 channels of air conduction sound data (Mic1-5) through a 5-channel air conduction microphone array; acquiring bone conduction sound data (VPU) through a bone conduction microphone; 2) Preprocessing the bone conduction sound data and air conduction sound data through a preprocessing module; 3) Using a beamforming module, performing voice enhancement in the direction of the device user and the direction of the non-device user respectively, with the intensity of voice enhancement for the device user being greater than that for the non-device user, thus increasing the difference in spectrum and energy between the device user's voice and the non-device user's voice, so that even in complex acoustic environments, a lightweight voice activity detection model can better identify the wearer and non-wearer; 4) Extracting acoustic features from the sound data through an acoustic feature extraction module. 4) Input the air-conducting acoustic features into the air-conducting speech activity detection module (air-conducting VAD module) on one side. The air-conducting speech activity detection module outputs the air-conducting speech activity probability. Through the air-conducting smoothing decision model, the air-conducting speech activity detection result is obtained based on the air-conducting speech activity probability and the air-conducting speech activity probability threshold. Input the bone-conducting acoustic features into the bone-conducting speech activity detection module (bone-conducting VAD module) on the other side. The bone-conducting speech activity detection module outputs the bone-conducting speech activity probability. Through the bone-conducting smoothing decision model, the bone-conducting speech activity detection result is obtained based on the bone-conducting speech activity probability and the bone-conducting speech activity probability threshold. 5) Through fusion decision, based on the bone-conducting speech activity detection result and the air-conducting speech activity detection result, determine whether someone is speaking and whether the speaker is a device user or a non-device user.

[0060] In one example, multiple fixed beamformers are used: one is specifically pointed toward the mouth of the device user (wearer), and the others correspond to different directions in the horizontal direction around the device. Figure 5 A schematic diagram of beamforming in this embodiment is shown. For the desired beam area (including: the device user and non-device users speaking face-to-face with the device user), in a 360-degree direction centered on the device user, a "first beam targeting the device user's mouth (device user direction)" and a "second beam targeting the area in front of the device user, within ±30 degrees in front (non-device user direction speaking face-to-face with the device user)" are designed for enhancement. At the same time, beam nulling is applied to the undesired area (including those interfering with the speaker).

[0061] In one example, a Linear Constrained Minimum Variance (LCMV) beamforming algorithm is employed, aiming to reduce the beamformer's output energy while maintaining the integrity of the desired signal. However, this method lacks explicit control over null directions during optimization, resulting in significant fluctuations in null characteristics at different frequencies. Furthermore, it does not consider white noise gain constraints, making it susceptible to broadband noise and limiting its robustness.

[0062] In another example, the Non-Linearly Constrained Minimum Variance (NLCMV) optimization criterion is employed. This criterion simultaneously introduces white noise gain constraints and null direction control mechanisms into the objective function, effectively improving the anti-interference capability and frequency stability of beamforming. Specifically, NLCMV optimizes the complex weight vector in each beam pointing direction to minimize the overall output power while satisfying the directional response constraints, thereby achieving high-fidelity enhancement of the target speech (device user and non-device user speaking to the device user) and directional suppression of interference sources (diffuse noise, interfering speaker). In the device wake-up scenario, the beam output specifically pointing towards the wearer's mouth can be selected as the enhanced signal. In the face-to-face translation scenario, the beam output specifically pointing towards the wearer's mouth and the beam output pointing towards the area in front of the wearer can be selected as the two enhanced signals.

[0063] Unlike centralized uniform microphone arrays, microphone arrays in wearable devices (such as smart glasses) can employ distributed, non-uniform structures. Therefore, beam modeling can be performed by combining convex optimization and unequal spacing separation methods. This approach, combined with normalization on top of the LCMV beamformer, results in stronger robustness and tolerance to steering vector errors. While suppressing interference and noise, it maintains a distortion-free response to signals in the target direction (the user's mouth, the area in front of the user), and normalization further enhances stability.

[0064] As can be seen from the above embodiments, the method provided in this application collects air-conducted sound data and bone-conducted sound data; obtains air-conducted speech activity detection results based on the air-conducted sound data; obtains bone-conducted speech activity detection results based on the bone-conducted sound data; obtains a combined speech activity detection result based on the air-conducted speech activity detection results and the bone-conducted speech activity detection results; the combined speech activity detection result includes: no one speaking, the device user speaking, or a non-device user speaking. This processing method allows for the differentiation between someone speaking and no one speaking based on air-conducted sound data, and when the air-conducted speech activity detection result indicates someone speaking, it determines whether the speaker is a non-device user or a device user based on the bone-conducted speech activity detection result, thereby achieving speech activity detection that can distinguish the speaker's identity. Compared to speech activity detection that relies solely on bone-conducted sound, this processing method allows for enhanced detection of device users based on air-conducted sound even when bone-conducted sound is weak (e.g., the device user speaks at a near-whisper volume with minimal skull vibration); therefore, it can effectively reduce the false negative rate for device users. Compared to speech activity detection relying solely on air conduction sound, this processing method enables enhanced detection of device users based on bone conduction sound in complex acoustic environments. Therefore, it effectively reduces the occurrence of misidentifying non-device users as device users in complex acoustic environments. In summary, the speech activity detection method provided in this application can effectively improve the robustness of speech activity detection, thereby enhancing the voice interaction experience of the device in real-world usage environments.

[0065] Second Embodiment In the above embodiments, a voice activity detection method is provided. Correspondingly, this application also provides a voice activity detection device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0066] This application also provides a voice activity detection device, comprising: a sound data acquisition unit for acquiring air conduction sound data and bone conduction sound data; a two-channel detection unit for acquiring air conduction voice activity detection results based on the air conduction sound data; and acquiring bone conduction voice activity detection results based on the bone conduction sound data; and a joint decision unit for acquiring a joint voice activity detection result based on the air conduction voice activity detection result and the bone conduction voice activity detection result; wherein the joint voice activity detection result includes: no one speaking, the device user speaking, or a non-device user speaking.

[0067] In one example, the two detection units are specifically used to obtain the depth features of air-conducted sound based on air-conducted sound data through the air-conducted sound activity detection module included in the speech activity detection model; and to obtain the air-conducted speech activity detection result based on the depth features of air-conducted sound; and to obtain the bone-conducted sound activity detection result based on the bone-conducted sound data through the bone-conducted sound activity detection module included in the speech activity detection model.

[0068] In one example, the speech activity detection model includes an air-conducting speech activity detection module, which obtains depth features of air-conducting sound based on air-conducting sound data. This includes: an air-conducting sound depth feature extraction submodule, which extracts depth features of air-conducting sound based on air-conducting sound data; and a bone-conducting speech activity detection module, which includes a bone-conducting speech activity detection model, which obtains depth features of bone-conducting sound based on bone-conducting sound data. This includes: a bone-conducting sound depth feature extraction submodule, which extracts depth features of bone-conducting sound based on bone-conducting sound data. Both the air-conducting sound depth feature extraction submodule and the bone-conducting sound depth feature extraction submodule employ neural network structures. The neural network size of the air-conducting sound depth feature extraction submodule is smaller than that of the bone-conducting sound depth feature extraction submodule.

[0069] In one example, the speech activity detection model includes an air-conducting speech activity detection module, which obtains air-conducting speech activity detection results based on the depth features of air-conducting sound. This includes: obtaining air-conducting speech activity probability based on the depth features of air-conducting sound using a speech activity probability acquisition submodule; obtaining air-conducting speech activity detection results based on the air-conducting speech activity probability and an air-conducting speech activity probability threshold using a smoothing decision submodule. Similarly, the speech activity detection model includes a bone-conducting speech activity detection module, which obtains bone-conducting speech activity detection results based on the depth features of bone-conducting sound. This includes: obtaining bone-conducting speech activity probability based on the depth features of bone-conducting sound using a speech activity probability acquisition submodule; obtaining bone-conducting speech activity detection results based on the bone-conducting speech activity probability and a bone-conducting speech activity probability threshold using a smoothing decision submodule. The device further includes a threshold adjustment unit for adjusting the air-conducting speech activity probability threshold and / or the bone-conducting speech activity probability threshold.

[0070] In one example, the threshold adjustment unit is specifically used to decrease the bone conduction speech activity probability threshold and / or increase the air conduction speech activity probability threshold in the case of device wake-up; and to increase the bone conduction speech activity probability threshold and / or decrease the air conduction speech activity probability threshold in the case of face-to-face translation.

[0071] In one example, the sound data acquisition unit is specifically used to acquire multiple air-conducting sound data through an air-conducting microphone array; the device further includes: a voice enhancement unit, used to enhance the voice of the device user based on the multiple air-conducting sound data; the step of obtaining the depth features of the air-conducting sound based on the air-conducting sound data includes: obtaining the depth features of the air-conducting sound based on the enhanced air-conducting sound data of the device user and the air-conducting sound data of non-device users.

[0072] In one example, the sound data acquisition unit is specifically used to acquire multiple air-conducting sound data through an air-conducting microphone array; the device further includes: a voice enhancement unit, used to enhance the voice of the device user and the voice of non-device users to different degrees based on the multiple air-conducting sound data; the step of obtaining the depth features of the air-conducting sound based on the air-conducting sound data includes: obtaining the depth features of the air-conducting sound based on the enhanced air-conducting sound data of the device user and the enhanced air-conducting sound data of the non-device user.

[0073] In one example, the step of enhancing the voice of the device user and the voice of the non-device user to different degrees based on multiple air-conducting sound data includes: enhancing the voice in the direction of the device user by means of a first beamforming and enhancing the voice in the direction of the non-device user by means of a second beamforming; wherein the angle of the device user direction targeted by the first beamforming is different from the angle of the non-device user direction targeted by the second beamforming.

[0074] In one example, the objective function of beamforming includes a white noise gain constraint term and a null direction control term. By optimizing the weight vectors in multiple beam pointing directions, the output energy of beamforming is reduced while maintaining the integrity of the desired signal in the target direction.

[0075] Third Embodiment In the above embodiments, a method for detecting speech activity is provided. Correspondingly, this application also provides a method for constructing a speech activity detection model. This method corresponds to the embodiments of the above method, so it is described simply. For relevant details, please refer to the description of the method embodiment one. The method embodiments described below are merely illustrative.

[0076] The speech activity detection model construction method in this embodiment includes the following steps: Step 1: Obtain the training dataset.

[0077] The training dataset includes multiple training data sets. Each training data set may include: air-conducted sound data, bone-conducted sound data, and labeled air-conducted speech activity detection results. The labeled air-conducted speech activity detection results include whether someone is speaking or not; the labeled bone-conducted speech activity detection results also include whether someone is speaking or not. In practice, the air-conducted and bone-conducted speech activity detection results can be labeled according to the actual situation of the speech activity.

[0078] In practice, residual non-stationary noise from real-world scenarios (such as wind noise, traffic noise, and other people's voices) can be introduced during training to improve the model's generalization ability in complex environments.

[0079] Step 2: Learn the speech activity detection model from the training dataset.

[0080] The speech activity detection model includes an air-conduction speech activity detection module and a bone-conduction speech activity detection module. The air-conduction speech activity detection module is used to obtain the depth features of air-conduction sound data and, based on these depth features, obtain the air-conduction speech activity detection result. The bone-conduction speech activity detection module is used to obtain the depth features of bone-conduction sound data and, based on these depth features, obtain the bone-conduction speech activity detection result.

[0081] The speech activity detection model construction method provided in this application adopts a supervised learning approach, using two types of sound data in the training data as the input data of the model and two types of labeled data in the training data as the output data of the model. By adjusting the model parameters, the difference between the two detection results predicted by the model and the two labeled data is reduced, thereby realizing the speech activity detection model learned from the training dataset.

[0082] In one example, the air-conducted sound data includes multiple air-conducted sound data streams; the method further includes: performing speech enhancement on device user speech and non-device user speech to different degrees based on the multiple air-conducted sound data streams; obtaining the depth features of the air-conducted sound based on the air-conducted sound data includes: obtaining the depth features of the air-conducted sound based on the enhanced air-conducted sound data of the device user and the enhanced air-conducted sound data of the non-device user. This processing approach allows for speech enhancement on both device users and non-device users, with different enhancement intensities applied to amplify the differences between the two speech streams (such as differences in spectrum, energy, etc.). This not only reduces the model learning difficulty and allows for faster model construction, but also enables a lightweight speech activity detection model to better identify wearers and non-wearers even in complex acoustic environments. Therefore, it effectively improves the robustness of wearer speech activity detection while maintaining a small model size and low learning difficulty, making it suitable for device wake-up, face-to-face translation, and other processing on resource-constrained devices.

[0083] In practical applications, the trained dual-channel VAD model can be quantized and deployed on a resource-constrained embedded microcontroller (ARM). Leveraging its low parameter count and efficient inference architecture, it meets the real-time processing requirements of small battery-powered devices such as smart glasses (e.g., frame latency less than 3ms). Simultaneously, by combining the target direction enhancement signal output from the pre-stage beamforming module, the overall anti-interference capability of the system is further improved.

[0084] As can be seen from the above embodiments, the speech activity detection model construction method provided in this application obtains a training dataset; the training data includes: air-conducted sound data, bone-conducted sound data, air-conducted speech activity detection result annotation data, and bone-conducted speech activity detection result annotation data; the air-conducted speech activity detection result annotation data includes: whether someone is speaking or not; the bone-conducted speech activity detection result annotation data includes: whether someone is speaking or not; a speech activity detection model is learned from the training dataset; the model includes an air-conducted speech activity detection module and a bone-conducted speech activity detection module; the air-conducted speech activity detection module is used to obtain the depth features of the air-conducted sound based on the air-conducted sound data; and obtain the air-conducted speech activity detection result based on the depth features of the air-conducted sound; the bone-conducted speech activity detection module is used to obtain the depth features of the bone-conducted sound based on the bone-conducted sound data; and obtain the bone-conducted speech activity detection result based on the depth features of the bone-conducted sound. This processing approach enables the fusion modeling of air conduction and bone conduction sound characteristics. The constructed end-to-end model includes speech activity detection based on bone conduction sound and speech activity detection based on air conduction sound. The two detection results are fused to make speaker identity decisions, achieving lightweight speaker identity discrimination. Therefore, it can effectively reduce the computational resource consumption of the model and is suitable for devices with limited computing resources.

[0085] Fourth embodiment In the above embodiments, a method for constructing a speech activity detection model is provided. Correspondingly, this application also provides a device for constructing a speech activity detection model. This device corresponds to the embodiments of the method described above. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0086] This application also provides a speech activity detection model construction device, comprising: a training data acquisition unit and a model learning unit. The training data acquisition unit is used to acquire a training dataset; the training data includes: air-conducted sound data, bone-conducted sound data, air-conducted speech activity detection result annotation data, and bone-conducted speech activity detection result annotation data; the air-conducted speech activity detection result annotation data includes whether someone is speaking or not; the bone-conducted speech activity detection result annotation data includes whether someone is speaking or not; the model learning unit is used to learn a speech activity detection model from the training dataset; the model includes an air-conducted speech activity detection module and a bone-conducted speech activity detection module; the air-conducted speech activity detection module is used to acquire depth features of the air-conducted sound based on the air-conducted sound data; and acquire air-conducted speech activity detection results based on the depth features of the air-conducted sound; the bone-conducted speech activity detection module is used to acquire depth features of the bone-conducted sound based on the bone-conducted sound data; and acquire bone-conducted speech activity detection results based on the depth features of the bone-conducted sound.

[0087] In one example, the air-conducted sound data includes multiple air-conducted sound data; the device further includes: a speech enhancement unit, used to enhance the speech of the device user and the speech of the non-device user to different degrees based on the multiple air-conducted sound data; the step of obtaining the depth features of the air-conducted sound based on the air-conducted sound data includes: obtaining the depth features of the air-conducted sound based on the enhanced air-conducted sound data of the device user and the enhanced air-conducted sound data of the non-device user.

[0088] Fifth embodiment In the above embodiments, a speech activity detection method is provided. Correspondingly, this application also provides a face-to-face translation method. This method corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0089] This application also provides a face-to-face translation method, including the following steps: Step 1: Collect air conduction sound data and bone conduction sound data.

[0090] Step 2: Obtain the air conduction speech activity detection results based on the air conduction sound data; and obtain the bone conduction speech activity detection results based on the bone conduction sound data.

[0091] Step 3: If the air conduction speech activity detection result indicates that someone is speaking and the bone conduction speech activity detection result indicates that no one is speaking, then the joint speech activity detection result is determined to be that the non-device user is speaking; based on the air conduction sound data, obtain the translation data of the non-device user's speech content.

[0092] In one example, obtaining the air-conducted speech activity detection result based on the air-conducted sound data includes: obtaining the depth features of the air-conducted sound based on the air-conducted sound data using the air-conducted speech activity detection module included in the speech activity detection model; and obtaining the air-conducted speech activity detection result based on the depth features of the air-conducted sound. Similarly, obtaining the bone-conducted speech activity detection result based on the bone-conducted sound data includes: obtaining the depth features of the bone-conducted sound based on the bone-conducted sound data using the bone-conducted speech activity detection module included in the speech activity detection model; and obtaining the bone-conducted speech activity detection result based on the depth features of the bone-conducted sound.

[0093] In one example, the acquisition of air-conducted sound data includes: acquiring multiple channels of air-conducted sound data through an air-conducted microphone array; the method further includes: performing different degrees of speech enhancement on the device user's speech and the non-device user's speech based on the multiple channels of air-conducted sound data; the acquisition of depth features of air-conducted sound based on the air-conducted sound data includes: acquiring depth features of air-conducted sound based on the enhanced air-conducted sound data of the device user and the non-device user.

[0094] In one example, the step of enhancing the voice of the device user and the voice of the non-device user to different degrees based on multiple air-conducting sound data includes: enhancing the voice in the direction of the device user by means of a first beamforming and enhancing the voice in the direction of the non-device user by means of a second beamforming; wherein the angle of the device user direction targeted by the first beamforming is different from the angle of the non-device user direction targeted by the second beamforming.

[0095] As can be seen from the above embodiments, the face-to-face translation method provided in this application collects air-conducted sound data and bone-conducted sound data; obtains air-conducted speech activity detection results based on the air-conducted sound data; and obtains bone-conducted speech activity detection results based on the bone-conducted sound data; if the air-conducted speech activity detection result indicates someone is speaking and the bone-conducted speech activity detection result indicates no one is speaking, then the joint speech activity detection result is determined to be that the speaker is not a device user; and translation data of the non-device user's speech content is obtained based on the air-conducted sound data. This processing method enables the translation of the non-device user's speech content during a dialogue between a device user and a non-device user, and can also display the translated content when the speaker is identified as a non-device user. Because it can better distinguish between device users and non-device users, it can effectively improve the accuracy of face-to-face translation.

[0096] Sixth Embodiment In the above embodiments, a face-to-face translation method is provided. Correspondingly, this application also provides a face-to-face translation device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0097] This application also provides a face-to-face translation device, including: a sound data acquisition unit, a two-channel detection unit, a speaker recognition unit, and a translation unit. The sound data acquisition unit is used to acquire air-conduction sound data and bone-conduction sound data; the two-channel detection unit is used to obtain air-conduction speech activity detection results based on the air-conduction sound data; and to obtain bone-conduction speech activity detection results based on the bone-conduction sound data; the speaker recognition unit is used to determine that the joint speech activity detection result is a non-device user speaking if the air-conduction speech activity detection result indicates someone is speaking and the bone-conduction speech activity detection result indicates no one is speaking; the translation unit is used to obtain translation data of the non-device user's speech content based on the air-conduction sound data.

[0098] In one example, obtaining the air-conducted speech activity detection result based on the air-conducted sound data includes: obtaining the depth features of the air-conducted sound based on the air-conducted sound data using the air-conducted speech activity detection module included in the speech activity detection model; and obtaining the air-conducted speech activity detection result based on the depth features of the air-conducted sound. Similarly, obtaining the bone-conducted speech activity detection result based on the bone-conducted sound data includes: obtaining the depth features of the bone-conducted sound based on the bone-conducted sound data using the bone-conducted speech activity detection module included in the speech activity detection model; and obtaining the bone-conducted speech activity detection result based on the depth features of the bone-conducted sound.

[0099] In one example, the acquisition of air-conducted sound data includes: acquiring multiple channels of air-conducted sound data through an air-conducted microphone array; the device further includes: a voice enhancement unit, used to enhance the voice of the device user and the voice of non-device users to different degrees based on the multiple channels of air-conducted sound data; the acquisition of depth features of air-conducted sound based on the air-conducted sound data includes: acquiring depth features of air-conducted sound based on the enhanced air-conducted sound data of the device user and the enhanced air-conducted sound data of non-device users.

[0100] In one example, the step of enhancing the voice of the device user and the voice of the non-device user to different degrees based on multiple air-conducting sound data includes: enhancing the voice in the direction of the device user by means of a first beamforming and enhancing the voice in the direction of the non-device user by means of a second beamforming; wherein the angle of the device user direction targeted by the first beamforming is different from the angle of the non-device user direction targeted by the second beamforming.

[0101] Seventh Embodiment In the above embodiments, a voice activity detection method is provided. Correspondingly, this application also provides a device wake-up method. This method corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0102] This application also provides a device wake-up method, including: Step 1: Collect air conduction sound data and bone conduction sound data.

[0103] Step 2: Obtain the air conduction speech activity detection results based on the air conduction sound data; and obtain the bone conduction speech activity detection results based on the bone conduction sound data.

[0104] Step 3: If the air conduction speech activity detection result is that someone is speaking, and the bone conduction speech activity detection result is that someone is speaking, then the joint speech activity detection result is determined to be that the device user is speaking; based on the air conduction sound data, device wake-up recognition is performed.

[0105] In one example, obtaining the air-conducted speech activity detection result based on the air-conducted sound data includes: obtaining the depth features of the air-conducted sound based on the air-conducted sound data using the air-conducted speech activity detection module included in the speech activity detection model; and obtaining the air-conducted speech activity detection result based on the depth features of the air-conducted sound. Similarly, obtaining the bone-conducted speech activity detection result based on the bone-conducted sound data includes: obtaining the depth features of the bone-conducted sound based on the bone-conducted sound data using the bone-conducted speech activity detection module included in the speech activity detection model; and obtaining the bone-conducted speech activity detection result based on the depth features of the bone-conducted sound.

[0106] In one example, the acquisition of air-conducted sound data includes: acquiring multiple channels of air-conducted sound data through an air-conducted microphone array; the method further includes: performing speech enhancement on the device user's speech based on the multiple channels of air-conducted sound data; the acquisition of depth features of air-conducted sound based on the air-conducted sound data includes: acquiring depth features of air-conducted sound based on the enhanced air-conducted sound data of the device user and the air-conducted sound data of non-device users.

[0107] In another example, the acquisition of air-conducted sound data includes: acquiring multiple channels of air-conducted sound data through an air-conducted microphone array; the method further includes: performing different degrees of speech enhancement on the device user's speech and the non-device user's speech based on the multiple channels of air-conducted sound data; the acquisition of depth features of air-conducted sound based on the air-conducted sound data includes: acquiring depth features of air-conducted sound based on the enhanced air-conducted sound data of the device user and the non-device user.

[0108] In one example, the step of enhancing the voice of the device user and the voice of the non-device user to different degrees based on multiple air-conducting sound data includes: enhancing the voice in the direction of the device user by means of a first beamforming and enhancing the voice in the direction of the non-device user by means of a second beamforming; wherein the angle of the device user direction targeted by the first beamforming is different from the angle of the non-device user direction targeted by the second beamforming.

[0109] In one example, the step of performing device wake-up recognition based on air conduction sound data includes: obtaining a wake-up score based on the air conduction sound data using a wake-up model; and waking up the device if the wake-up score meets the device wake-up conditions.

[0110] like Figure 6 As shown, in specific implementation, before obtaining the wake-up score based on the air conduction sound data through the wake-up model, the device user's voice can be enhanced to obtain the enhanced voice data of the device user. Figure 6The wake-up integrated decision strategy is as follows: if the joint voice activity detection result indicates that the device user is speaking and the wake-up score meets the device wake-up conditions, then the device needs to be woken up. The role of the joint voice activity detection result is to perform wake-up verification; the device can be woken up through the wake-up module.

[0111] As can be seen from the above embodiments, the device wake-up method provided in this application determines that the joint speech activity detection result is that the device user is speaking if both the air conduction speech activity detection result and the bone conduction speech activity detection result indicate that someone is speaking; and performs device wake-up recognition based on the air conduction sound data. This processing method ensures that device wake-up recognition is performed when the speaker is detected as a device user. Because it can better distinguish between device users and non-device users, it can effectively reduce the false negative rate for device users and the false positive rate for non-device users being mistaken for device users.

[0112] Eighth embodiment In the above embodiments, a device wake-up method is provided. Correspondingly, this application also provides a device wake-up apparatus. This apparatus corresponds to the embodiments of the above method. Since the apparatus embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The apparatus embodiments described below are merely illustrative.

[0113] This application also provides a device wake-up device, comprising: a sound data acquisition unit, a two-channel detection unit, a speaker recognition unit, and a device wake-up recognition unit. The sound data acquisition unit is used to acquire air conduction sound data and bone conduction sound data; the two-channel detection unit is used to obtain air conduction speech activity detection results based on the air conduction sound data; and to obtain bone conduction speech activity detection results based on the bone conduction sound data; the speaker recognition unit is used to determine that the joint speech activity detection result is that the device user is speaking if both the air conduction speech activity detection result and the bone conduction speech activity detection result indicate that someone is speaking; and the device wake-up recognition unit is used to perform device wake-up recognition based on the air conduction sound data.

[0114] In one example, obtaining the air-conducted speech activity detection result based on the air-conducted sound data includes: obtaining the depth features of the air-conducted sound based on the air-conducted sound data using the air-conducted speech activity detection module included in the speech activity detection model; and obtaining the air-conducted speech activity detection result based on the depth features of the air-conducted sound. Similarly, obtaining the bone-conducted speech activity detection result based on the bone-conducted sound data includes: obtaining the depth features of the bone-conducted sound based on the bone-conducted sound data using the bone-conducted speech activity detection module included in the speech activity detection model; and obtaining the bone-conducted speech activity detection result based on the depth features of the bone-conducted sound.

[0115] In one example, the acquisition of air-conducted sound data includes: acquiring multiple channels of air-conducted sound data through an air-conducted microphone array; the device further includes: a voice enhancement unit for enhancing the voice of the device user based on the multiple channels of air-conducted sound data; the acquisition of depth features of air-conducted sound based on the air-conducted sound data includes: acquiring depth features of air-conducted sound based on the enhanced air-conducted sound data of the device user and the air-conducted sound data of non-device users.

[0116] In another example, the acquisition of air-conducted sound data includes: acquiring multiple channels of air-conducted sound data through an air-conducted microphone array; the device further includes: a voice enhancement unit, used to enhance the voice of the device user and the voice of non-device users to different degrees based on the multiple channels of air-conducted sound data; the acquisition of depth features of air-conducted sound based on the air-conducted sound data includes: acquiring depth features of air-conducted sound based on the enhanced air-conducted sound data of the device user and the enhanced air-conducted sound data of non-device users.

[0117] In one example, the step of enhancing the voice of the device user and the voice of the non-device user to different degrees based on multiple air-conducting sound data includes: enhancing the voice in the direction of the device user by means of a first beamforming and enhancing the voice in the direction of the non-device user by means of a second beamforming; wherein the angle of the device user direction targeted by the first beamforming is different from the angle of the non-device user direction targeted by the second beamforming.

[0118] Ninth Embodiment In the above embodiments, a method for detecting speech activity is provided. Correspondingly, this application also provides a method for detecting speech activity. This method corresponds to Embodiment 1 of the above method. Since this embodiment is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0119] This application also provides a method for detecting speech activity, including: Step 1: Collect air conduction sound data and bone conduction sound data.

[0120] Step 2: Using the feature processing module included in the speech activity detection model, obtain deep fusion features of air conduction sound and bone conduction sound based on the air conduction sound data and bone conduction sound data.

[0121] Deep fusion features refer to abstract features that combine the characteristics of air-conducted sound and bone-conducted sound, extracted by the feature processing module (such as a neural network structure) using air-conducted sound data and bone-conducted sound data as input data.

[0122] Step 3: Using the classification module included in the voice activity detection model, obtain the voice activity detection results based on the deep fusion features. The voice activity detection results include: no one speaking, device user speaking, or non-device user speaking.

[0123] The method provided in this embodiment differs from Method Embodiment 1 in that the fusion methods for the two types of audio data are different. Method Embodiment 1 fuses the two types of audio data by performing decision-level fusion of air conduction-based speech activity detection results and bone conduction-based speech activity detection results; the speech activity detection model includes two-channel VAD detection. In contrast, the method provided in this embodiment fuses the features of the two types of data by fusing them and identifying the speaker based on the fused features; the speech activity detection model includes only one-channel VAD detection.

[0124] In practice, a speech activity detection model can be learned from the training dataset. The training data includes: air conduction sound data and bone conduction sound data, as well as labeled data for speech activity detection. The labeled data can be speech from an unattended audience, a device user, or a non-device user.

[0125] In one example, the feature processing module included in the speech activity detection model obtains deep fusion features of air-conducted sound and bone-conducted sound based on air-conducted sound data and bone-conducted sound data, including: obtaining depth features of air-conducted sound based on air-conducted sound data through an air-conducted feature extraction submodule included in the feature processing module; obtaining depth features of bone-conducted sound based on bone-conducted sound data through a bone-conducted feature extraction submodule included in the feature processing module; and obtaining deep fusion features of air-conducted sound and bone-conducted sound based on the depth features of air-conducted sound and bone-conducted sound through a feature fusion submodule included in the feature processing module.

[0126] In another example, the feature processing module included in the speech activity detection model obtains deep fusion features of air-conducted sound and bone-conducted sound based on air-conducted sound data and bone-conducted sound data. This includes obtaining deep fusion features of air-conducted sound and bone-conducted sound through multiple feature extraction layers included in the feature processing module. The input data of the first feature extraction layer includes air-conducted sound data and bone-conducted sound data, and the abstract features output by different levels of feature extraction layers are fusion features of different depth levels.

[0127] As can be seen from the above embodiments, the speech activity detection method provided in this application collects air-conducted sound data and bone-conducted sound data; through the feature processing module included in the speech activity detection model, it obtains deep fusion features of air-conducted sound and bone-conducted sound based on the air-conducted sound data and bone-conducted sound data; through the classification module included in the speech activity detection model, it obtains speech activity detection results based on the deep fusion features, and the speech activity detection results include: no one speaking, the device user speaking, or a non-device user speaking. This processing method integrates air-conducted and bone-conducted sound to achieve speech activity detection that can distinguish the speaker's identity. Compared to speech activity detection that relies solely on bone-conducted sound, this processing method allows for enhanced detection of the device user based on air-conducted sound even when bone-conducted sound is weak (e.g., the device user speaks at a near-whisper volume with slight skull vibration); therefore, it can effectively reduce the false negative rate for device users. Compared to speech activity detection that relies solely on air conduction sound, this processing method enables enhanced detection of device users based on bone conduction sound in complex acoustic environments. Therefore, it effectively reduces the occurrence of misidentifying non-device users as device users in complex acoustic environments. In summary, the speech activity detection method provided in this application can effectively improve the robustness of speech activity detection, thereby enhancing the voice interaction experience of the device in real-world usage environments.

[0128] Tenth Embodiment In the above embodiments, a voice activity detection method is provided. Correspondingly, this application also provides a voice activity detection device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0129] This application also provides a speech activity detection device, comprising: a sound data acquisition unit, a feature processing unit, and a classification unit. The sound data acquisition unit is used to acquire air-conduction sound data and bone-conduction sound data; the feature processing unit is used to obtain deep fusion features of air-conduction sound and bone-conduction sound based on the air-conduction sound data and bone-conduction sound data using the feature processing module included in the speech activity detection model; the classification unit is used to obtain speech activity detection results based on the deep fusion features using the classification module included in the speech activity detection model, wherein the speech activity detection results include: no one speaking, the device user speaking, or a non-device user speaking.

[0130] In one example, the feature processing module included in the speech activity detection model obtains deep fusion features of air-conducted sound and bone-conducted sound based on air-conducted sound data and bone-conducted sound data, including: obtaining depth features of air-conducted sound based on air-conducted sound data through an air-conducted feature extraction submodule included in the feature processing module; obtaining depth features of bone-conducted sound based on bone-conducted sound data through a bone-conducted feature extraction submodule included in the feature processing module; and obtaining deep fusion features of air-conducted sound and bone-conducted sound based on the depth features of air-conducted sound and bone-conducted sound through a feature fusion submodule included in the feature processing module.

[0131] In another example, the feature processing module included in the speech activity detection model obtains deep fusion features of air-conducted sound and bone-conducted sound based on air-conducted sound data and bone-conducted sound data. This includes obtaining deep fusion features of air-conducted sound and bone-conducted sound through multiple feature extraction layers included in the feature processing module. The input data of the first feature extraction layer includes air-conducted sound data and bone-conducted sound data, and the abstract features output by different levels of feature extraction layers are fusion features of different depth levels.

[0132] Eleventh Embodiment In the above embodiments, a voice activity detection method is provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0133] The electronic device of this embodiment includes: a memory and a processor; the memory is used to store a program that implements any of the above methods, and the device is powered on and runs the program of the above methods through the processor.

[0134] Memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0135] In specific implementations, the electronic device may further include one or more of the following components: a power supply component, an input / output (I / O) interface, and a communication component. The power supply component provides power to various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device. The I / O interface provides an interface between the processor 503 and peripheral interface modules, which may be a keyboard, click wheel, buttons, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and user devices (such as smartphones, tablets, etc.).

[0136] Twelfth Embodiment This application also provides a computer-readable storage medium. Since the embodiments of the computer-readable storage medium are substantially similar to the method embodiments, the description is relatively simple; relevant details can be found in the description of the method embodiments. The computer-readable storage medium embodiments described below are merely illustrative.

[0137] In this embodiment, a non-transitory computer-readable storage medium including instructions is provided, such as a memory including instructions, which can be executed by a processor of an electronic device to complete any of the methods provided in this disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0138] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0139] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

[0140] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0141] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0142] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0143] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A method for detecting speech activity, characterized in that, include: Collect air conduction sound data and bone conduction sound data; Based on the air-conducted sound data, obtain the air-conducted speech activity detection results; In addition, based on bone conduction sound data, bone conduction speech activity detection results are obtained; Based on the air-guided speech activity detection results and the bone-guided speech activity detection results, a joint speech activity detection result is obtained; The joint voice activity detection results include: no one speaking, the device user speaking, or a non-device user speaking.

2. The method according to claim 1, characterized in that, The step of obtaining air-conducted speech activity detection results based on air-conducted sound data includes: The speech activity detection model includes an air-conducted speech activity detection module, which obtains the depth features of air-conducted sound based on the air-conducted sound data; and obtains the air-conducted speech activity detection results based on the depth features of the air-conducted sound. The step of obtaining bone conduction speech activity detection results based on bone conduction sound data includes: The speech activity detection model includes a bone conduction speech activity detection module, which obtains the depth features of bone conduction sound data based on the bone conduction sound data; and obtains the bone conduction speech activity detection results based on the depth features of the bone conduction sound.

3. The method according to claim 2, characterized in that, The speech activity detection model includes an air-conducted speech activity detection module, which obtains the depth features of air-conducted sound based on air-conducted sound data, including: The air-conducted speech activity detection module includes an air-conducted sound depth feature extraction submodule, which obtains the depth features of air-conducted sound based on the air-conducted sound data. The speech activity detection model includes a bone conduction speech activity detection module, which obtains depth features of bone conduction sound based on bone conduction sound data, including: The bone conduction speech activity detection module includes a bone conduction sound depth feature extraction submodule, which obtains the depth features of bone conduction sound based on the bone conduction sound data. Both the air-conducting sound depth feature extraction submodule and the bone-conducting sound depth feature extraction submodule adopt a neural network structure; the neural network size of the air-conducting sound depth feature extraction submodule is smaller than that of the bone-conducting sound depth feature extraction submodule.

4. The method according to claim 2, characterized in that, The speech activity detection model includes an air-conducted speech activity detection module, which obtains air-conducted speech activity detection results based on the depth features of air-conducted sound. These results include: The air conduction speech activity detection module includes a speech activity probability acquisition submodule, which acquires the air conduction speech activity probability based on the depth features of the air conduction sound. The smooth decision submodule included in the air-conducted speech activity detection module obtains the air-conducted speech activity detection result based on the air-conducted speech activity probability and the air-conducted speech activity probability threshold. The speech activity detection model includes a bone conduction speech activity detection module, which obtains bone conduction speech activity detection results based on the depth features of bone conduction sound. These results include: The bone conduction speech activity detection module includes a speech activity probability acquisition submodule, which acquires the bone conduction speech activity probability based on the depth features of the bone conduction sound. The smooth decision submodule included in the bone-guided speech activity detection module obtains the bone-guided speech activity detection results based on the bone-guided speech activity probability and the bone-guided speech activity probability threshold. The method further includes: adjusting the air conduction speech activity probability threshold and / or the bone conduction speech activity probability threshold.

5. The method according to claim 4, characterized in that, The adjustment of the air conduction speech activity probability threshold and / or bone conduction speech activity probability threshold shall be performed in at least one of the following ways: If it is a device wake-up scenario, then reduce the bone conduction speech activity probability threshold and / or increase the air conduction speech activity probability threshold; In face-to-face translation scenarios, increase the bone conduction speech activity probability threshold and / or decrease the air conduction speech activity probability threshold.

6. The method according to any one of claims 2 to 5, characterized in that, The acquisition of air conduction sound data includes: acquiring multiple channels of air conduction sound data through an air conduction microphone array; The method further includes: Based on multi-channel air-conducting sound data, the user's voice is enhanced. The step of obtaining the depth features of air-conducted sound based on air-conducted sound data includes: Depth features of air conduction sound are obtained based on enhanced air conduction sound data from equipment users and air conduction sound data from non-equipment users.

7. The method according to any one of claims 2 to 5, characterized in that, The acquisition of air conduction sound data includes: acquiring multiple channels of air conduction sound data through an air conduction microphone array; The method further includes: Based on multi-channel air-conducted sound data, different degrees of voice enhancement are applied to the voice of equipment users and the voice of non-equipment users; The step of obtaining the depth features of air-conducted sound based on air-conducted sound data includes: Depth features of air conduction sound are obtained based on enhanced air conduction sound data from equipment users and non-equipment users.

8. The method according to claim 7, characterized in that, The process of enhancing the voice of device users and non-device users to different degrees based on multi-channel air-conducting sound data includes: Through the first beamforming, voice enhancement is performed in the direction of the device user based on multi-channel air-guided sound data; By using second beamforming, voice enhancement is performed in the direction of non-equipment users based on multi-channel air-guided sound data; The angle targeted by the first beamforming towards the device user is different from the angle targeted by the second beamforming towards the non-device user.

9. A method for constructing a speech activity detection model, characterized in that, include: Obtain the training dataset; The training data includes: air conduction sound data, bone conduction sound data, and labeled data of air conduction speech activity detection results and bone conduction speech activity detection results. The labeled data of air conduction speech activity detection results includes whether someone is speaking or not. The labeled data of bone conduction speech activity detection results includes whether someone is speaking or not. A speech activity detection model is learned from the training dataset; the model includes an air-conducting speech activity detection module and a bone-conducting speech activity detection module; the air-conducting speech activity detection module is used to obtain the depth features of air-conducting sound data and obtain the air-conducting speech activity detection result based on the depth features of air-conducting sound; the bone-conducting speech activity detection module is used to obtain the depth features of bone-conducting sound data and obtain the bone-conducting speech activity detection result based on the depth features of bone-conducting sound.

10. A face-to-face translation method, characterized in that, include: Collect air conduction sound data and bone conduction sound data; Based on the air-conducted sound data, obtain the air-conducted speech activity detection results; In addition, based on bone conduction sound data, bone conduction speech activity detection results are obtained; If the air conduction speech activity detection result indicates that someone is speaking, and the bone conduction speech activity detection result indicates that no one is speaking, then the joint speech activity detection result is determined to be that the speaker is not a device user. Based on air conduction sound data, translation data of the speech content spoken by non-device users is obtained.

11. A device wake-up method, characterized in that, include: Collect air conduction sound data and bone conduction sound data; Based on the air-conducted sound data, obtain the air-conducted speech activity detection results; In addition, based on bone conduction sound data, bone conduction speech activity detection results are obtained; If both the air conduction speech activity detection result and the bone conduction speech activity detection result indicate that someone is speaking, then the combined speech activity detection result is determined to be that the device user is speaking. Device wake-up identification is performed based on air conduction sound data.

12. A method for detecting speech activity, characterized in that, include: Collect air conduction sound data and bone conduction sound data; The feature processing module included in the speech activity detection model obtains deep fusion features of air conduction sound and bone conduction sound based on air conduction sound data and bone conduction sound data. The voice activity detection model includes a classification module, which obtains voice activity detection results based on the deep fusion features. The voice activity detection results include: no one speaking, device user speaking, or non-device user speaking.

13. A voice activity detection device, characterized in that, include: The sound data acquisition unit is used to acquire air conduction sound data and bone conduction sound data; Two detection units are used to obtain the air conduction speech activity detection results based on air conduction sound data; In addition, based on bone conduction sound data, bone conduction speech activity detection results are obtained; The joint decision-making unit is used to obtain the joint speech activity detection result based on the air-guided speech activity detection result and the bone-guided speech activity detection result; The joint voice activity detection results include: no one speaking, the device user speaking, or a non-device user speaking.

14. An electronic device, characterized in that, include: Air conduction pickup device; Bone conduction microphone; processor; as well as A memory for storing a program for implementing the method according to any one of claims 1 to 12, wherein the device is powered on and the program for running the method is executed by the processor.