Generative speech restoration using vibration sensor data
A GSR machine learning network processes voice vibration sensor data to restore high-quality speech data, addressing the complexity and cost issues of multi-microphone setups by enhancing bandwidth without acoustic microphones.
Patent Information
- Application Number
- PCT/US2024/059217
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2024-12-09
- Publication Date
- 2025-08-07
AI Technical Summary
Existing headset devices for voice communications utilize multiple acoustic microphones in combination with voice vibration sensors, increasing cost and complexity, and there is a need for systems and techniques to provide high-quality speech data without relying on acoustic microphones.
Implementing a General Speech Restoration (GSR) machine learning network, specifically a generative neural network, to generate a restoration mask based on voice vibration sensor data, allowing for the recovery of high-quality speech data from low-quality and band-limited inputs without using acoustic microphone data.
The GSR network effectively increases the bandwidth of voice vibration sensor speech output to match or exceed the frequency range of human voice, providing high-quality speech data without the need for additional acoustic microphones, thereby reducing device complexity and cost.
Smart Images

Figure US2024059217_07082025_PF_FP_ABST
Abstract
Description
Qualcomm Ref. No.2307828WO GENERATIVE SPEECH RESTORATION USING VIBRATION SENSOR DATA FIELD
[0001] The present disclosure generally relates to audio signal processing. For example, aspects of the present disclosure relate to general speech restoration (GSR) using vibration sensor data (e.g., bone conduction microphone data). BACKGROUND
[0002] In some examples, when a user speaks (e.g., generates a self-voice signal), the user’s voice may travel along two paths, including an acoustic path and a bone conduction path. Acoustic microphones can be used to pick up an acoustic path-based input audio signal using the acoustic path. The acoustic path-based input audio signal can include the user’s self-voice signal and may additionally include distortion patterns from external or background signals, noise, etc. A voice vibration sensor (e.g., bone conduction microphone (BCM), voice accelerometer (VA), etc.) can be used to pick up a bone conduction path-based input audio signal using the bone conduction path. The bone conduction path-based input audio signal can include the user’s self-voice signal at an improved signal-to-noise ratio (SNR). For example, the bone conduction path-based input audio signal may include a lesser and / or negligible contribution from external or background signals, noise, etc.
[0003] Voice vibration sensors are devices that can be used to sense or detect human speech (e.g., a user voice) based on sensing the bone-conducted vibrations caused by the vocal cords. Voice vibration sensors can be used to capture mechanical vibrations through a wearer’s skin, using the bone conduction path, and to convert the captured mechanical vibrations into electrical signals indicative of or including the wearer’s self-voice signal. In some examples, the terms “voice vibration sensor” and “BCM” may be used interchangeably. For instance, voice vibration sensors may also be referred to as BCMs, and vice versa; a voice vibration sensor can be used to implement a BCM, and vice versa. Voice vibration sensors are not designed to sense air-conducted sound, as a traditional acoustic microphone would. Instead, a voice vibration sensor can be designed to sense bone-conducted and / or soft tissue- conducted vibrations that are caused by, and propagate from, the user’s vocal cords. To sense these bone or soft tissue-conducted vibrations, a voice vibration sensor can be coupled (e.g.,Qualcomm Ref. No.2307828WO brought into physical contact, either directly or indirectly) with some portion of the user’s body. For instance, a voice vibration sensor can be placed directly on the skin, often on (or near) the head or neck. BRIEF SUMMARY
[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0005] Disclosed are systems, methods, apparatuses, and computer-readable media for processing audio data. According to at least one illustrative example, a method of processing audio data is provided, the method comprising: obtaining an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generating, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and generating, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
[0006] In another example, an apparatus for processing audio data is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: obtain an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generate, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in theQualcomm Ref. No.2307828WO voice vibration speech signal; and generate, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
[0007] In another example, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at least one processor to: obtain an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generate, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and generate, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
[0008] In another example, an apparatus for processing audio data is provided. The apparatus includes: means for obtaining an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; means for generating, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and means for generating, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
[0009] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifyingQualcomm Ref. No.2307828WO or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages, will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims.
[0010] While aspects are described in the present disclosure by illustration to some examples, those skilled in the art will understand that such aspects may be implemented in many different arrangements and scenarios. Techniques described herein may be implemented using different platform types, devices, systems, shapes, sizes, and / or packaging arrangements. For example, some aspects may be implemented via integrated chip examples or implementations, or other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, and / or artificial intelligence devices). Aspects may be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and / or system-level components. Devices incorporating described aspects and features may include additional components and features for implementation and practice of claimed and described aspects. For example, transmission and reception of wireless signals may include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and / or summers). It is intended that aspects described herein may be practiced in a wide variety of devices, components, systems, distributed arrangements, and / or end-user devices of varying size, shape, and constitution.
[0011] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.Qualcomm Ref. No.2307828WO
[0012] Other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art based on the accompanying drawings and detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Illustrative aspects of the present application are described in detail below with reference to the following figures:
[0014] FIG.1 is a diagram illustrating an example of an audio signaling using one or more voice vibration sensors, bone conduction microphones (BCMs), and / or voice accelerometers (VAs), in accordance with some examples;
[0015] FIG.2 is a diagram illustrating an example of a wearable device that can be used to perform audio signal processing using one or more voice vibration sensors to sense bone conducted voice or speech signals using one or more audio frequency bands within the voice vibration frequency range of human speech, in accordance with some examples;
[0016] FIG. 3 is a diagram of an example audio signal processing system including a wearable device with one or more voice vibration sensors for sensing bone conducted voice or speech signals, in accordance with some examples;
[0017] FIG. 4 is a diagram illustrating an example of a voice communications device including a voice vibration sensor and one or more outward-facing acoustic microphones, in accordance with some examples;
[0018] FIG. 5A is a diagram illustrating an example of an audio processing system for general speech restoration (GSR) using one or more generative neural networks and voice vibration sensor data, in accordance with some examples;
[0019] FIG. 5B is a diagram illustrating an example of general speech restoration using voice vibration sensor data and one or more outward-facing microphone signals, in accordance with some examples;
[0020] FIG.6 is a diagram of an example GSR machine learning architecture for generative speech restoration, in accordance with some examples;Qualcomm Ref. No.2307828WO
[0021] FIG.7 is a flow chart illustrating an example of a process for processing audio data, in accordance with some examples; and
[0022] FIG. 8 is a block diagram illustrating an example of a computing system, in accordance with some examples. DETAILED DESCRIPTION
[0023] Certain aspects and aspects of this disclosure are provided below. Some of these aspects and aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.
[0024] The ensuing description provides exemplary aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary aspects will provide those skilled in the art with an enabling description for implementing an exemplary aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the scope of the application as set forth in the appended claims.
[0025] Voice vibration sensors are devices that can be used to sense or detect human speech (e.g., voice) based on the detection of vibration in solid materials. For instance, voice vibration sensors can be configured as a contact device or sensor (e.g., placed in contact with a user’s jaw, skin, etc.) that may be used to sense the bone-conducted vibrations caused by the user’s vocal cords. Voice vibration sensors may include and / or be interchangeably referred to as bone conduction microphones (BCMs) and / or bone conduction sensors. Whereas acoustic microphones are designed to generate an audio signal based on sensing air- conducted sound waves, voice vibration sensors are designed to sense bone-conducted and / or soft tissue-conducted vibrations that are caused by, and propagate from, the user’s vocal cords. To sense these bone or soft tissue-conducted vibrations, a voice vibration sensor can be coupled (e.g., brought into physical contact, either directly or indirectly) with some portionQualcomm Ref. No.2307828WO of the user’s body. For instance, a voice vibration sensor can be placed directly on the skin, often on (or near) the user’s head or neck.
[0026] In some examples, one or more voice vibration sensors can be included in various wearable devices and / or other audio and / or electronic devices. Voice vibration sensors can be used for various purposes, which can include one or more of self-speech pick-up, imposter rejection, noise suppression, voice activity detection (VAD), vibration pick-up for noise cancellation, etc., among various others. For instance, one or more voice vibration sensors can be included in wearable devices such as a pair of in-ear true wireless stereo (TWS) earbuds, AR / VR headsets, smart glasses, etc., and can be used to obtain bone conducted audio signals that capture the user’s voice (e.g., the wearer’s voice) against unwanted external sound, etc. In another example, one or more voice vibration sensors can be used to provide covert communications, based on the voice vibration sensors having a lower threshold of audibility or detectability of vocal sounds produced by a user. Unlike acoustic microphones, voice vibration sensors are not designed to sense air-conducted sound. In some examples, voice vibration sensors can be used for voice enhancement processing for mobile communication or various other use cases, for instance based on the voice vibration sensor having greater sensitivity to bone-conducted speech rather than air-conducted background noise and other unwanted external sound(s). One or more voice vibration sensors included in earbuds or headphones can be used to capture the wearer’s voice as vibration, rather than any air-conducted external sound (e.g., the one or more voice vibration sensors can be used to capture the wearer’s voice vibration while rejecting sounds that are air-conducted). In another example, one or more voice vibration sensors can be used to perform voice activity detection (VAD), such as for detecting a wake-up word to activate a device that includes or is associated with the one or more voice vibration sensors (e.g., where improved accuracy of the detection of the user utterance of the wake-up word can prevent false activations).
[0027] In some examples, TWS earbuds or other headset devices for voice communications may include multiple sensors and / or microphones. For example, a multi-microphone headset device for voice communications may include one or more primary outward-facing voice microphones (e.g., outward-facing acoustic microphones), one or more feed-forward acoustic microphones (e.g., used for active noise cancellation (ANC), and one or more voice vibrationQualcomm Ref. No.2307828WO sensors for bone conduction pickup. In some cases, voice or speech pickup can be implemented using various combinations of the respective signals from the one or more acoustic microphones and the one or more voice vibration sensors. For instance, voice vibration sensors may have better SNR in lower frequency ranges and acoustic microphones may have better SNR in higher frequency ranges, and a multi-microphone may implement voice or speech pickup based on low-pass filtering a voice vibration sensor output and high- pass filtering an acoustic microphone output.
[0028] Voice vibration sensors may be associated with frequency-dependent sensitivity characteristics relative to the bone-conducted voice vibration frequency range (e.g., the range of voice vibration frequencies that can be bone-conducted above a detection threshold). In some examples, the bone-conducted voice vibration frequency range can include frequencies from 100 Hz – 1 kHz. In some cases, the frequency range of the bone-conducted voice vibration (e.g., also referred to as the “voice band”) can include frequencies below 100 Hz and / or frequencies above 1 kHz. For instance, the frequency range of bone-conducted voice vibration can be based at least in part on the respective voice vibration sensor implementation used to sense or detect the bone-conducted voice vibration frequencies. In some examples, the bone-conducted voice vibration range can be based on a location on the head where the vibrations are sensed (e.g., a location of the voice vibration sensor on the head or body). The bone-conducted voice vibration range can additionally be based on the respective noise floor of the voice vibration sensor implementation used to sense or detect the bone-conducted voice vibration frequencies. In some cases, the bone-conducted voice vibration range can correspond to the voice vibration sensor’s noise floor relative to the amplitude of the vibrations (e.g., the amplitude of the bone-conducted voice vibrations). For instance, in some examples of voice vibration sensors located at the ear, vibration energy above 1 kHz drops into the voice vibration sensor’s noise floor and may be undetectable.
[0029] As noted above, many current implementations of headset devices for voice communications utilize various combinations of acoustic microphones for sensing relatively higher frequencies of human speech and voice vibration sensors for sensing relatively lower frequencies of human speech. Implementing multiple acoustic microphones in combination with one or more voice vibration sensors can increase the cost and complexity of the headsetQualcomm Ref. No.2307828WO devices for voice communications. There is a need for systems and techniques that can be used to provide high-quality speech data using a voice vibration sensor. There is a further need for systems and techniques that can be used to provide high-quality speech data without using multiple acoustic microphones in combination with the voice vibration sensor.
[0030] Systems, apparatuses, processes (also referred to as methods), and computer- readable media (collectively referred to as “systems and techniques”) are described herein that can be used to perform general speech restoration (GSR) using vibration sensor data. For instance, the systems and techniques can be used to perform GSR to recover or restore high- quality speech data from relatively low-quality and / or band-limited voice vibration sensor data. In some aspects, the systems and techniques can perform GSR using voice vibration sensor data without utilizing acoustic microphone data or signals to augment the voice vibration sensor data.
[0031] According to some aspects, one or more GSR machine learning networks can be used to generate high-quality speech data from a relatively low-quality speech data input. In some cases, the one or more GSR machine learning networks can be implemented as one or more generative neural networks. For example, generative neural network can be configured to receive an input spectrogram corresponding to the relatively low-quality speech data input. The generative neural network (and / or various other GSR machine learning network implementations) can generate a restoration mask based on the spectrogram of the low- quality speech data input. The restoration mask can be used to generate (e.g., recover and / or restore etc.) high-quality speech data from the relatively low-quality speech data input.
[0032] In one illustrative example, a GSR machine learning network and / or GSR generative neural network can be trained and implemented to generate a restoration mask based on input speech data obtained using a voice vibration sensor. For instance, the GSR machine learning network and / or GSR generative neural network can be trained and implemented to generate restoration masks for input speech data obtained using only a voice vibration sensor (e.g., the GSR machine learning network and / or GSR generative neural network can generate a restoration mask without utilizing acoustic microphone audio data or audio signals as input).Qualcomm Ref. No.2307828WO
[0033] In some aspects, the GSR machine learning network (e.g., GSR generative neural network) can be configured to generate a restored speech output corresponding to a relatively low-quality speech input obtained using a voice vibration sensor speech. In one illustrative example, the restored speech output can have an increased bandwidth relative to the voice vibration sensor speech input. In some aspects, the input bandwidth to the GSR network can correspond to the band-limited bone-conducted voice vibration range of the voice vibration sensor used to obtain the audio input signal. The GSR network can generate the restored speech output to have an increased bandwidth that includes the full bone-conducted voice vibration range and / or that is the same as or similar to the bandwidth (e.g., frequency range) of an acoustic microphone. In some cases, the GSR network can generate the restored speech output to have an increased bandwidth that is the same as or similar to the full bandwidth (e.g., full frequency range) of human voice or speech.
[0034] In some examples, the GSR machine learning network (e.g., GSR generative neural network) can be configured to utilize an input comprising a combination of audio data from a voice vibration sensor and audio data from an outward facing acoustic microphone, where the voice vibration sensor and the outward facing acoustic microphone are included in the same audio device (also referred to as a voice communications audio device). For instance, the voice vibration sensor and the outward facing acoustic microphone can be included in the same headset device for voice communications, the same pair of TWS earbuds, etc. In some aspects, the voice vibration sensor signal and the outward facing acoustic microphone signal can be synchronized prior to being fed to the GSR machine learning network as input. For instance, based on the relative positioning (e.g., relative separation distances, etc.) between the voice vibration sensor and the outward facing acoustic microphone on the same audio device, a respective delay can be applied to the voice vibration sensor signal or to the outward facing acoustic microphone signal, where the delayed signal is thereby synchronized with the other, non-delayed signal. In some cases, a first delay may be applied to the voice vibration sensor signal and a second delay may be applied to the outward facing acoustic microphone signal, where the first delay value is different from the second delay value.
[0035] Further aspects of the systems and techniques will be described with reference to the figures.Qualcomm Ref. No.2307828WO
[0036] FIG.1 is a diagram illustrating an example of an audio signaling scenario 100 using one or more voice vibration sensors (e.g., bone conduction sensors, bone conduction microphones (BCMs), and / or voice accelerometers (VAs), etc.) in accordance with some examples. For instance, the audio signaling scenario 100 may be associated with a user 105 using a wearable device 115 to experience a listen-through feature (e.g., among various other features and use cases that can be associated with and / or implemented using one or more voice vibration sensors).
[0037] For example, a user 105 may use a wearable device 115 (e.g., a wireless communication device, wireless headset, earbuds, in-ear true wireless stereo (TWS) earbuds, speaker, hearing assistance device, or the like), which may be worn by the user 105 in a hands-free manner. In some cases, the wearable device 115 may also be referred to as a hearing device. In some examples, the user 105 may continuously wear the wearable device 115, whether the wearable device 115 is currently in use (e.g., inputting an audio signal, outputting an audio signal, or both at one or more microphones 120) or not. In some examples, the wearable device 115 may include multiple microphones 120. For instance, the wearable device 115 may include one or more outer microphones 120, such as outer microphone 120a and outer microphone 120b. Wearable device 115 may also include one or more inner microphones 120, such as inner microphone 120c. The wearable device 115 may use the microphones 120 for noise detection, audio signal output, active noise cancellation, and the like. A wearable device (e.g., such as the wearable device 115) can include a greater or lesser number of microphones.
[0038] When the user 105 speaks, the user 105 may generate a unique audio signal (e.g., self-voice signal). For example, the user 105 may generate a self-voice signal that may travel along an acoustic path 125 (e.g., from the mouth of user 105 to the microphones 120 of the headset). The user 105 may also generate a self-voice signal that may follow a sound conduction path 130 created by vibrations via bone conduction between the vocal cords or mouth of the user 105 and the microphones 120 of the wearable device 115. In some examples, the wearable device 115 may perform self-voice activity detection (SVAD) based on the self-voice qualities. For instance, the wearable device 115 may identify inter channel phase and intensity differences (e.g., interaction between the outer microphones 120 and theQualcomm Ref. No.2307828WO inner microphones 120 of the wearable device 115). In some cases, the wearable device 115 may use the detected differences as qualifying features to contrast self-speech signals and external signals. For example, if one or more differences between channel phase and intensity between inner microphone 120c and outer microphone 120a are detected, or if one or more differences between channel phase and intensity between inner microphone 120c and outer microphone 120a satisfy a threshold value, then the wearable device 115 may determine that a self-voice signal is present in an input audio signal.
[0039] In some examples, the wearable device 115 may provide a listen-through feature for operating in a transparent mode. A listen-through feature may allow the user 105 to hear an output audio signal from the wearable device 115 as if the wearable device 115 were not present. The listen-through feature may allow the user 105 to wear the wearable device 115 in a hands-free manner regardless of the current use-case of the wearable device 115 (e.g., regardless of whether the wearable device 115 is outputting an audio signal, inputting an audio signal, or both using one or more microphones 120). For example, an audio source 110 (e.g., a person, audio from the surrounding environment, or the like) may generate an external audio signal 135. For example, a person may speak to the user 105, creating external audio signal 135. Without a listen-through feature, the external audio signal 135 may be blocked, muffled, or otherwise distorted by the wearable device 115. A listen-through feature may utilize outer microphone 120a, outer microphone 120b, inner microphone 120c, or a combination to receive an input audio signal (e.g., external audio signal 135), process the input audio signal, and output an audio signal (e.g., via inner microphone 120c) that sounds natural to the user 105 (e.g., sounds as if the user 105 were not wearing a device).
[0040] A self-voice audio signal following acoustic path 125 and the external audio signal 135 may have different distortion patterns. For instance, the external audio signal 135, self- voice audio signal following acoustic path 125, or both may have a first distortion pattern. But self-voice following sound conduction path 130, self-voice following acoustic path 125, or both may have a second distortion pattern. The microphones 120 of the wearable device 115 may detect the self-voice audio signal and the external audio signal 135 similarly. Thus, without different treatments for the different signal types, a user 105 may not experience a natural sounding input audio signal. That is, wearable device 115 may detect an input audioQualcomm Ref. No.2307828WO signal including a combination of external audio signal 135, self-voice via acoustic path 125, or self-voice via sound conduction path 130. Wearable device 115 may detect the input audio signal using the microphones 120.
[0041] In some examples, one or more (or all) of the microphones 120 can be implemented as voice vibration sensors and / or bone conduction microphones (BCMs). A voice vibration sensor can be configured to detect the bone conducted voice of a user (e.g., the bone conducted self-voice signal). In some cases, the wearable device 115 can include one or more voice vibration sensors 140. The voice vibration sensor 140 can be the same as or similar to any of the microphones 120 that may be implemented as BCMs, VAs, and / or various other types of voice vibration sensor(s). In some examples, the one or more voice vibration sensors 140 may be different from one or more of the microphones 120 and / or may be different from one or more BCMs or VAs used to implement the microphones 120. In some examples, the one or more voice vibration sensors 140 can be BCMs and / or VAs.
[0042] In some cases, a user 105 may experience bone conduction when speaking using wearable device 115. For example, bone conduction may be the conduction of sound to the inner ear through the bones of the skull, which may allow the user 105 to perceive audio (e.g., speech or self-voice, etc.) using vibrations in the bone. In some examples, bone may convey lower-frequency sounds better than higher-frequency sound. The voice vibration sensor 140 may include a transducer that outputs a signal based on the vibrations of the bone due to audio. Additionally or alternatively, the voice vibration sensor 140 may include any device (e.g., a sensor, or the like) that detects a vibration and outputs an electronic signal.
[0043] In some examples, the wearable device 115 may receive an input audio signal from outer microphone 120a, outer microphone 120b, or both (e.g., an external audio signal 135, the self-voice of the user 105, or both) and an input audio signal from an inner microphone 120c. The wearable device 115 may output an audio signal (e.g., the bone conduction signal, also referred to as the “voice vibration speech signal,” “voice vibration signal,” or “vibration signal”) to a speaker or other audio device (e.g., including various speakers or audio playback devices the user 105 can hear, etc.).
[0044] FIG. 2 is a diagram illustrating an example of a wearable device 205 that can be used to perform audio signal processing using one or more voice vibration sensors to senseQualcomm Ref. No.2307828WO bone conducted voice or speech signals using one or more audio frequency bands within the voice vibration frequency range of approximately 100 Hertz (Hz) to 1 kilohertz (kHz), in accordance with some examples. In some cases, the wearable device 205 may be an example of aspects of a wearable device 115 of FIG. 1. The wearable device 205 may include a receiver 210, a signal processing manager 215, and a speaker 220. The wearable device 205 may also include a processor. Each of these components may be in communication with one another (e.g., via one or more buses).
[0045] The receiver 210 may receive audio signals from a surrounding area (e.g., via an array of microphones, including one or more voice vibration sensors for sensing bone conducted voice or speech signals). Detected audio signals may be passed on to other components of the wearable device 205. The receiver 210 may utilize a single antenna or a set of antennas to communicate wirelessly with other devices and / or may utilize one or more wired connections to communicate with other devices.
[0046] The signal processing manager 215 may receive, at the wearable device 205 including at least one voice vibration sensor for bone conducted audio sensing, a corresponding one or more bone conducted audio signals (e.g., voice vibration signals). The voice vibration audio signals can correspond to the voice or speech of a user of the wearable device 205. In some cases, the voice vibration audio signals may be received in one or more frequency bands and / or using one or more frequency band groups or subsets of the voice vibration frequency range between 100 Hz – 1 kHz.
[0047] The actions performed by the signal processing manager 215 as described herein may be implemented to realize one or more potential advantages. One implementation may enable a wearable device (e.g., wearable device 115 of FIG.1, wearable device 205 of FIG. 2, etc.) to use a signal output of a voice vibration sensor or other bone conduction sensor to account for self-voice in an audio signal. The voice vibration sensor can be used to obtain a bone conducted audio signal (e.g., a bone conducted self-voice signal) of the user, which can be used for various downstream audio processing and / or audio output tasks, etc. For instance, the bone conducted audio signal can be used to implement filtering of one or more acoustic audio signals (e.g., non-bone conducted audio signals obtained from acoustic microphones), to provide a transparent mode to the user, to allow for a natural sounding self-voice as anQualcomm Ref. No.2307828WO output of the wearable device, to perform various other voice enhancement audio signal processing operations, and / or to perform various other noise reduction operations, etc., among various others. Using one or more voice vibration sensors (e.g., BCMs, VAs, etc.) to generate or sense bone conducted self-voice signals of the user, a processor of a wearable device (e.g., a processor controlling the receiver 210, the signal processing manager 215, the speaker 220, or a combination thereof) may improve user experience.
[0048] The signal processing manager 215, or its sub-components, may be implemented in hardware, code (e.g., software or firmware) executed by a processor, or any combination thereof. If implemented in code executed by a processor, the functions of the signal processing manager 215, or its sub-components may be executed by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate-array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described in the present disclosure.
[0049] The signal processing manager 215, or its sub-components, may be physically located at various positions, including being distributed such that portions of functions are implemented at different physical locations by one or more physical components. In some examples, the signal processing manager 215, or its sub-components, may be a separate and distinct component in accordance with various aspects of the present disclosure. In some examples, signal processing manager 215, or its sub-components, may be combined with one or more other hardware components, including but not limited to an input / output (I / O) component, a transceiver, a network server, another computing device, one or more other components described in the present disclosure, or a combination thereof in accordance with various aspects of the present disclosure.
[0050] The speaker 220 may provide output signals generated by other components of the wearable device 205. In some examples, the speaker 220 may be collocated with one or more microphones (e.g., voice vibration sensors, BCMs, VAs, and / or acoustic microphones) of wearable device 205.
[0051] FIG.3 is a diagram of an example audio signal processing system 300 including a wearable device 305 with one or more voice vibration sensors for sensing bone conductedQualcomm Ref. No.2307828WO voice or speech signals, in accordance with some examples. For instance, the example audio signal processing system 300 can be used to perform audio signal processing using one or more voice vibration sensors to sense bone conducted voice or speech signals using one or more audio frequency bands within the voice vibration frequency range of approximately 100 Hertz (Hz) to 1 kilohertz (kHz), in accordance with some examples.
[0052] The wearable device 305 may be an example of or include the components of wearable device 115 of FIG.1, wearable device 205 of FIG.2, etc. The wearable device 305 may include components for bi-directional voice and data communications including components for transmitting and receiving communications, including a signal processing manager 310, an input / output (I / O) controller 315, a transceiver 320, memory 330, and a processor 340. These components may be in electronic communication via one or more buses (e.g., bus 345).
[0053] The signal processing manager 310 may receive, at the wearable device including at least one voice vibration sensor 360 (e.g., a voice vibration sensor, a BCM, or other bone conduction sensor for bone conducted audio sensing), a corresponding one or more bone conducted audio signals. The bone conducted audio signals can correspond to the voice or speech of a user of the wearable device 205. In some cases, the bone conducted audio signals may be received in one or more frequency bands and / or using one or more frequency band groups or subsets of the voice vibration frequency range between 100 Hz – 1 kHz. In some cases, the wearable device 305 can additionally include one or more microphones 350, which may be provided as acoustic (e.g., non-bone conduction) microphones. In some examples, the signal processing manager 310 can receive acoustic audio signals from the one or more acoustic microphones 350 and can receive one or more bone conducted audio signals from the one or more voice vibration sensors 360.
[0054] The I / O controller 315 may manage input and output signals for the wearable device 305. The I / O controller 315 may also manage peripherals not integrated into the wearable device 305. In some cases, the I / O controller 315 may represent a physical connection or port to an external peripheral. In some cases, the I / O controller 315 may utilize an operating system such as iOS®, ANDROID®, MS-DOS®, MS-WINDOWS®, OS / 2®, UNIX®, LINUX®, or other known operating system(s). In some examples, the I / O controller 315 mayQualcomm Ref. No.2307828WO represent or interact with a modem, a keyboard, a mouse, a touchscreen, or a similar device. In some cases, the I / O controller 315 may be implemented as part of a processor. In some cases, a user may interact with the wearable device 305 via the I / O controller 315 or via hardware components controlled by the I / O controller 315.
[0055] The transceiver 320 may communicate bi-directionally, via one or more antennas, wired, or wireless links. For example, the transceiver 320 may represent a wireless transceiver and may communicate bi-directionally with another wireless transceiver. The transceiver 320 may also include a modem to modulate the packets and provide the modulated packets to the antennas for transmission, and to demodulate packets received from the antennas. In some examples, listen-through features implemented using the one or more voice vibration sensors 360 and corresponding bone conducted audio signals (e.g., bone conducted self-voice signals) described above may allow a user to experience natural sounding interactions with an environment while performing wireless communications or receiving data via transceiver 320.
[0056] The speaker 325 may provide an output audio signal to a user (e.g., with or without listen-through features and / or with or without combining the bone conducted audio signal(s) from the one or more voice vibration sensors 360 with the acoustic audio signal(s) from the one or more acoustic microphones 350 if present).
[0057] The memory 330 may include random-access memory (RAM) and read-only memory (ROM). The memory 330 may store computer-readable, computer-executable code 335 including instructions that, when executed, cause the processor to perform various functions described herein. In some cases, the memory 330 may contain, among other things, a basic I / O system (BIOS) which may control basic hardware or software operation such as the interaction with peripheral components or devices.
[0058] The processor 340 may include an intelligent hardware device or component (e.g., a general-purpose processor, a digital signal processor (DSP), a central processing unit (CPU), a microcontroller, an ASIC, an FPGA, a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, a neural processing unit (NPU), a neural signal processor (NSP), any combination thereof, and / or other hardware device or component). In some cases, the processor 340 may be configured to operate aQualcomm Ref. No.2307828WO memory array using a memory controller. In other cases, a memory controller may be integrated into the processor 340. The processor 340 may be configured to execute computer- readable instructions stored in a memory (e.g., the memory 330) to cause the wearable device 305 to perform various functions (e.g., functions or tasks supporting speech restoration using a voice vibration sensor and / or a bone conduction sensor).
[0059] The code 335 may include instructions to implement aspects of the present disclosure, including instructions to support signal processing. In some cases, aspects of the signal processing manager 310, the I / O controller 315, and / or the transceiver 320 may be implemented by portions of the code 335 executed by the processor 340 or another device. The code 335 may be stored in a non-transitory computer-readable medium such as system memory or other type of memory. In some cases, the code 335 may not be directly executable by the processor 340 but may cause a computer (e.g., when compiled and executed) to perform functions described herein.
[0060] FIG.4 is a diagram illustrating an example of a voice communications device 400 including a voice vibration sensor and one or more outward-facing acoustic microphones, in accordance with some examples. For example, the voice communications device 400 can be a headset for voice communications, a TWS earbud, etc. In some aspects, the voice communications device 400 can be an example of or include the components of the wearable device 115 of FIG.1, the wearable device 205 of FIG.2, and / or the wearable device 305 of FIG.3, etc.
[0061] For instance, the voice communications device 400 can include a speaker 410 (e.g., among various other audio output devices, transducers, components, etc.) configured to output an audio signal to a user of the voice communications device 400. In some examples, the speaker 410 can be the same as or similar to one or more of the speaker 220 of FIG. 2, the speaker 325 of FIG.3, etc.
[0062] The voice communications device 400 can include a first voice microphone 422 and a second voice microphone 424. In some examples, the first voice microphone 422 and the second voice microphone 424 may both be outward-facing microphones (e.g., outward- facing relative to a housing 402 of the voice communications device 400). The first voice microphone 422 and the second voice microphone 424 may be examples of acousticQualcomm Ref. No.2307828WO microphones, and may be the same as or similar to one another. In some aspects, the first voice microphone 422 and the second voice microphone 424 can be outward-facing acoustic voice microphones provided on or within the housing 402 of the voice communications device 400. The first voice microphone 422 and / or the second voice microphone 424 may be the same as or similar to one or more of the microphones 120 of FIG. 1 (e.g., the outer microphone 120a, outer microphone 120b, inner microphone 120c, etc.), and / or the microphones 350 of FIG.3.
[0063] In some examples, the first microphone 422 can be an outward-facing, acoustic microphone configured for voice pickup for the voice communications device 400, and the second microphone 424 can be an outward-facing feedforward microphone. For instance, the first microphone 422 can be utilized as a primary outward-facing voice microphone, and the second microphone 424 can be a feedforward acoustic microphone used for and / or associated with active noise cancelling (ANC) implemented by the voice communications device 400.
[0064] The voice communications device 400 can include one or more voice vibration sensors 440. In some aspects, the voice vibration sensor 440 can be an internally mounted voice vibration sensor (e.g., the voice vibration sensor 440 can be provided internal to the voice communications device 400, within an interior or enclosed volume of the housing 402, etc.). The voice vibration sensor 440 can be the same as or similar to one or more of the voice vibration sensor 140 of FIG.1, and / or the voice vibration sensor 360 of FIG.3, etc. In some aspects, the voice vibration sensor 440 can be a BCM, a bone conduction sensor, a voice accelerometer (VA), etc.
[0065] In some aspects, the noise sub-space between voice vibration sensors and acoustic microphones may be uncorrelated (e.g., environmental noise in the acoustic domain does not affect the voice vibration sensor). For example, the noise sub-space between the voice vibration sensor 440 and the acoustic voice microphones 422, 424 may be uncorrelated for voice communications or other speech signal pickup performed by the example voice communications device 400 of FIG.4.
[0066] Based on the uncorrelated noise sub-space between voice vibration sensors and acoustic microphones, various types of noise and / or signal degradation may be present in an acoustic audio signal obtained using the first microphone 422 or the second microphone 424,Qualcomm Ref. No.2307828WO but may be absent from (or present in a reduced amount) the bone-conducted vibration signal obtained using the voice vibration sensor 440 or BCM. For example, various types of noise or distortion that may be present in an acoustic microphone signal but absent or reduced in prominence in a voice vibration sensor signal can include one or more of ambient or background noise, reverb, distortion, clipping, wind-noise, lowered resolution, etc., among various others.
[0067] However, as noted previously, voice vibration sensors (e.g., BCMs, etc.) can be associated with a relatively limited spectrum or bandwidth within the frequency range corresponding to human voice or speech. For example, voice vibration sensors may have a restricted frequency response, where the voice vibration sensor captures frequencies below 800-1,000 Hz effectively, but may experience attenuation of frequencies above 800-1,000 Hz. For instance, bone conducted voice signals (e.g., voice vibration signals) may typically be confined to the relatively low frequencies of human speech (e.g., corresponding to frequencies within the bone conducted voice vibration range between approximately 100 Hz and 1 kHz). The bone conducted voice vibration range is narrower than the frequency range of human speech, which may generally fall between 100 Hz and 8 kHz.
[0068] Some voice vibration sensor implementations may have a frequency response with a single resonance peak that is located outside of the 100 Hz – 1 kHz voice vibration band (e.g., a resonance peak greater than 1 kHz). The resonance peak of the voice vibration sensor frequency response is indicative of the resonance frequency where the voice vibration sensor exhibits the highest sensitivity. For instance, a voice vibration sensor with a resonance peak at 4 kHz will have higher sensitivity at frequencies near the 4 kHz resonance frequency and lower sensitivity at frequencies away from the 4 kHz resonance frequency. Some voice vibration sensor implementations are highly sensitive in a fixed narrow band at higher frequencies (e.g., outside of the 100 Hz – 1 kHz voice vibration band, such as in a fixed narrow band at a 4 kHz resonance frequency, etc.). In some examples, a voice vibration sensor may be implemented with a resonance peak near 4 kHz and with a high Q factor. A resonance peak in the higher frequencies can be implemented with a goal of boosting the high frequency of the measured voice signal towards improved clarity. In at least some examples, the high magnitude of the resonance peak can cause the voice vibration sensor toQualcomm Ref. No.2307828WO become sensitive to the air-conducted sound within its narrow frequency range. Additionally, the frequency of the resonance peak can be much higher than the voice vibration band frequencies, and the signal captured at the resonance peak does not represent the true bone- conducted voice, but additionally includes an air-conducted component, along with any unwanted external noise that may be present.
[0069] Based on the bone conducted voice vibration range representing a subset of the wider frequency range of human speech, a voice vibration sensor implementation may be unable to natively capture (e.g., capture without audio processing or audio post-processing operations, augmentation, etc.) the whole frequency range of the wearer’s voice as it would have been captured by an acoustic microphone. For instance, the voice captured using a voice vibration sensor may sound muffled (although free from air-conducted external sound) and / or may lack clarity. To compensate for the reduced voice clarity that may be associated with voice vibration sensor signals (e.g., voice vibration sensor voice or speech audio data), there is a need for implementations that can be configured to extend the bandwidth of the speech data obtained from and / or generated based on the voice vibration sensor data obtained from a voice vibration sensor.
[0070] In one illustrative example, and as noted previously, the systems and techniques described herein can be configured to perform general speech restoration (GSR) to enhance voice vibration sensor audio data. For instance, the systems and techniques can use one or more machine learning GSR networks to generate a reconstructed or high quality speech signal that includes a wider frequency range or bandwidth than the voice vibration sensor audio data. The reconstructed high-quality speech signal generated by the one or more machine learning GSR networks can additionally remove one or more distortions or degradations that may be present in the input voice vibration sensor audio data.
[0071] GSR and other speech restoration techniques can be implemented and used to restore degraded speech signals (e.g., the input speech signal) to high-quality speech signals (e.g., the output speech signal). For example, GSR and other speech restoration techniques can be used to reconstruct and / or enhance speech signals that are degraded, incomplete, and / or noisy, etc. The restored, high-quality speech signals can have improved intelligibilityQualcomm Ref. No.2307828WO and perceived quality (e.g., by a listener) relative to the relatively low-quality input speech signal provided to the GSR or speech restoration system.
[0072] In some aspects, one or more machine learning networks can be used to perform GSR and / or various other speech recognition tasks, and are referred to herein as “GSR machine learning networks” or “GSR networks.” In some examples, GSR networks can be implemented based on one or more generative neural networks that are configured to restore or enhance the low-quality input speech signal based on generating one or more missing portions of audio that that were not captured in the original input and / or that experienced degradation or distortion as represented in the capture of the original input, etc.
[0073] For example, generative neural network-based techniques for GSR machine learning can be based on generative adversarial networks (GANs), where a generator neural network and a discriminator neural network are configured to compete against one another. The generator neural network can be used to generate speech signals that are analyzed by the discriminator neural network, which learns to distinguish or differentiate between real (e.g., non-restored) and generated (e.g., restored) speech signals. The adversarial process of the interactions between the generator and discriminator networks can enhance the quality of the generated or restored speech signals. In another example, GSR machine learning can be based on one or more autoencoders and / or variational autoencoders (VAEs). These networks can be configured to model the distribution of speech data, and may be used to generate high- quality speech signals based on learning latent space representations of clean speech. In some examples, GSR machine learning may be based on WaveNet or related architectures. Originally designed for speech synthesis tasks, WaveNet and related machine learning architectures can be configured to use dilated convolutions to model temporal relationships in audio signals, which may be adapted for use in performing GSR and other speech restoration tasks. General neural networks and / or other GSR machine learning networks can be used to restore high-quality speech from a capture that has been degraded by several simultaneous distortions (e.g., such as noise, reverb, distortion, clipping, wind noise, low- resolution, etc.).
[0074] In some cases, GSR techniques and / or GSR machine learning networks can be implemented based on the biological mechanisms of human hearing when restoring distortedQualcomm Ref. No.2307828WOspeech. For example, given an input speech signal ^^ ∈ ℝ^, where ^^ is the number of samples,a distortion process acting on the input speech signal can be modeled as ^^^∙^, such that: ^^ ൌ ^^^^^^ Eq. (1)
[0075] Here, ^^ ∈ ℝ^ represents the degraded speech signal. The distortion processmodeled as ^^^∙^ can be a cumulative or multi-distortion process. For example, ^^^∙^ can capture the effects of one or multiple types of simultaneous distortions that may be present within or otherwise acting upon the input speech signal s (e.g., noise, reverb, distortion, clipping, wind noise, low-resolution, etc., as noted previously above).
[0076] The GSR speech restoration task can be to restore a high-quality speech signal ^^^ from the degraded input speech signal ^^: ^^^ ൌ ^^^^^^ Eq. (2)
[0077] In Eq. (2), the term f^∙^represents a restoration function than can be viewed as the reverse process of ^^^∙^. For example, one or more restoration functions can be learned and implemented by the one or more GSR machine learning networks to perform the GSR or other speech restoration task, as will be described below with respect to FIGS.5A-6.
[0078] In one illustrative example, the systems and techniques described herein can utilize a two-stage GSR system. The two-stage GSR system can be implemented using a single GSR machine learning network configured to perform the first stage and the second stage, and / or can be implemented using two or more GSR machine learning networks (e.g., one or more GSR machine learning networks configured to perform the first stage and one or more GSR machine learning networks configured to perform the second stage).
[0079] For instance, in a first GSR stage, the distorted speech signal ^^ can be mapped into a representation ^^ (e.g., Mel spectrogram): ^^ ൌ ^^^^^^ Eq. (3)
[0080] In one illustrative example, the intermediate representation ^^ can be a Mel spectrogram corresponding to (e.g., generated from) the distorted speech signal ^^. A MelQualcomm Ref. No.2307828WO spectrogram is a visual representation of the spectrum of frequencies in a sound signal as they vary with time, adapted to the Mel scale (e.g., where the Mel scale corresponds to the non-linear perceptual characteristics of human hearing, for improved alignment with human perception of pitch changes, etc.). For instance, frequencies in a Mel spectrogram can be converted to the Mel scale, where higher frequencies may be scaled down and lower frequencies may be scaled up to better match the non-linear perceptual characteristics of human hearing. Like a standard spectrogram, a Mel spectrogram can represent the intensity or power of different frequencies over time, as a two-dimensional (2D) representation with time provided on the horizontal axis and Mel frequency provided on the vertical axis. Color, brightness, or other values can be used to indicate the amplitude or power of each Mel frequency at each point in time represented within the Mel spectrogram. In some examples, a Mel spectrogram can be generated using a Fast Fourier Transform (FFT) applied to a time- domain audio signal, where the FFT converts the time-domain audio into the frequency- domain and the frequency-domain audio is subsequently mapped onto the Mel scale. For instance, a Mel spectrogram or other intermediate representation ^^ can be generated based on the FFT of the time-domain distorted speech signal ^^ and a subsequent mapping onto the Mel scale. In some aspects, the Mel spectrogram can be provided as input to the one or more GSR machine learning networks and / or generative neural networks as a feature input for the GSR or other speech restoration task. For example, the Mel spectrogram is configured to mimic human auditory perception, and can better adapt the GSR machine learning network(s) to perform GSR speech restoration with human-like hearing capabilities.
[0081] In a second GSR stage, high-quality speech can be restored based on the Mel spectrogram or other intermediate representation ^^ corresponding to the first GSR stage and / or generated using Eq. (3). In one illustrative example, the second GSR stage can also be referred to as a synthesis stage, where the high-quality speech signal is restored or synthesized according to: ^^^ ൌ ^^^^^^ Eq. (4)
[0082] Here, ^^^ represents the restored, high-quality speech signal generated as output by the GSR process and / or GSR machine learning network(s). The term g^∙^represents aQualcomm Ref. No.2307828WO generative function configured to generate the high-quality speech signal based on the Mel spectrogram or other intermediate representation ^^ of Eq. (3).
[0083] FIG. 5A is a diagram illustrating an example of an audio processing system 500a for general speech restoration (GSR) including one or more generative neural networks configured to generate a restored, high-quality speech signal based on relatively low-quality voice vibration sensor data, in accordance with some examples.
[0084] In one illustrative example, the audio processing system 500a can be included in and / or implemented by a wearable audio device, headset device for voice communications, TWS earbuds, etc. For instance, the audio processing system 500a can be included in and / or implemented by the wearable device 115 of FIG.1, the wearable device 205 of FIG.2, the wearable device 305 of FIG.3, and / or the voice communications device 400 of FIG.4.
[0085] The audio processing system 500a can include one or more voice vibration sensors 510. In some aspects, the voice vibration sensor 510 can be a BCM or other voice vibration sensor. In some cases, the voice vibration sensor 510 can be the same as or similar to one or more of the voice vibration sensor 140 of FIG.1, the voice vibration sensor 360 of FIG. 3, the voice vibration sensor 440 of FIG.4, etc.
[0086] The voice vibration sensor 510 can be associated with one or more feedforward microphones 520 (e.g., “FF mic”) included in or implemented by the same audio processing system 500a. For instance, the one or more feedforward microphones 520 of FIG.5A can be the same as or similar to the second microphone 424 of FIG.4. The feedforward microphone 520 can be an acoustic microphone.
[0087] The audio processing system 500a can include an audio output engine 590, which in some aspects can include one or more speakers or various other audio transducers. For instance, the audio output engine 590 can include one or more speakers that are the same as or similar to the speaker 410 of FIG. 4. In some cases, the audio output engine 590 can comprise an output buffer that is associated with or communicatively coupled to a speaker or other audio output device. For instance, the audio output engine 590 can receive a restored high-quality speech signal generated by the GSR machine learning network 540. In some examples, the restored high-quality speech signal from the GSR machine learning networkQualcomm Ref. No.2307828WO 540 can be immediately output by a speaker or audio output device included in the audio output engine 590. In some examples, the restored high-quality speech signal from the GSR machine learning network 540 can be stored in an audio output buffer implemented by the audio output engine 590.
[0088] In one illustrative example, the audio processing system 500a includes a GSR machine learning network 540. In some aspects, the GSR machine learning network 540 can include one or more generative neural networks configured to generate restored high-quality speech data from relatively low-quality input speech data. For instance, the GSR machine learning network 540 can generate high-quality restored speech data based on relatively low- quality input speech data obtained using the voice vibration sensor 510. The GSR machine learning network 540 can be implemented using a single machine learning network, can be implemented using two machine learning networks, and / or can be implemented using more than two machine learning networks. In some aspects, the GSR machine learning network 540 can implement single-stage GSR speech restoration according to Eqs. (1) and (2) above, without utilizing the Mel spectrogram or other intermediate representation ^^ of Eq. (3). In some aspects, the GSR machine learning network 540 can implement two-stage GSR speech restoration according to Eqs. (1), (3), and (4), for example based on utilizing the Mel spectrogram or other intermediate representation ^^ of Eq. (3).
[0089] In some aspects, the GSR machine learning network 540 can be implemented based on the GSR machine learning architecture 600 of FIG.6. In one illustrative example, the GSR machine learning network 540 can be the same as or similar to the GSR machine learning architecture 600 of FIG.6.
[0090] For instance, low-quality speech 605 can be obtained using a voice vibration sensor, such as the voice vibration sensor 510 of FIG. 5A. The low-quality speech 605 may be interchangeably referred to herein as low-quality speech data, a low-quality speech signal, and / or low-quality audio data (e.g., where the low-quality audio data is indicative or representative of low-quality speech).
[0091] The low-quality speech data 605 can be provided to a spectrogram engine 610 configured to generate a spectrogram corresponding to the low-quality speech data 605. For instance, the low-quality speech data 605 can be time-domain audio data obtained by theQualcomm Ref. No.2307828WO voice vibration sensor 510 of FIG. 5A, and the spectrogram engine 610 can generate a spectrogram 615 comprising a frequency-domain representation of the time-domain low- quality speech data 605. The spectrogram 615 can also be referred to as a low-quality spectrogram (e.g., based on the spectrogram 615 being a frequency-domain representation of the low-quality speech data 605). For instance, the low-quality spectrogram 615 can have a narrower frequency range or bandwidth relative to the larger frequency range or bandwidth corresponding to the restored spectrogram 655 that will be generated by the GSR machine learning network(s) using the low-quality spectrum 615.
[0092] In some aspects, the spectrogram engine 610 can be implemented as a Mel spectrogram engine and the spectrogram 615 can be a Mel spectrogram corresponding to the low-quality speech data 605 (e.g., and may also be referred to as a low-quality Mel spectrogram). For instance, the spectrogram engine 610 and the respective input and output of the low-quality speech data 605 and low-quality spectrogram 615 can correspond to Eq. (3) above (e.g., where the spectrogram engine 610 implements the restoration function f ^∙^ of Eq. (3), the low-quality spectrogram 615 corresponds to the intermediate representation ^^ of Eq. (3), and the low-quality speech data 605 corresponds to the distorted speech ^^ of Eq. (3)).
[0093] In one illustrative example, the GSR machine learning architecture 600 can include an analysis engine 640 and a synthesis engine 660. In some aspects, the analysis engine 640 can be implemented using a first machine learning network and / or generative neural network, and the synthesis engine can be implemented using a second machine learning network and / or generative neural network. In some cases, the analysis engine 640 and the synthesis engine 660 can be implemented or combined into a single machine learning network and / or single generative neural network.
[0094] In some aspects, the low-quality spectrogram 615 (e.g., low-quality Mel spectrogram or other intermediate representation ^^ of Eq. (3)) can be provided to a restoration neural network 642 included in the analysis engine 640. The restoration neural network 642 can be configured to generate a restoration mask 645 based on the input of low-quality spectrogram 615. For instance, the restoration mask 645 can comprise generated audio data that represents one or more portions of restored speech data for the low-quality spectrogram 615.Qualcomm Ref. No.2307828WO
[0095] In some examples, the restoration mask 645 includes generated audio data that represents restored speech data corresponding to times and / or frequencies that are not included in the input low-quality spectrogram 615. For instance, in examples where the low- quality spectrogram 615 is based on low-quality speech data 605 obtained from a voice vibration sensor (e.g., voice vibration sensor 510 of FIG.5A), the voice vibration sensor data 605 may represent frequencies up to approximately 800-1,000 Hz (e.g., based on attenuation of frequencies greater than 800-1,000 Hz when sensing bone-conducted sound or vibrations with a voice vibration sensor, as noted previously above). The restoration mask 645 can include generated audio data that represents restored speech for frequencies above the 800- 1,000 Hz attenuation threshold of the voice vibration sensor associated with the low-quality speech data 605 and the low-quality spectrogram 615. In another example, the restoration mask 645 can include generated audio data that represents a refinement of existing speech data at times and / or frequencies that are already included in the input low-quality spectrogram 615. For example, the restoration mask 645 can include generated audio data that increases the amplitude and / or clarity of frequencies already corresponding to human speech within the input low-quality spectrogram 615, generated audio data that masks or removes noise or distortion such as crackling or muffling, etc.
[0096] The analysis engine 640 can include or implement a summation operation 647 to combine the restoration mask 645 (e.g., generated by the restoration neural network 642) with the input low-quality spectrogram 615. Combining the restoration mask 645 and the input low-quality spectrogram 615 (e.g., using the summation operation 647) can be performed to generate the restored spectrogram 655, which includes and / or represents the restored high-quality speech data that is based on the low-quality speech data 605 from the voice vibration sensor.
[0097] In some aspects, the restored spectrogram 655 can also be referred to as a high- quality spectrogram and / or a restored high-quality spectrogram 655. In examples where the input low-quality spectrogram 615 is a Mel spectrogram, the output high-quality (e.g., restored) spectrogram 655 is also a Mel spectrogram.
[0098] The restored spectrogram 655 can be provided to a synthesis engine 660 included in the GSR machine learning architecture 600. The synthesis engine 660 can include a neuralQualcomm Ref. No.2307828WO vocoder 662 configured to transform the restored spectrogram 655 from the frequency- domain to a time-domain audio data signal. For instance, the neural vocoder can generate a time-domain high-quality speech signal 690 based on converting the restored spectrogram 655 from the frequency-domain to the time-domain. In some cases, the neural vocoder 662 can be implemented using one or more neural networks and / or other machine learning networks. In some cases, the neural vocoder 662 can be configured to implement operations that are the reverse of the spectrogram engine 610 (e.g., where the spectrogram engine 610 generates a frequency-domain spectrogram data from an input time-domain audio data, and where the neural vocoder 662 generates a time-domain audio data from an input frequency- domain spectrogram data).
[0099] The high-quality speech data 690 can be a restored speech signal generated by the GSR machine learning architecture 600 from the low-quality and / or distorted speech signal input obtained from a voice vibration sensor (e.g., such as the voice vibration sensor 510 of FIG.5A). In some aspects, the high-quality speech data 690 of FIG.6 can be the same as or similar to the restored speech signal ^^^ of Eq. (2) and / or Eq. (4). In one illustrative example, the GSR machine learning network 540 of FIG. 5A can be the same as the GSR machine learning architecture 600 of FIG.6, and the output of the GSR machine learning network 540 can be the restored high-quality speech 69i0. For instance, the restored high-quality speech 690 can be the same as the signal provided to the audio output engine 590 of FIG.5A.
[0100] In some aspects, the restoration neural network 642 and / or one or more other machine learning networks associated with or implemented by the analysis engine 640 can be based on a fully-connected deep neural network (DNN). In some aspects, the restoration neural network 642 and / or one or more other machine learning networks associated with or implemented by the analysis engine 640 can be based on a bidirectional gated recurrent unit (BiGRU) machine learning model architecture and / or a bidirectional long short-term memory (BiLSTM) machine learning model architecture. In some examples, the restoration neural network 642 and / or one or more other machine learning networks associated with or implemented by the analysis engine 640 can be based on a convolutional neural network (CNN) machine learning model architecture, such as U-Net and / or Residual U-Net (ResU- Net), etc.Qualcomm Ref. No.2307828WO
[0101] In some aspects, the neural vocoder 662 and / or one or more other machine learning networks associated with or implemented by the synthesis engine 660 can be based on a distillation network machine learning model architecture, such as WaveNet, etc. In some examples, the neural vocoder 662 and / or one or more other machine learning networks associated with or implemented by the synthesis engine 660 can be based on a likelihood flow machine learning model architecture (e.g., WaveGlow, etc.), an autoencoder machine learning model architecture (e.g., WaveVAE), a generative adversarial network (GAN) machine learning model architecture (e.g., WaveGAN), and / or one or more diffusion machine learning model architectures (e.g., DiffWave, etc.).
[0102] In some aspects, the systems and techniques described herein can be used to restore full-band clean (e.g., high-quality) speech using only a relatively low-quality speech data or signal obtained using a voice vibration sensor (e.g., voice vibration sensor 510 of FIG.5A). By restoring full-band clean speech from just the voice vibration sensor signal(s) on a headset or other audio device, the systems and techniques can eliminate the need for additional voice microphones on the headset or other audio device. For instance, the example voice communications device 400 of FIG. 4 includes a voice vibration sensor 440 and multiple outward-facing acoustic / voice microphones 422 and 424 that are used in combination to generate a voice signal for output to the speaker 410. In one illustrative example, the systems and techniques can be used to generate a high-quality voice signal for output to the speaker 410 of the voice communications device 400, using only the voice vibration sensor 440 (e.g., without utilizing or requiring the additional, acoustic microphones 422 or 424).
[0103] Removing outward-facing acoustic (e.g., voice) microphones from a headset or other voice communications device (e.g., such as the voice communications device 400 of FIG. 4) can increase the efficiency, increase the manufacturing yield, and simplify the mechanical and acoustic design of the headset or other voice communications device. The systems and techniques can also be used to provide voice communications that are more robust to environmental and / or wind noise, as environmental noise and wind noise are not detected in the bone-conducted vibration frequency band(s) corresponding to the sensitivity range of the voice vibration sensor (e.g., and the high-quality restored speech 605 can be generated based on performing GSR using only the voice vibration sensor data as input).Qualcomm Ref. No.2307828WO
[0104] In some aspects, a headset or other voice communications device implementing the systems and techniques described herein may still include at least one feed-forward acoustic microphone, such as the feed-forward acoustic microphone 520 of FIG.5A. For instance, the feed-forward acoustic microphone 520 can be included in an acoustic echo cancellation (AEC) system or sub-system of the headset or other voice communications device.
[0105] FIG.5B is a diagram illustrating an example of a GSR system 500b configured to perform general speech restoration using a vibration sensor signal 562 and an outward-facing microphone signal 564. For instance, the vibration sensor signal 562 of FIG. 5B can be a voice vibration sensor signal obtained from the voice vibration sensor 510 of FIG.5A. The outward-facing microphone signal 564 of FIG. 5B can be an acoustic microphone signal obtained from the feed-forward microphone 520 of FIG.5A.
[0106] In one illustrative example, the systems and techniques can generate a restored high-quality speech signal based on combining the voice vibration signal 562 and the outward-facing microphone signal 564 prior to the GSR machine learning network 542. The GSR machine learning network 542 of FIG. 5B can be the same as or similar to the GSR machine learning network 540 of FIG.5A (and / or can be the same as or similar to the GSR machine learning architecture 600 of FIG.6).
[0107] In some aspects, a combination operation 580 (e.g., also referred to as a summation or summation operation) can be used to combine the voice vibration sensor signal 562 with the outward-facing acoustic microphone signal 564, prior to the input of GSR machine learning network 542 of FIG.5B. For instance, the GSR machine learning network 542 can be configured to receive as input the combined voice vibration and acoustic microphone signal that is generated by the combination operation 580. In some aspects, the combination operation 580 of FIG.5B can be the same as or similar to the combination operation 530 of FIG.5A.
[0108] In some examples, the voice vibration sensor signal 562 and the outward facing acoustic microphone signal 564 can be synchronized prior to being fed to the GSR machine learning network 542 as input and / or can be synchronized prior to the combination operation 580 of FIG.5B and / or the combination operation 530 of FIG.5A. For instance, based on the relative positioning (e.g., relative separation distances, etc.) between the voice vibrationQualcomm Ref. No.2307828WO sensor 510 and the outward facing acoustic microphone 520 on the same audio device, a respective delay can be applied to the voice vibration sensor signal 562 or to the outward facing acoustic microphone signal 564, where the delayed signal is synchronized with the other, non-delayed signal.
[0109] In some aspects, the delay can be applied to the voice vibration sensor signal 562. For instance, FIG. 5B can implement a delay 575 on the transmission path of the voice vibration sensor signal 562, prior to the combination operation 580. The delay 575 can apply a configured delay value (e.g., ^^) to the voice vibration sensor signal 562, such that the delayed voice vibration sensor signal 562 (e.g., output from the delay operation 575) is synchronized with the non-delayed outward-facing microphone signal 564 of FIG.5B.
[0110] In another example, the delay can be applied to the outward-facing microphone signal 564. For instance, FIG.5A includes a delay operation 525 that is implemented on the transmission path between the feedforward microphone 520 and the combination operation 530. The delay 525 can apply a configured delay value to the acoustic or voice microphone signal from the feedforward microphone 520, such that the delayed feedforward microphone signal 564 is synchronized with the non-delayed voice vibration sensor signal 562 from voice vibration sensor 510.
[0111] In some cases, a first delay may be applied to the voice vibration sensor signal 562 and a second delay may be applied to the outward facing acoustic microphone signal 564, where the first delay value is different from the second delay value. The respectively delayed voice vibration sensor signal 562 and the respectively delayed outward-facing microphone signal 564 can be synchronized with each other before being combined at the combination operation 580 of FIG.5B.
[0112] The synchronized vibration sensor signal 562 and outward-facing microphone signal 564 can be combined using the combination operation 580 of FIG. 5B, and the combined vibration-acoustic signal can be provided as input to the GSR machine learning network 542. The GSR machine learning network 542 can be configured (e.g., trained) to generate a restored high-quality speech signal as output, where the restored high-quality speech signal is generated based on both the vibration sensor signal 562 and the acoustic microphone signal 564 represented in the combined input to the GSR machine learningQualcomm Ref. No.2307828WO network 542. The restored high-quality speech signal can be output from the GSR machine learning network 542 to an output buffer 592, which may be the same as the audio output engine 590 of FIG.5A and / or may be included in or implemented by the audio output engine 590 of FIG.5A.
[0113] FIG.7 is a flow chart illustrating an example of a process 700 for processing audio data. The process 700 can be performed by a computing device or apparatus or a component or system (e.g., one or more chipsets, one or more processors such as one or more CPUs, DSPs, NPUs, NSPs, microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc., any combination thereof, and / or other component or system) of the computing device or apparatus. In one example, the processes described herein may be performed by a wireless communication device. In one example, the processes described herein may be performed by an audio device and / or wearable device that includes one or more voice vibration sensors, BCMs, bone conduction sensors, etc. For instance, the audio device and / or wearable device can be the same as or similar to one or more of the device 115 of FIG.1, the device 205 of FIG.2, the device 305 of FIG. 3, the voice communications system 400 of FIG. 4, etc. In another example, the processes described herein may be performed by a computing device with the computing system 800 shown in FIG. 8. For instance, an audio device with the computing architecture shown in FIG. 8 may include the components of the audio devices described herein and may implement the operations of processes described herein.
[0114] At block 702, the computing device (or component thereof) can obtain an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor. For example, the wireless communication device can be the same as or similar to the voice communications system 400 of FIG.4 and the voice vibration sensor can be the same as or similar to the voice vibration sensor 440 of FIG. 4. In some examples, the voice vibration speech signal is a bone- conducted speech signal and the voice vibration sensor is a bone conduction microphone (BCM).
[0115] In some cases, the voice vibration sensor can be the same as or similar to the voice vibration sensor 510 of FIG.5A, and the voice vibration speech signal can be the same as orQualcomm Ref. No.2307828WO similar to the voice vibration sensor signal 562 of FIG. 5B. In some examples, the audio signal and / or the voice vibration speech signal can be the same as or similar to the low-quality speech signal 605 of FIG.6. In some cases, the audio signal and the voice vibration speech signal are the same. For instance, the audio signal and the voice vibration speech signal can be obtained using the voice vibration sensor can be the same as or similar to the voice vibration sensor 440 of FIG. 4 and / or the voice vibration sensor 510 of FIG. 5A, etc. The audio signal and the voice vibration speech signal can be the same as or similar to an output signal of the voice vibration sensor 510 of FIG.5A and / or the low-quality speech signal 605 of FIG.6.
[0116] In some cases, the audio signal includes the voice vibration speech signal obtained by the voice vibration sensor, and further includes one or more acoustic microphone signals obtained by one or more acoustic microphones associated with the voice vibration sensor. For instance, the one or more acoustic microphone signals can be obtained by one or more acoustic microphones that are the same as or similar to the microphone 424 and / or the microphone 422 of FIG.4, associated with the voice vibration sensor 440 of FIG.4. In some cases, the one or more acoustic microphones can include a feed-forward microphone. In some examples, the one or more acoustic microphones can be the same as or similar to the microphone 520 of FIG.5A. The one or more acoustic microphone signals can be the same as or similar to an output signal of the microphone 520 of FIG.5A, and / or can be the same as or similar to the outward-facing microphone signal 564 of FIG.5B.
[0117] In some examples, the voice vibration sensor and the one or more acoustic microphones are included in a voice communications audio device, and at least one acoustic microphone of the one or more acoustic microphones is an outward-facing microphone of the voice communications audio device. For instance, the voice vibration sensor and the one or more acoustic microphones can be included in one or more of the device 115 of FIG.1, the device 205 of FIG.2, the device 305 of FIG.3, the voice communications system 400 of FIG.4, etc.
[0118] In some cases, to obtain the audio signal, the wireless communication device (or component thereof) is further configured to apply a delay to the voice vibration speech signal obtained by the voice vibration sensor to thereby generate a delayed voice vibration speechQualcomm Ref. No.2307828WO signal, wherein the delayed voice vibration speech signal is synchronized with the one or more acoustic microphone signals. For instance, the delay can be applied based on or using the delay 525 of FIG.5A and / or the delay 575 of FIG.5B.
[0119] The computing device (or component thereof) can be further configured to combine the delayed voice vibration speech signal with the one or more acoustic microphone signals to thereby obtain the audio signal. For instance, the combining can be based on one or more of the combining operation 530 of FIG.5A and / or the combining operation 580 of FIG.5B.
[0120] At block 704, the computing device (or component thereof) can generate, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal.
[0121] For example, the GSR machine learning network can be the same as or similar to one or more of the GSR machine learning network 540 of FIG. 5A, 542 of FIG. 5B, the restoration neural network 642 of FIG. 6, the analysis engine 640 of FIG. 6, etc. In some cases, the restoration mask can be the same as or similar to the restoration mask 645 of FIG. 6. The plurality of audio samples associated with the one or more frequencies not represented in the voice vibration speech signal can be included in the restored spectrogram output 655 of FIG.6.
[0122] In some cases, the GSR machine learning network can include one or more generative neural networks. For instance, the GSR machine learning network can be the same as or similar to the analysis engine 640 of FIG. 6 and the one or more generative neural networks can be the same as or similar to the restoration neural network 642 of FIG.6.
[0123] In some examples, to generate the restoration mask, the computing device (or component thereof) can be configured to generate a Mel spectrogram based on the voice vibration speech signal. For instance, the Mel spectrogram can be the same as or similar to the low-quality spectrogram 615 of FIG. 6, generated based on the voice vibration speech signal 605 of FIG. 6. The computing device (or component thereof) can be configured to process, using a generative neural network of the GSR machine learning network, the Mel spectrogram to generate the restoration mask. For instance, the restoration neural networkQualcomm Ref. No.2307828WO 642 can be used to process the Mel spectrogram 615 of FIG. 6 to generate the restoration mask 645.
[0124] At block 706, the computing device (or component thereof) can generate, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal. For example, the restored speech signal can be the same as or similar to one or more of the restored spectrogram 655 of FIG.6 (e.g., generated based on the restoration mask 645 of FIG.6) and / or the high-quality speech signal 690 of FIG.6.
[0125] In some cases, a bandwidth associated with the voice vibration speech signal is a subset of a full-band bandwidth associated with the restored speech signal. For example, a bandwidth associated with the low-quality speech 605 and low-quality spectrogram 615 of FIG.6 can be a subset of a full-band bandwidth associated with the restored speech signal corresponding to the restored spectrogram 655 and / or the high-quality restored speech signal 690 of FIG.6.
[0126] In some cases, the audio signal and the voice vibration speech signal are the same, and the restored speech signal (e.g., restored spectrogram 655 and / or high-quality restored speech signal 690 of FIG. 6) is a full-band speech signal generated without using acoustic microphone audio data. In some examples, to generate the restoration mask, the computing device (or component thereof) is configured to generate a Mel spectrogram based on the voice vibration speech signal and process, using a generative neural network of the GSR machine learning network, the Mel spectrogram to generate the restoration mask.
[0127] In some examples, to generate the restored speech signal, the computing device (or component thereof) can be configured to generate a restored Mel spectrogram based on combining the Mel spectrogram and the restoration mask. For instance, the restored Mel spectrogram can be the same as or similar to the restored spectrogram 655 of FIG. 6, generated based on combining (e.g., at combining operation 647 of FIG.6) the low-quality Mel spectrogram 615 and the restoration mask 645 of FIG.6.Qualcomm Ref. No.2307828WO
[0128] The computing device (or component thereof) can process the restored Mel spectrogram using a neural vocoder machine learning network to generate the restored speech signal. For instance, the neural vocoder machine learning network can be the same as or similar to the neural vocoder machine learning network 662 of FIG. 6, and / or a neural vocoder machine learning network included in the synthesis engine 660 of FIG.6. In some cases, the restored speech signal is the same as or similar to the high-quality restored speech signal 690 of FIG.6. The restored speech signal can correspond to the voice vibration speech signal and the plurality of audio samples associated with the one or more frequencies not represented in the voice vibration speech signal.
[0129] In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) that are configured to carry out the steps of processes described herein. In some examples, the computing device may include a display, one or more network interfaces configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data according to the 3G, 4G, 5G, and / or other cellular standard, data according to the WiFi (802.11x) standards, data according to the BluetoothTMstandard, data according to the Internet Protocol (IP) standard, and / or other types of data.
[0130] The components of the computing device may be implemented in circuitry. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0131] The processes described herein can include a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored onQualcomm Ref. No.2307828WO one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the processes.
[0132] Additionally, the processes described herein, may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer- readable or machine-readable storage medium may be non-transitory.
[0133] FIG. 8 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. In particular, FIG.8 illustrates an example of computing system 800, which may be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 805. Connection 805 may be a physical connection using a bus, or a direct connection into processor 810, such as in a chipset architecture. Connection 805 may also be a virtual connection, networked connection, or logical connection.
[0134] In some aspects, computing system 800 is a distributed system in which the functions described in this disclosure may be distributed within a datacenter, multiple data centers, a peer network, etc. In some aspects, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some aspects, the components may be physical or virtual devices.
[0135] Example system 800 includes at least one processing unit (CPU or processor) 810 and connection 805 that communicatively couples various system components includingQualcomm Ref. No.2307828WO system memory 815, such as read-only memory (ROM) 820 and random-access memory (RAM) 825 to processor 810. Computing system 800 may include a cache 812 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 810.
[0136] Processor 810 may include any general-purpose processor and a hardware service or software service, such as services 832, 834, and 836 stored in storage device 830, configured to control processor 810 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 810 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0137] To enable user interaction, computing system 800 includes an input device 845, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 800 may also include output device 835, which may be one or more of a number of output mechanisms. In some instances, multimodal systems may enable a user to provide multiple types of input / output to communicate with computing system 800.
[0138] Computing system 800 may include communications interface 840, which may generally govern and manage the user input and system output. The communication interface may perform or facilitate receipt and / or transmission wired or wireless communications using wired and / or wireless transceivers, including those making use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, an AppleTMLightningTMport / plug, an Ethernet port / plug, a fiber optic port / plug, a proprietary wired port / plug, 3G, 4G, 5G and / or other cellular data network wireless signal transfer, a BluetoothTMwireless signal transfer, a BluetoothTMlow energy (BLE) wireless signal transfer, an IBEACONTMwireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, ad-hoc network signal transfer,Qualcomm Ref. No.2307828WO radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof. The communications interface 840 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to determine a location of the computing system 800 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0139] Storage device 830 may be a non-volatile and / or non-transitory and / or computer- readable memory device and may be a hard disk or other types of computer readable media which may store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu-ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read- only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., Level 1 (L1) cache, Level 2 (L2) cache, Level 3 (L3) cache, Level 4 (L4) cache, Level 5 (L5) cache, or other (L#) cache), resistive random-access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.Qualcomm Ref. No.2307828WO
[0140] The storage device 830 may include software services, servers, services, etc., that when the code that defines such software is executed by the processor 810, it causes the system to perform a function. In some aspects, a hardware service that performs a particular function may include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 810, connection 805, output device 835, etc., to carry out the function. The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data may be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0141] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather thanQualcomm Ref. No.2307828WO restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.
[0142] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.
[0143] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0144] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a processQualcomm Ref. No.2307828WO corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.
[0145] Processes and methods according to the above-described examples may be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions may include, for example, instructions and data which cause or otherwise configure a general-purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used may be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0146] In some aspects the computer-readable storage devices, mediums, and memories may include a cable or wireless signal containing a bitstream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0147] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof, in some cases depending in part on the particular application, in part on the desired design, in part on the corresponding technology, etc.
[0148] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-Qualcomm Ref. No.2307828WO readable or machine-readable medium. A processor(s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also may be embodied in peripherals or add-in cards. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0149] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
[0150] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random-access memory (RAM) such as synchronous dynamic random-access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that may be accessed, read, and / or executed by a computer, such as propagated signals or waves.Qualcomm Ref. No.2307828WO
[0151] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0152] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein may be replaced with less than or equal to (“^”) and greater than or equal to (“^”) symbols, respectively, without departing from the scope of this description.
[0153] Where components are described as being “configured to” perform certain operations, such configuration may be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0154] The phrase “coupled to” or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0155] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A andQualcomm Ref. No.2307828WO B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.
[0156] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
[0157] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub- functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.Qualcomm Ref. No.2307828WO
[0158] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function).
[0159] Illustrative aspects of the disclosure include:
[0160] Aspect 1. An apparatus for processing audio data, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generate, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and generate, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
[0161] Aspect 2. The apparatus of Aspect 1, wherein the voice vibration speech signal is a bone-conducted speech signal and the voice vibration sensor is a bone conduction microphone (BCM).
[0162] Aspect 3. The apparatus of any of Aspects 1 to 2, wherein the GSR machine learning network includes one or more generative neural networks.Qualcomm Ref. No.2307828WO
[0163] Aspect 4. The apparatus of any of Aspects 1 to 3, wherein the audio signal and the voice vibration speech signal are the same, and wherein the restored speech signal is a full- band speech signal generated without using acoustic microphone audio data.
[0164] Aspect 5. The apparatus of any of Aspects 1 to 4, wherein the audio signal includes the voice vibration speech signal obtained by the voice vibration sensor, and further includes one or more acoustic microphone signals obtained by one or more acoustic microphones associated with the voice vibration sensor.
[0165] Aspect 6. The apparatus of Aspect 5, wherein: the voice vibration sensor and the one or more acoustic microphones are included in a voice communications audio device; and at least one acoustic microphone of the one or more acoustic microphones is an outward- facing microphone of the voice communications audio device.
[0166] Aspect 7. The apparatus of any of Aspects 5 to 6, wherein, to obtain the audio signal, the at least one processor is further configured to: apply a delay to the voice vibration speech signal obtained by the voice vibration sensor to thereby generate a delayed voice vibration speech signal, wherein the delayed voice vibration speech signal is synchronized with the one or more acoustic microphone signals; and combine the delayed voice vibration speech signal with the one or more acoustic microphone signals to thereby obtain the audio signal.
[0167] Aspect 8. The apparatus of any of Aspects 1 to 7, wherein, to generate the restoration mask, the at least one processor is configured to: generate a Mel spectrogram based on the voice vibration speech signal; and process, using a generative neural network of the GSR machine learning network, the Mel spectrogram to generate the restoration mask.
[0168] Aspect 9. The apparatus of Aspect 8, wherein, to generate the restored speech signal, the at least one processor is configured to: generate a restored Mel spectrogram based on combining the Mel spectrogram and the restoration mask; and process the restored Mel spectrogram using a neural vocoder machine learning network to generate the restored speech signal, wherein the restored speech signal corresponds to the voice vibration speech signal and the plurality of audio samples associated with the one or more frequencies not represented in the voice vibration speech signal.Qualcomm Ref. No.2307828WO
[0169] Aspect 10. The apparatus of any of Aspects 1 to 9, wherein a bandwidth associated with the voice vibration speech signal is a subset of a full-band bandwidth associated with the restored speech signal.
[0170] Aspect 11. A method for processing audio data, comprising: obtaining an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generating, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and generating, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
[0171] Aspect 12. The method of Aspect 11, wherein the voice vibration speech signal is a bone-conducted speech signal and the voice vibration sensor is a bone conduction microphone (BCM).
[0172] Aspect 13. The method of any of Aspects 11 to 12, wherein the GSR machine learning network includes one or more generative neural networks.
[0173] Aspect 14. The method of any of Aspects 11 to 13, wherein the audio signal and the voice vibration speech signal are the same, and wherein the restored speech signal is a full- band speech signal generated without using acoustic microphone audio data.
[0174] Aspect 15. The method of any of Aspects 11 to 14, wherein the audio signal includes the voice vibration speech signal obtained by the voice vibration sensor, and further includes one or more acoustic microphone signals obtained by one or more acoustic microphones associated with the voice vibration sensor.
[0175] Aspect 16. The method of Aspect 15, wherein: the voice vibration sensor and the one or more acoustic microphones are included in a voice communications audio device; and at least one acoustic microphone of the one or more acoustic microphones is an outward- facing microphone of the voice communications audio device.Qualcomm Ref. No.2307828WO
[0176] Aspect 17. The method of any of Aspects 15 to 16, wherein obtaining the audio signal includes: applying a delay to the voice vibration speech signal obtained by the voice vibration sensor to thereby generate a delayed voice vibration speech signal, wherein the delayed voice vibration speech signal is synchronized with the one or more acoustic microphone signals; and combining the delayed voice vibration speech signal with the one or more acoustic microphone signals to thereby obtain the audio signal.
[0177] Aspect 18. The apparatus of any of Aspects 1 to 17, wherein generating the restoration mask includes: generating a Mel spectrogram based on the voice vibration speech signal; and processing, using a generative neural network of the GSR machine learning network, the Mel spectrogram to generate the restoration mask.
[0178] Aspect 19. The method of Aspect 18, wherein generating the restored speech signal includes: generating a restored Mel spectrogram based on combining the Mel spectrogram and the restoration mask; and processing the restored Mel spectrogram using a neural vocoder machine learning network to generate the restored speech signal, wherein the restored speech signal corresponds to the voice vibration speech signal and the plurality of audio samples associated with the one or more frequencies not represented in the voice vibration speech signal.
[0179] Aspect 20. The method of any of Aspects 11 to 19, wherein a bandwidth associated with the voice vibration speech signal is a subset of a full-band bandwidth associated with the restored speech signal.
[0180] Aspect 21. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any of Aspects 1 to 10.
[0181] Aspect 22. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any of Aspects 11 to 20.
[0182] Aspect 23. An apparatus for processing audio data comprising one or more means for performing operations according to any of Aspects 1 to 10.Qualcomm Ref. No.2307828WO
[0183] Aspect 24. An apparatus for processing audio data comprising one or more means for performing operations according to any of Aspects 11 to 20.
Claims
Qualcomm Ref. No.2307828WO CLAIMS What is claimed is:
1. An apparatus for processing audio data, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generate, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and generate, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
2. The apparatus of claim 1, wherein the voice vibration speech signal is a bone- conducted speech signal and the voice vibration sensor is a bone conduction microphone (BCM).
3. The apparatus of claim 1, wherein the GSR machine learning network includes one or more generative neural networks.
4. The apparatus of claim 1, wherein the audio signal and the voice vibration speech signal are the same, and wherein the restored speech signal is a full-band speech signal generated without using acoustic microphone audio data.Qualcomm Ref. No.2307828WO 5. The apparatus of claim 1, wherein the audio signal includes the voice vibration speech signal obtained by the voice vibration sensor, and further includes one or more acoustic microphone signals obtained by one or more acoustic microphones associated with the voice vibration sensor.
6. The apparatus of claim 5, wherein: the voice vibration sensor and the one or more acoustic microphones are included in a voice communications audio device; and at least one acoustic microphone of the one or more acoustic microphones is an outward-facing microphone of the voice communications audio device.
7. The apparatus of claim 5, wherein, to obtain the audio signal, the at least one processor is further configured to: apply a delay to the voice vibration speech signal obtained by the voice vibration sensor to thereby generate a delayed voice vibration speech signal, wherein the delayed voice vibration speech signal is synchronized with the one or more acoustic microphone signals; and combine the delayed voice vibration speech signal with the one or more acoustic microphone signals to thereby obtain the audio signal.
8. The apparatus of claim 1, wherein, to generate the restoration mask, the at least one processor is configured to: generate a Mel spectrogram based on the voice vibration speech signal; and process, using a generative neural network of the GSR machine learning network, the Mel spectrogram to generate the restoration mask.
9. The apparatus of claim 8, wherein, to generate the restored speech signal, the at least one processor is configured to: generate a restored Mel spectrogram based on combining the Mel spectrogram and the restoration mask; andQualcomm Ref. No.2307828WO process the restored Mel spectrogram using a neural vocoder machine learning network to generate the restored speech signal, wherein the restored speech signal corresponds to the voice vibration speech signal and the plurality of audio samples associated with the one or more frequencies not represented in the voice vibration speech signal.
10. The apparatus of claim 1, wherein a bandwidth associated with the voice vibration speech signal is a subset of a full-band bandwidth associated with the restored speech signal.
11. A method for processing audio data, comprising: obtaining an audio signal associated with a voice vibration sensor, wherein the audio signal includes a voice vibration speech signal obtained by the voice vibration sensor; generating, using a general speech restoration (GSR) machine learning network, a restoration mask based on the audio signal, wherein the restoration mask includes a plurality of audio samples associated with one or more frequencies not represented in the voice vibration speech signal; and generating, based on the audio signal and the restoration mask, a restored speech signal, wherein the restored speech signal includes the one or more frequencies not represented in the voice vibration speech signal and does not include one or more distortions represented in the voice vibration speech signal.
12. The method of claim 11, wherein the voice vibration speech signal is a bone- conducted speech signal and the voice vibration sensor is a bone conduction microphone (BCM).
13. The method of claim 11, wherein the GSR machine learning network includes one or more generative neural networks.Qualcomm Ref. No.2307828WO 14. The method of claim 11, wherein the audio signal and the voice vibration speech signal are the same, and wherein the restored speech signal is a full-band speech signal generated without using acoustic microphone audio data.
15. The method of claim 11, wherein the audio signal includes the voice vibration speech signal obtained by the voice vibration sensor, and further includes one or more acoustic microphone signals obtained by one or more acoustic microphones associated with the voice vibration sensor.
16. The method of claim 15, wherein: the voice vibration sensor and the one or more acoustic microphones are included in a voice communications audio device; and at least one acoustic microphone of the one or more acoustic microphones is an outward-facing microphone of the voice communications audio device.
17. The method of claim 15, wherein obtaining the audio signal includes: applying a delay to the voice vibration speech signal obtained by the voice vibration sensor to thereby generate a delayed voice vibration speech signal, wherein the delayed voice vibration speech signal is synchronized with the one or more acoustic microphone signals; and combining the delayed voice vibration speech signal with the one or more acoustic microphone signals to thereby obtain the audio signal.
18. The method of claim 11, wherein generating the restoration mask includes: generating a Mel spectrogram based on the voice vibration speech signal; and processing, using a generative neural network of the GSR machine learning network, the Mel spectrogram to generate the restoration mask.
19. The method of claim 18, wherein generating the restored speech signal includes: generating a restored Mel spectrogram based on combining the Mel spectrogram and the restoration mask; andQualcomm Ref. No.2307828WO processing the restored Mel spectrogram using a neural vocoder machine learning network to generate the restored speech signal, wherein the restored speech signal corresponds to the voice vibration speech signal and the plurality of audio samples associated with the one or more frequencies not represented in the voice vibration speech signal.
20. The method of claim 11, wherein a bandwidth associated with the voice vibration speech signal is a subset of a full-band bandwidth associated with the restored speech signal.
Citation Information
Patent Citations
Deep learning speech extraction and noise reduction method fusing signals of bone vibration sensor and microphone
EP4044181A1