Audio enhancement method and device and earphone
By combining feature extraction and fusion models of air conduction and bone conduction audio signals, the problem of audio enhancement in headphones under extremely low signal-to-noise ratio and interference with human voices was solved, achieving high-quality audio signal enhancement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANKER INNOVATIONS TECH CO LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are ineffective in enhancing headphone audio under conditions of extremely low signal-to-noise ratio and the presence of interfering human voices. In particular, multi-microphone audio enhancement technology cannot be applied to single-microphone situations, and simple splicing of VPU signals and Talk signals results in poor removal of interfering human voices and uneven processing effects.
A method combining air conduction and bone conduction audio signals is adopted. By training a feature extraction model and an amplitude prediction model, the voiceprint features of the bone conduction audio signal are extracted and fused with the features of the air conduction audio signal in the trained model to generate predicted amplitude to enhance the audio signal.
It effectively enhances audio signals in environments with extremely low signal-to-noise ratios and interference from human voices, achieving noise reduction without losing frequency bands and improving audio quality.
Smart Images

Figure CN121967950A_ABST
Abstract
Description
Audio enhancement methods, devices and headphones Technical Field
[0001] This application relates to the field of audio processing technology, specifically to an audio enhancement method, apparatus, and headphones. Background Technology
[0002] This section is intended to provide background or context for implementing the embodiments of the invention as set forth in the claims and detailed description. The description herein is not intended to imply that it is prior art simply because it is included in this section.
[0003] With the widespread use of mobile communication devices, people can use these devices to make calls or record audio anytime, anywhere. However, during calls or recordings, the clarity of the audio signal captured by the device deteriorates due to ambient noise and interfering human voices, affecting call or recording quality. Therefore, it is necessary to reduce audio noise to eliminate noise or interference from calls or recordings and improve audio quality. Summary of the Invention
[0004] Therefore, it is necessary to provide an audio enhancement method, device, and headphones that can improve audio quality in response to the above-mentioned technical problems.
[0005] In a first aspect, this application provides an audio enhancement method, comprising:
[0006] Acquire air conduction audio signals and bone conduction audio signals;
[0007] Acquire the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal;
[0008] The bone conduction audio signal is subjected to feature extraction using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal.
[0009] The voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal are input into a trained amplitude prediction model to obtain the predicted amplitude.
[0010] The target audio signal is obtained based on the predicted amplitude.
[0011] Secondly, this application provides an audio enhancement device, comprising:
[0012] The first acquisition module is used to acquire air conduction audio signals and bone conduction audio signals;
[0013] The second acquisition module is used to acquire the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal;
[0014] The feature extraction module is used to extract features from the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal.
[0015] The amplitude prediction module is used to input the voiceprint features of the bone conduction audio signal, the first audio feature, and the second audio feature into the trained amplitude prediction model to obtain the predicted amplitude.
[0016] An audio enhancement module is used to obtain a target audio signal based on the predicted amplitude.
[0017] Thirdly, this application also provides an earphone, including an air conduction microphone, a bone conduction microphone, a memory, and a processor. The air conduction microphone is used to collect air conduction signals, the bone conduction microphone is used to collect bone conduction signals, the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0018] Acquire air conduction audio signals and bone conduction audio signals;
[0019] Acquire the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal;
[0020] The bone conduction audio signal is subjected to feature extraction using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal.
[0021] The voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal are input into a trained amplitude prediction model to obtain the predicted amplitude.
[0022] The target audio signal is obtained based on the predicted amplitude.
[0023] The aforementioned audio enhancement method, device, and headphones acquire air-conducted audio signals and bone-conducted audio signals; acquire a first audio feature of the air-conducted audio signal and a second audio feature of the bone-conducted audio signal; extract features from the bone-conducted audio signal using a trained feature extraction model to obtain the voiceprint features of the bone-conducted audio signal; input the voiceprint features of the bone-conducted audio signal, the first audio feature, and the second audio feature into a trained amplitude prediction model to obtain the predicted amplitude; and obtain the target audio signal based on the predicted amplitude.
[0024] Audio signals acquired via air conduction have a wide frequency range, while audio signals acquired via bone conduction are almost unaffected by environmental noise. This application utilizes a primary model to extract voiceprint features from bone conduction audio signals. The extracted voiceprint features retain the advantage of low noise in bone conduction, thus aiding in noise reduction. Voiceprint features, along with audio features from both air and bone conduction audio signals, are fused in a pre-trained secondary model. The resulting predicted amplitude can be used to generate the enhanced audio signal. This dual-model mechanism combines the advantages of both bone and air conduction, resulting in a target audio signal that retains its frequency bands while achieving noise reduction, thereby improving audio quality. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 is an application environment diagram of the audio enhancement method and the fusion training method in one embodiment;
[0027] Figure 2 is a flowchart illustrating an audio enhancement method in one embodiment;
[0028] Figure 3 is a flowchart illustrating the fusion training method in one embodiment;
[0029] Figure 4 is a flowchart illustrating the audio enhancement method and fusion training method in another embodiment;
[0030] Figure 5 is a structural block diagram of an audio enhancement device in one embodiment;
[0031] Figure 6 is another structural block diagram of the audio enhancement device in one embodiment;
[0032] Figure 7 is a structural block diagram of a fusion training device in one embodiment;
[0033] Figure 8 is an internal structure diagram of a computer device in one embodiment;
[0034] Figure 9 is an internal structural diagram of a computer device in another embodiment. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to be used in the overall description of this application. The use of terms like "first" and "second" is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, the number of indicated technical features, or the sequential relationship between indicated technical features.
[0036] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.
[0037] Audio enhancement, also known as audio noise reduction, refers to the process of removing noise components from an audio signal while retaining the desired audio signal. Audio enhancement methods can be divided into traditional audio enhancement algorithms and neural network-based audio enhancement. Neural network-based audio enhancement has performance advantages, especially for unexpected noise types such as sudden noise. Due to the complexity of the noise destruction process, neural network-based audio enhancement shows significant advantages in environments with low signal-to-noise ratios and non-stationary noise.
[0038] In the field of headphone audio enhancement, neural network-based audio enhancement technology combines traditional beamforming technology with single-channel audio enhancement technology to improve call quality and intelligibility. It can achieve good results in some scenarios, but the effect will deteriorate in cases of extremely low signal-to-noise ratio and the presence of interfering human voices.
[0039] In some cases, multi-microphone audio enhancement technology can be used to address human voice interference in the audio, but its effectiveness remains poor in low signal-to-noise ratio scenarios. Furthermore, multi-microphone audio enhancement technology is not suitable for situations with only a single microphone.
[0040] With the development of VPU (bone conduction microphone), the high signal-to-noise ratio (SNR) of VPU signals at low frequencies (below 1kHz) has greatly promoted the development of audio enhancement technology for wearable devices such as headphones. Audio enhancement based on VPU signals remains effective in extremely low SNR scenarios and in situations with interfering human voices, ensuring call quality. Taking advantage of the high SNR of VPU signals at low frequencies, the low-frequency VPU signal can be spliced with the high-frequency headphone talk signal, and then processed through a single-channel audio enhancement network. However, this simple splicing of VPU and talk signals still has some problems, such as poor effectiveness in eliminating interfering human voices, a "bandwidth cutoff" phenomenon affecting the listening experience, and poor high-frequency recovery.
[0041] To address the aforementioned problems, this application proposes an audio enhancement method and a fusion training method. Please refer to Figure 1, which illustrates an application environment for the audio enhancement method and fusion training method in one embodiment. The audio enhancement method and fusion training method provided in this application embodiment can be applied to the application environment shown in Figure 1. In this environment, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on other servers.
[0042] Both the terminal and the server can be used independently to execute the audio enhancement method and fusion training method provided in the embodiments of this application.
[0043] For example, the terminal acquires air conduction audio signal samples and bone conduction audio signal samples; it acquires the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; the terminal uses the bone conduction audio signal samples as input data for the feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model, and performs fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
[0044] The terminal acquires air conduction audio signals and bone conduction audio signals; the terminal acquires the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal; the terminal extracts features from the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal; the terminal inputs the voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal into a trained amplitude prediction model to obtain the predicted amplitude; the terminal obtains the target audio signal based on the predicted amplitude.
[0045] In addition, the terminal and server can also work together to execute the audio enhancement method and fusion training method provided in the embodiments of this application.
[0046] For example, the server acquires air conduction audio signal samples and bone conduction audio signal samples; the server acquires the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; the server uses the bone conduction audio signal samples as input data for the feature extraction model, uses the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, uses the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model, and performs fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
[0047] The server provides an interface for the terminal to call the model. The terminal acquires air conduction audio signals and bone conduction audio signals. The terminal acquires the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal. The terminal extracts features from the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal. The terminal inputs the voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal into a trained amplitude prediction model to obtain the predicted amplitude. The terminal obtains the target audio signal based on the predicted amplitude.
[0048] The terminal and server can also work together to execute the audio enhancement method provided in the embodiments of this application, and work together to execute the fusion training method provided in the embodiments of this application.
[0049] For example, the terminal acquires air conduction audio signal samples and bone conduction audio signal samples; the terminal acquires the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; the terminal sends the bone conduction audio signal samples, the third audio feature, and the fourth audio feature to the server. The server uses the bone conduction audio signal samples as input data for the feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model. The server then performs fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
[0050] The terminal acquires air conduction audio signals and bone conduction audio signals; the terminal acquires a first audio feature of the air conduction audio signal and a second audio feature of the bone conduction audio signal; the terminal sends the bone conduction audio signal, the first audio feature, and the second audio feature to the server. The server extracts features from the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal; the server inputs the voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal into a trained amplitude prediction model to obtain the predicted amplitude; the server obtains the target audio signal based on the predicted amplitude; the server sends the target audio signal to the terminal.
[0051] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, IoT device, or portable wearable device. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, or smart glasses.
[0052] Server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0053] Terminal 102 and server 104 can be connected via Bluetooth, USB (Universal Serial Bus) or network, etc., and this application does not impose any restrictions.
[0054] In an exemplary embodiment, as shown in FIG2, an audio enhancement method is provided. Taking the application of this method to the terminal in FIG1 as an example, the method includes the following steps S202 to S210. Wherein:
[0055] Step S202: Acquire air conduction audio signal and bone conduction audio signal.
[0056] The air conduction audio signal and bone conduction audio signal obtained in this application refer to air conduction audio signal and bone conduction audio signal obtained based on the same audio. The air conduction audio signal and bone conduction audio signal of the same audio contain the same or similar audio content, for example, they can be audio signals obtained for the same scene in the same time period.
[0057] The computer device can be a wearable device, which may include an air conduction microphone and a bone conduction microphone. The bone conduction microphone is positioned at the wearing location of the wearable device, close to the user. When the wearer emits a sound, the sound is transmitted to the bone conduction microphone via bone conduction and is captured by the bone conduction microphone as a raw bone conduction signal.
[0058] In scenarios such as voice calls, the wearer emits sound while wearing the wearable device. The wearable device collects the wearer's voice signal as the raw bone conduction signal through a bone conduction microphone and as the raw air conduction signal through an air conduction microphone. Optionally, the air conduction microphone can also simultaneously collect the wearer's voice signal and the ambient sound signal of the wearer's environment, and use the simultaneously collected voice signal and ambient sound signal as the raw air conduction signal.
[0059] The air conduction audio signal can be the raw air conduction signal acquired by the air conduction microphone, or the audio signal obtained after preprocessing the raw air conduction signal acquired by the air conduction microphone; the bone conduction audio signal can be the raw bone conduction signal acquired by the bone conduction microphone, or the audio signal obtained after preprocessing the raw bone conduction signal acquired by the bone conduction microphone. Preprocessing may include, but is not limited to, feature transformation, signal segmentation, and signal alignment.
[0060] For example, a raw air conduction signal can be acquired using an air conduction microphone, and a feature transformation can be performed on the raw air conduction signal to obtain its frequency domain signal. Based on the raw air conduction signal and its frequency domain signal, an air conduction audio signal can be obtained. Similarly, a raw bone conduction signal can be acquired using a bone conduction microphone, and a feature transformation can be performed on the raw bone conduction signal to obtain its frequency domain signal. Based on the raw bone conduction signal and its frequency domain signal, a bone conduction audio signal can be obtained. The feature transformation is a transformation that converts the signal from the time domain to the frequency domain, including but not limited to short-time Fourier transform, wavelet transform, continuous wavelet transform, and S-transform.
[0061] This application does not limit the number of air conduction microphones and bone conduction microphones. There can be one or more raw air conduction signals and one or more raw bone conduction signals, correspondingly resulting in one or more bone conduction audio signals and one or more air conduction audio signals. For example, multiple air conduction microphones and one bone conduction microphone can be set in the device. Multiple raw air conduction signals can be acquired through multiple air conduction microphones, and each raw air conduction signal can be preprocessed to obtain multiple air conduction audio signals. Similarly, one raw bone conduction signal can be acquired through one bone conduction microphone, and this bone conduction signal can be preprocessed to obtain multiple bone conduction audio signals.
[0062] Step S204: Obtain the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal.
[0063] Audio features are those that contain audio information and reflect the characteristics of an audio signal. Audio features can include amplitude and phase, as well as other features that reflect the characteristics of the audio signal, such as duration and source.
[0064] It should be noted that the terms "first" and "second" in "first audio feature" and "second audio feature" are only used to distinguish the source of the audio features, namely, they originate from air conduction audio signals and bone conduction audio signals, respectively. The first and second audio features can be the same type of audio signals obtained by performing the same processing on air conduction audio signals and bone conduction audio signals.
[0065] Step S206: Extract features from the bone conduction audio signal using the trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal.
[0066] Since the bone conduction audio signal originates directly or indirectly from the raw bone conduction signal collected by the bone conduction microphone, thanks to the sound propagation method of bone conduction, the raw bone conduction signal contains less noise, and consequently, the bone conduction audio signal contains less interference. To achieve the strongest possible noise reduction effect, this application first performs feature processing on the bone conduction audio signal to obtain the voiceprint feature of the bone conduction audio signal. This voiceprint feature can play an auxiliary role in noise reduction during subsequent feature fusion.
[0067] To achieve feature extraction from bone conduction audio signals, this application pre-trains a feature extraction model. This model can be constructed based on a neural network, containing multiple layers such as input, output, hidden, and normalization layers, interconnected by weights and activation functions. The feature extraction model is trained using bone conduction audio signal samples and their corresponding voiceprint feature labels as a training set, enabling it to learn how to extract voiceprint features from bone conduction audio signals. The trained feature extraction model receives bone conduction audio signals as input data, performs feature extraction on the signals, and obtains the voiceprint features of the bone conduction audio signals.
[0068] In one embodiment, when the neural network is a convolutional neural network, the feature extraction model may also include convolutional layers, such as convolutional processing of the received data through convolutional layers to extract the voiceprint features of the bone conduction audio signal.
[0069] Step S208: Input the voiceprint features, first audio features, and second audio features of the bone conduction audio signal into the trained amplitude prediction model to obtain the predicted amplitude.
[0070] In addition to the feature extraction model, this application includes a pre-trained amplitude prediction model. This amplitude prediction model is used to predict the amplitude based on the voiceprint features of the input bone conduction audio signal and the audio features of both the bone conduction and air conduction audio signals. The predicted amplitude includes amplitude information of the expected enhancement result of the acquired sound.
[0071] To achieve the prediction of amplitude based on the voiceprint features of the input bone conduction audio signal and the audio features of the bone conduction audio signal and the air conduction audio signal, this application pre-constructs an amplitude prediction model based on a neural network, and uses the learning ability of the neural network to train the amplitude prediction model to learn the correspondence between these input data and the predicted amplitude.
[0072] The amplitude prediction model can contain multiple layers, such as an input layer, output layer, hidden layer, and normalization layer, which are interconnected by weights and activation functions. The model is trained using the voiceprint features of bone conduction audio signal samples and the audio features of both bone conduction and air conduction audio signal samples as input data, and the amplitude of either the audio signal sample or the air conduction audio signal sample as the target data. Training ends when a preset termination condition is met, resulting in the trained amplitude prediction model.
[0073] By inputting the voiceprint features, first audio features, and second audio features of the bone conduction audio signal into a trained amplitude prediction model, the predicted amplitude can be obtained. The pre-output predicted amplitude will vary depending on the target data used during model training. By selecting different target data during training, the amplitude prediction model can selectively output predicted amplitudes for enhancing bone conduction audio signals or for enhancing air conduction audio signals.
[0074] For example, if the amplitude of the bone conduction audio signal sample is used as the target data, the amplitude prediction model learns the relationship between the input data and the amplitude of the bone conduction audio signal sample during training. After training ends and a trained amplitude prediction model is obtained, the voiceprint features, first audio features and second audio features of the bone conduction audio signal are input into the trained amplitude prediction model, and the obtained predicted amplitude is the predicted amplitude used to enhance the bone conduction audio signal.
[0075] If the amplitude of the air conduction audio signal sample is used as the target data, the amplitude prediction model learns the relationship between the input data and the amplitude of the air conduction audio signal sample during training. After training ends and a trained amplitude prediction model is obtained, the voiceprint features, first audio features and second audio features of the air conduction audio signal are input into the trained amplitude prediction model, and the obtained predicted amplitude is used to enhance the predicted amplitude of the air conduction audio signal.
[0076] Step S210: Obtain the target audio signal based on the predicted amplitude.
[0077] In signal processing, amplitude and phase are two fundamental properties describing a signal wave. Amplitude refers to the maximum distance the signal wave deviates from a reference value, while phase refers to the position of the signal wave at a given moment relative to a reference time point within the signal wave. The product of amplitude and phase is the negative representation of the signal wave. After obtaining the predicted amplitude, multiplying the predicted amplitude by a specific phase allows for the reconstruction of the sound signal wave, thus achieving audio enhancement of the corresponding audio signal.
[0078] Based on the audio enhancement requirements of different enhancement objects, the predicted amplitude can be multiplied by different phases to reconstruct the signal wave of the enhancement object as the enhancement result of the audio signal of the enhancement object, which is the target audio signal.
[0079] For example, the enhancement target can be bone conduction-related audio signals, such as the original bone conduction signal or bone conduction audio signal. In the training phase, the phase of the bone conduction audio signal sample is used as the target data to train the amplitude prediction model. When the target audio signal is obtained based on the predicted amplitude, the predicted amplitude is multiplied by the phase of the bone conduction audio signal, and the resulting target audio signal is the enhancement result of the bone conduction-related audio signal.
[0080] For example, the enhancement target can be air conduction-related audio signals, such as the original air conduction signal or air conduction audio signal. In the training phase, the phase of the air conduction audio signal sample is used as the target data to train the amplitude prediction model. When the target audio signal is obtained based on the predicted amplitude, the predicted amplitude is multiplied by the phase of the air conduction audio signal, and the resulting target audio signal is the enhanced result of the air conduction-related audio signal.
[0081] The aforementioned audio enhancement method acquires air-conduction audio signals and bone-conduction audio signals; acquires first audio features of the air-conduction audio signals and second audio features of the bone-conduction audio signals; extracts features from the bone-conduction audio signals using a trained feature extraction model to obtain voiceprint features; inputs the voiceprint features, first audio features, and second audio features of the bone-conduction audio signals into a trained amplitude prediction model to obtain predicted amplitude; and obtains the target audio signal based on the predicted amplitude. Audio signals acquired through air conduction have a wide frequency range, while audio signals acquired through bone conduction are almost unaffected by environmental noise. This application utilizes a first-level model to extract voiceprint features from the bone-conduction audio signals. The extracted voiceprint features retain the advantage of low noise in bone conduction, thus aiding in noise reduction. The voiceprint features, along with the audio features of the air-conduction and bone-conduction audio signals, are fused in a pre-trained second-level model. The resulting predicted amplitude can be used to generate the enhanced audio signal. This dual-model mechanism combines the advantages of both bone and air conduction, resulting in a target audio signal that retains frequency bands while achieving noise reduction, thereby improving audio quality.
[0082] In one embodiment, acquiring air conduction audio signals and bone conduction audio signals includes:
[0083] The original air conduction signal was acquired using an air conduction microphone, and the original bone conduction signal was acquired using a bone conduction microphone. A short-time Fourier transform was performed on the original air conduction signal to obtain its frequency domain signal. An air conduction audio signal was obtained based on the original air conduction signal and its frequency domain signal. A short-time Fourier transform was performed on the original bone conduction signal to obtain its frequency domain signal. A bone conduction audio signal was obtained based on the original bone conduction signal and its frequency domain signal.
[0084] The system simultaneously acquires raw air conduction signals and raw bone conduction signals using air conduction and bone conduction microphones. These acquired raw air conduction and raw bone conduction signals are time-domain signals. For example, during a user's call, the system acquires the user's audio signal in real-time using the air conduction microphone to obtain the raw air conduction signal, and the system acquires the user's audio signal in real-time using the bone conduction microphone to obtain the raw bone conduction signal.
[0085] To introduce frequency domain transformation, the original air conduction and bone conduction signals can be converted to the frequency domain using a short-time Fourier transform, yielding their frequency domain signals. Since the original air conduction and bone conduction signals are time domain signals, combining these two signals yields the time-frequency domain air conduction audio signal, and vice versa. Thus, the original air conduction and bone conduction signals are converted to the time-frequency domain, allowing for the incorporation of frequency domain information in subsequent model processing.
[0086] In the above embodiments, by using short-time Fourier transform, the original bone conduction signal and original air conduction signal in the time domain can be converted into bone conduction audio signal and air conduction audio signal in the time-frequency domain that contain more information before being input into the model. This allows the model to learn the frequency domain information of the bone conduction audio signal and air conduction audio signal and output more accurate prediction results.
[0087] In one embodiment, obtaining the target audio signal based on the predicted amplitude includes:
[0088] The predicted amplitude is multiplied by the phase of the air-conducted audio signal, and the result of the multiplication is used as the enhanced signal in the time-frequency domain. The enhanced signal is then subjected to a short-time inverse Fourier transform to obtain the target audio signal.
[0089] If a feature transformation is used when converting the original air conduction signal and the original bone conduction signal to the time-frequency domain, then after obtaining the enhanced signal in the time-frequency domain, the inverse feature transformation is used to process the enhanced signal to convert the enhanced signal in the time-frequency domain to the time domain, thereby obtaining the target audio signal that can be directly heard.
[0090] If the feature transformation method is short-time Fourier transform, then after obtaining the enhanced signal in the time-frequency domain, the enhanced signal is subjected to inverse short-time Fourier transform to obtain the target audio signal.
[0091] In the above embodiments, the signals of bone conduction and air conduction can be converted back and forth between the time domain and the time-frequency domain through short-time Fourier transform and inverse short-time Fourier transform. After the model obtains the expected prediction amplitude, the inverse time-frequency domain to time domain transformation is performed to obtain a clear target audio signal.
[0092] In one embodiment, there are multiple air conduction microphones. When acquiring raw air conduction signals based on the air conduction microphones, raw air conduction signals can be acquired from different directions based on multiple air conduction microphones to obtain multiple raw air conduction signals.
[0093] This application does not limit the number of air conduction microphones. When there are multiple air conduction microphones, each air conduction microphone can be installed in different orientations to collect raw air conduction signals from different directions, thus obtaining multiple raw air conduction signals from different directions.
[0094] Taking headphones as an example, multiple air conduction microphones can be installed on the same headphone in different orientations (such as front, back, left, right, up, and down). These microphones can then collect sound signals from multiple directions within the headphone, resulting in multiple raw air conduction signals. Due to the different sampling directions, the sounds contained in these raw air conduction signals, as well as their intensity and location, will vary, thus enriching the amount of sound information contained within the raw air conduction signals. In some cases, this also facilitates more accurate localization of individual sounds, thereby improving audio quality and enabling audio processing such as stereo reproduction.
[0095] In one embodiment, the plurality of air conduction microphones includes unidirectional air conduction microphones and omnidirectional air conduction microphones. Based on the acquisition of raw air conduction signals from different directions using the plurality of air conduction microphones, the resulting plurality of raw air conduction signals include:
[0096] The target direction where the sound signal is greater than a preset decibel threshold is determined, and the sound signal in the target direction is collected directionally using a unidirectional air conduction microphone to obtain a unidirectional raw air conduction signal; the sound signal in all directions is collected using an omnidirectional air conduction microphone to obtain an omnidirectional raw air conduction signal.
[0097] In this embodiment, the device may include unidirectional and omnidirectional air conduction microphones among its multiple air conduction microphones. The omnidirectional air conduction microphone collects sound signals from all directions to obtain an omnidirectional raw air conduction signal. For example, during a user's call, the user's voice signal is collected using a bone conduction microphone to obtain a raw bone conduction signal. A unidirectional air conduction microphone is used to directionally collect sound signals from the user's location to obtain a unidirectional raw air conduction signal. Finally, an omnidirectional raw air conduction microphone collects sound signals from all directions within the user's environment to obtain an omnidirectional raw air conduction microphone signal. The sound signals from all directions within the user's environment may include both the user's voice signal and background noise signals from the user's surroundings.
[0098] In one scenario, when using a unidirectional air conduction microphone to directionally acquire sound signals from one or more directions, the target direction of the sound-emitting object (such as a user) can be determined by the volume of the sound. For example, a decibel threshold can be preset, and sound signals from different directions can be continuously monitored. When a sound signal exceeding the preset decibel threshold is detected, the target direction of the sound signal exceeding the preset decibel threshold is determined. The unidirectional air conduction microphone is then used to directionally acquire the sound signal from the target direction, thereby obtaining the sound signal of the sound-emitting object as the unidirectional raw air conduction signal.
[0099] In some cases, it is not necessary to set a specific acquisition target. Sound signals from different directions are continuously monitored. When a sound signal exceeding a preset decibel threshold is detected, this sound signal is used as the acquisition target. A unidirectional air conduction microphone is used to directionally acquire the sound signal exceeding the preset decibel threshold, obtaining a unidirectional raw air conduction signal.
[0100] In the above embodiments, the unidirectional air conduction microphone serves as the main air conduction microphone, acquiring relatively pure audio signals and ensuring a high signal-to-noise ratio, so that the model can easily extract the target sound during processing. The omnidirectional air conduction microphone serves as the auxiliary air conduction microphone, utilizing the characteristic of air conduction that does not miss any frequency bands to ensure the integrity of the audio signal input to the model. The two work together to improve the accuracy of the predicted amplitude of the model output.
[0101] In one embodiment, the air conduction audio signal includes multiple signals. A short-time Fourier transform is performed on the original air conduction signal to obtain its frequency domain signal. Based on the original air conduction signal and its frequency domain signal, the air conduction audio signal is obtained, including:
[0102] Short-time Fourier transform is performed on multiple original air conduction signals to obtain the frequency domain signal of each original air conduction signal. Based on each original air conduction signal and its corresponding frequency domain signal, multiple air conduction audio signals are obtained.
[0103] Among them, multiple air conduction audio signals include unidirectional air conduction audio signals converted from unidirectional original air conduction signals.
[0104] Unidirectional and omnidirectional air conduction microphones can simultaneously acquire multiple raw air conduction signals. When performing a short-time Fourier transform on the raw air conduction signals to obtain their frequency domain signals, and then obtaining the air conduction audio signal based on the raw air conduction signals and their frequency domain signals, each raw air conduction signal is processed separately. Taking one microphone acquiring one signal as an example, the unidirectional raw air conduction signal acquired by each unidirectional air conduction microphone is converted into a frequency domain signal using a short-time Fourier transform, and the unidirectional air conduction audio signal is obtained based on each unidirectional raw air conduction signal and its frequency domain signal; similarly, the omnidirectional raw air conduction signal acquired by each omnidirectional air conduction microphone is converted into a frequency domain signal using a short-time Fourier transform, and the omnidirectional air conduction audio signal is obtained based on each omnidirectional raw air conduction signal and its frequency domain signal.
[0105] It is understood that this application does not limit the number of unidirectional air conduction microphones, omnidirectional air conduction microphones, unidirectional raw air conduction signals, omnidirectional raw air conduction signals, unidirectional air conduction audio signals, and omnidirectional air conduction audio signals.
[0106] In one embodiment, multiplying the predicted amplitude by the phase of the air-conducting audio signal and using the result of the multiplication as an enhanced signal in the time-frequency domain includes: multiplying the predicted amplitude by the phase of the unidirectional air-conducting audio signal and using the result of the multiplication as an enhanced signal in the time-frequency domain.
[0107] When the device is equipped with both a unidirectional air conduction microphone and an omnidirectional air conduction microphone, the time-frequency domain enhancement signal obtained based on the predicted amplitude includes: multiplying the predicted amplitude by the phase of the unidirectional air conduction audio signal, and using the result of the multiplication as the time-frequency domain enhancement signal.
[0108] If there are multiple unidirectional air conduction microphones, one of them is selected as the master air conduction microphone. The predicted amplitude is multiplied by the phase of the unidirectional air conduction audio signal corresponding to the master air conduction microphone, and the result is used as the time-frequency domain enhancement signal. For example, the distance from each unidirectional air conduction microphone to the sound source can be obtained, and the unidirectional air conduction microphone closest to the sound source can be selected as the master air conduction microphone. Alternatively, the average value of the original air conduction signals acquired by each unidirectional air conduction microphone can be determined, and the unidirectional air conduction microphone with the largest average value of the acquired original air conduction signals can be selected as the master air conduction microphone.
[0109] In the above embodiments, multiple air conduction microphones are used to collect the original air conduction signal. The different collection directions of multiple air conduction microphones can be utilized, and combined with multi-microphone audio enhancement technology, the amount of information that the model can process can be further increased, so that the model can output a more accurate prediction amplitude based on more information.
[0110] In one embodiment, before inputting the voiceprint features, first audio features, and second audio features of the bone conduction audio signal into a trained amplitude prediction model to obtain the predicted amplitude, the method further includes:
[0111] Acquire air conduction audio signal samples and bone conduction audio signal samples; acquire the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; use the bone conduction audio signal samples as input data for the feature extraction model, use the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, use the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model, and perform fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
[0112] Before using the amplitude prediction model to obtain the predicted amplitude, this application first constructs and trains the amplitude prediction model. Specifically, the data preprocessing steps during amplitude prediction model training are the same as those used when using the amplitude prediction model. For details on acquiring air conduction audio signal samples and bone conduction audio signal samples, please refer to the relevant descriptions in the foregoing embodiments. For details on acquiring the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples, please refer to the relevant descriptions in the foregoing embodiments on acquiring the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal.
[0113] For example, when acquiring air conduction audio signal samples and bone conduction audio signal samples, original air conduction signal samples are acquired based on air conduction microphones, and original bone conduction signal samples are acquired based on bone conduction microphones. Specifically, unidirectional original air conduction signal samples are acquired directionally based on unidirectional air conduction microphones, and omnidirectional original air conduction signal samples are acquired based on omnidirectional air conduction microphones. Short-time Fourier transform is performed on the original air conduction signal samples to obtain frequency domain signal samples of the original air conduction signal samples. Based on the original air conduction signal samples and the frequency domain signal samples of the original air conduction signal samples, air conduction audio signal samples are obtained. Short-time Fourier transform is performed on the original bone conduction signal samples to obtain frequency domain signal samples of the original bone conduction signal samples. Based on the original bone conduction signal samples and the frequency domain signal samples of the original bone conduction signal samples, bone conduction audio signal samples are obtained.
[0114] The third and fourth audio features are the audio features of the air conduction audio signal sample and the bone conduction audio signal sample, respectively, and the third and fourth audio features are the same type of audio features as the first and second audio features. For example, they are all amplitude and phase.
[0115] In one embodiment, a fusion training method is used to train the dual model of this application. Specifically, bone conduction audio signal samples are used as input data for the feature extraction model, and the output of the feature extraction model, along with a third and fourth audio feature, are used as input data for the amplitude prediction model. The feature extraction model and the amplitude prediction model are then fused and trained. The data processing methods during the training of the feature extraction model and the amplitude prediction model can be found in the relevant descriptions in the aforementioned embodiments of the audio enhancement method, and will not be repeated here.
[0116] The target output of the fusion training can be determined according to actual needs. If it is necessary to enhance the bone conduction audio signal, the amplitude of the bone conduction audio signal sample is used as the target output of the amplitude prediction model; if it is necessary to enhance the air conduction audio signal, the amplitude of the air conduction audio signal sample is used as the target output of the amplitude prediction model. When there are multiple air conduction microphones in the device and multiple raw air conduction signal samples are collected, the amplitude of the air conduction audio signal sample corresponding to the raw air conduction signal sample of the main air conduction microphone is used as the target output.
[0117] When training the feature extraction model and the amplitude prediction model together, only one set of loss functions is set for each model. The amplitude prediction model continuously obtains its predicted output based on the output of the feature extraction model. The predicted output of the feature extraction model is compared with the target output corresponding to the input data based on the set loss function. By continuously reducing the loss value of the loss function, the feature extraction model and the amplitude prediction model are optimized synchronously to achieve the purpose of fusion training.
[0118] In the above embodiments, the feature extraction model and the amplitude prediction model are trained together to learn the relationship between the output of the feature extraction model and the audio features of the air conduction audio signal samples and bone conduction audio signal samples and the amplitude of the air conduction audio signal samples. The trained dual model can be used to predict the desired amplitude of audio in audio enhancement, thereby achieving audio enhancement and improving audio quality.
[0119] In an exemplary embodiment, as shown in FIG3, a fusion training method is provided. Taking the application of this method to the terminal in FIG1 as an example, the method includes the following steps S302 to S306. Wherein:
[0120] Step S302: Obtain air conduction audio signal samples and bone conduction audio signal samples.
[0121] Step S304: Obtain the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample.
[0122] Step S306: Use the bone conduction audio signal sample as input data for the feature extraction model, use the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, use the amplitude of the air conduction audio signal sample as the target output of the amplitude prediction model, and perform fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
[0123] In one embodiment, acquiring air conduction audio signal samples and bone conduction audio signal samples includes:
[0124] Raw air conduction signal samples were acquired using an air conduction microphone and raw bone conduction signal samples were acquired using a bone conduction microphone. Specifically, unidirectional raw air conduction signal samples were acquired using a unidirectional air conduction microphone, and omnidirectional raw air conduction signal samples were acquired using an omnidirectional air conduction microphone. Short-time Fourier transform was performed on the raw air conduction signal samples to obtain frequency domain signal samples. Based on the raw air conduction signal samples and the frequency domain signal samples, air conduction audio signal samples were obtained. Short-time Fourier transform was performed on the raw bone conduction signal samples to obtain frequency domain signal samples. Based on the raw bone conduction signal samples and the frequency domain signal samples, bone conduction audio signal samples were obtained.
[0125] The aforementioned fusion training method acquires air conduction audio signal samples and bone conduction audio signal samples; it acquires the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; it uses the bone conduction audio signal samples as input data for the feature extraction model, and the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model. The feature extraction model and the amplitude prediction model are then fused and trained to obtain the trained feature extraction model and the trained amplitude prediction model. Audio signals acquired through air conduction have a wide frequency range, while audio signals acquired through bone conduction are almost unaffected by environmental noise. This application designs and trains a dual-model mechanism. The first-level model can extract voiceprint features from the bone conduction audio signal, and the extracted voiceprint features continue the advantage of low noise in bone conduction, which can assist in noise reduction. The second-level model can fuse the voiceprint features output by the first-level model with the audio features of the air conduction audio signal and the bone conduction audio signal. The primary and secondary models are trained together to learn the output of the feature extraction model and the relationship between the audio features of air-conducted and bone-conducted audio signal samples and the amplitude of the air-conducted audio signal samples. The trained dual model can be used to predict the desired amplitude of audio in audio enhancement, thereby achieving audio enhancement and improving audio quality.
[0126] This application also provides a specific embodiment, please refer to Figure 4. The specific embodiment of the audio enhancement method and fusion training method in this application scenario is as follows:
[0127] S1. Collect raw air conduction signal samples based on an air conduction microphone and raw bone conduction signal samples based on a bone conduction microphone.
[0128] Among them, unidirectional raw air conduction signal samples are collected based on unidirectional air conduction microphones, and omnidirectional raw air conduction signal samples are collected based on omnidirectional air conduction microphones.
[0129] S2. Perform a short-time Fourier transform on the original air-conduction signal sample to obtain the frequency-domain signal sample of the original air-conduction signal sample. Based on the original air-conduction signal sample and the frequency-domain signal sample of the original air-conduction signal sample, obtain the air-conduction audio signal sample.
[0130] S3. Obtain the amplitude and phase of the air-conduction audio signal sample.
[0131] S4. Perform a short-time Fourier transform on the original bone-conduction signal sample to obtain the frequency-domain signal sample of the original bone-conduction signal sample. Based on the original bone-conduction signal sample and the frequency-domain signal sample of the original bone-conduction signal sample, obtain the bone-conduction audio signal sample.
[0132] S5. Obtain the amplitude and phase of the bone-conduction audio signal sample.
[0133] S6. Use the bone-conduction audio signal sample as the input data of the feature extraction model. Use the output of the feature extraction model, the amplitude and phase of the air-conduction audio signal sample, and the amplitude and phase of the bone-conduction audio signal sample as the input data of the amplitude prediction model. Use the amplitude of the air-conduction audio signal sample as the target output of the amplitude prediction model. Perform fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
[0134] S7. Determine the target direction where the audio signal is greater than the preset decibel threshold, and based on the unidirectional air-conduction microphone, perform directional acquisition on the audio signal in the target direction to obtain the unidirectional original air-conduction signal.
[0135] S8. Based on the omnidirectional air-conduction microphone, collect the audio signals in all directions to obtain the omnidirectional original air-conduction signal.
[0136] S9. Perform a short-time Fourier transform on multiple original air-conduction signals to obtain the frequency-domain signals of each original air-conduction signal. Based on each original air-conduction signal and the corresponding frequency-domain signal, obtain multiple air-conduction audio signals.
[0137] Among them, the multiple air-conduction audio signals include the unidirectional air-conduction audio signal converted from the unidirectional original air-conduction signal. <H
[0138] The original air-conduction signal m with a length of T in the time domain q can be expressed as m q (t), where t represents time, 0 < t ≤ T. Then, through the short-time Fourier transform, the original air-conduction signal m q (t) in the time domain is converted to the time-frequency domain, and the obtained air-conduction audio signal can be expressed by formula (1):
[0139] M q (n,k)=STFT(m q (t)) (1)
[0140] Where n is the frame sequence, 0 < n ≤ N, N is the total number of frames, k is the center frequency sequence, 0 < k ≤ K, K is the total number of frequency points. Here, q represents the air conduction microphone, 0 < q ≤ Q, Q is the total number of original air conduction signals. In some cases, Q is also equal to the total number of air conduction microphones.
[0141] When there is one air conduction microphone in the device, Q = 1. The original air conduction signal is collected through one air conduction microphone, and the original air conduction signal is transformed into the time-frequency domain through short-time Fourier transform to obtain an air conduction audio signal; when there are multiple air conduction microphones in the device, Q > 1. The original air conduction signals are collected through multiple air conduction microphones, and each original air conduction signal is transformed into the time-frequency domain through short-time Fourier transform to obtain multiple air conduction audio signals.
[0142] S10. Obtain the amplitudes and phases of multiple air conduction audio signals.
[0143] Obtain multiple air conduction audio signals M q The amplitude Mag of M(n,k) can be expressed by formula (2), and obtain multiple air conduction audio signals M q The phase Pha of M(n,k) can be expressed by formula (3):
[0144] MagM q M(n,k) = abs(M q (n,k)) (2)
[0145] PhaM q M(n,k)= M q (n,k) / abs(M q (n,k)) (3)
[0146] S11. Based on the air conduction microphone to collect the original air conduction signal, and based on the bone conduction microphone to collect the original bone conduction signal.
[0147] S12. Perform short-time Fourier transform on the original bone conduction signal to obtain the frequency domain signal of the original bone conduction signal, and based on the original bone conduction signal and the frequency domain signal of the original bone conduction signal, obtain the bone conduction audio signal.
[0148] The original bone conduction signal v with a length of T in the time domain can be expressed as v(t), where t represents time, 0 < t ≤ T. Then, after transforming the original bone conduction signal v(t) in the time domain into the time-frequency domain through short-time Fourier transform, the obtained bone conduction audio signal can be expressed by formula (z):
[0149] V(n,k)=STFT(v(t)) (4)
[0150] where n is the frame sequence of the bone-conducted audio signal, 0 < n ≤ N, N is the total number of frames, k is the central frequency sequence, 0 < k ≤ K, and K is the total number of frequency points.
[0151] S13. Obtain the amplitude and phase of the bone-conducted audio signal.
[0152] The amplitude Mag of the bone-conducted audio signal V(n,k) can be expressed by Equation (5), and the phase Pha of the bone-conducted audio signal V(n,k) can be expressed by Equation (6):
[0153] MagV(n,k) = abs(V(n,k)) (5)
[0154] PhaV(n,k) = V(n,k) / abs(V(n,k)) (6)
[0155] S14. Perform feature extraction on the bone-conducted audio signal through the trained feature extraction model to obtain the voiceprint feature of the bone-conducted audio signal.
[0156] The voiceprint feature of the extracted bone-conducted audio signal V(n,k) can be expressed as DNN1(V(n,k)).
[0157] S15. Input the voiceprint feature, the first audio feature, and the second audio feature of the bone-conducted audio signal into the trained amplitude prediction model to obtain the predicted amplitude.
[0158] Input the voiceprint feature DNN1(V(n,k)), the amplitude MagV(n,k), phase PhaV(n,k) of the bone-conducted audio signal, the amplitude MagM q (n,k), and phase PhaM q (n,k) of the air-conducted audio signal into the trained amplitude prediction model to obtain the predicted amplitude Mag(n,k). This process can be expressed by Equation (7):
[0159] Mag(n,k) = DNN2
M_q (n,k), V(n,k), DNN1(V(n,k))
[0160] S16. Multiply the predicted amplitude by the phase of the unidirectional air-conducted audio signal, and use the multiplication result as the enhanced signal in the time-frequency domain.
[0161] S17. Perform inverse short-time Fourier transform on the enhanced signal to obtain the target audio signal.
[0162] Multiply the predicted amplitude Mag(n,k) by the phase PhaM1(n,k) of the unidirectional air-conducted audio signal. This process can be expressed by Equation (8):
[0163] X(t)=ISTFT(Mag(n,k)*PhaM1(n,k)) (8)
[0164] X(t) is the final enhanced result, i.e., the target audio signal.
[0165] In the above embodiments, audio signals acquired via air conduction have a wide frequency range, while audio signals acquired via bone conduction are almost unaffected by environmental noise. This application designs and trains a dual-model mechanism. A primary model and a secondary model are fused and trained to learn the output of a feature extraction model and the relationship between the audio features of air-conducted and bone-conducted audio signal samples and the amplitude of the air-conducted audio signal samples. The trained primary model can extract voiceprint features from the bone-conducted audio signal. These extracted voiceprint features retain the advantage of low noise from bone conduction, thus aiding in noise reduction. The voiceprint features are fused with the audio features of the air-conducted and bone-conducted audio signals in the trained secondary model. The resulting predicted amplitude can be used to generate an enhanced audio signal. By fusing the advantages of bone conduction and air conduction through the trained dual model, a suitable predicted amplitude is output. The target audio signal obtained based on this predicted amplitude neither loses frequency bands nor reduces noise, thus improving audio quality.
[0166] For the parts of the above-mentioned fusion training method not described in detail, please refer to the relevant description in the audio enhancement method of this application, which will not be repeated here.
[0167] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0168] Based on the same inventive concept, this application also provides an audio enhancement apparatus for implementing the audio enhancement method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio enhancement apparatus embodiments provided below can be found in the limitations of the audio enhancement method described above, and will not be repeated here.
[0169] In an exemplary embodiment, as shown in FIG5, an audio enhancement device is provided, including: a first acquisition module 701, a second acquisition module 702, a feature extraction module 703, an amplitude prediction module 704, and an audio enhancement module 705, wherein:
[0170] The first acquisition module 701 is used to acquire air conduction audio signals and bone conduction audio signals;
[0171] The second acquisition module 702 is used to acquire the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal;
[0172] The feature extraction module 703 is used to extract features from the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal.
[0173] The amplitude prediction module 704 is used to input the voiceprint features, first audio features and second audio features of the bone conduction audio signal into the trained amplitude prediction model to obtain the predicted amplitude.
[0174] Audio enhancement module 705 is used to obtain the target audio signal based on the predicted amplitude.
[0175] In one embodiment, when acquiring air conduction audio signals and bone conduction audio signals, the first acquisition module 701 is further configured to:
[0176] Raw air conduction signals were acquired using an air conduction microphone, and raw bone conduction signals were acquired using a bone conduction microphone.
[0177] The original air conduction signal is subjected to short-time Fourier transform to obtain the frequency domain signal of the original air conduction signal. Based on the original air conduction signal and the frequency domain signal of the original air conduction signal, the air conduction audio signal is obtained.
[0178] A short-time Fourier transform is performed on the original bone conduction signal to obtain its frequency domain signal. Based on the original bone conduction signal and its frequency domain signal, the bone conduction audio signal is obtained.
[0179] In one embodiment, when enhancing audio based on predicted amplitude to obtain a target audio signal, the audio enhancement module 705 is further configured to:
[0180] The predicted amplitude is multiplied by the phase of the air-conducted audio signal, and the result of the multiplication is used as the enhanced signal in the time-frequency domain.
[0181] The enhanced signal is subjected to a short-time inverse Fourier transform to obtain the target audio signal.
[0182] In one embodiment, when acquiring raw air conduction signals based on an air conduction microphone, the first acquisition module 701 is further configured to:
[0183] Multiple raw air conduction signals are obtained by acquiring raw air conduction signals from different directions using multiple air conduction microphones.
[0184] In one embodiment, the plurality of air conduction microphones includes unidirectional air conduction microphones and omnidirectional air conduction microphones. When acquiring multiple original air conduction signals from different directions based on the multiple air conduction microphones, the first acquisition module 701 is further configured to:
[0185] The target direction where the audio signal is greater than a preset decibel threshold is determined, and the audio signal in the target direction is collected in a directional manner based on a unidirectional air conduction microphone to obtain the unidirectional raw air conduction signal.
[0186] The omnidirectional air conduction microphone is used to collect audio signals from all directions to obtain the omnidirectional raw air conduction signal.
[0187] In one embodiment, the air conduction audio signal includes multiple signals. When performing a short-time Fourier transform on the original air conduction signal to obtain its frequency domain signal, based on the original air conduction signal and its frequency domain signal, the first acquisition module 701 is further configured to:
[0188] Short-time Fourier transform is performed on multiple original air conduction signals to obtain the frequency domain signal of each original air conduction signal. Based on each original air conduction signal and its corresponding frequency domain signal, multiple air conduction audio signals are obtained, including a unidirectional air conduction audio signal converted from a unidirectional original air conduction signal.
[0189] Multiplying the predicted amplitude by the phase of the air-conducted audio signal, and using the result as the time-frequency domain augmented signal includes:
[0190] The predicted amplitude is multiplied by the phase of the unidirectional air-conducted audio signal, and the result of the multiplication is used as the enhanced signal in the time-frequency domain.
[0191] In one embodiment, referring to Figure 6, the audio enhancement device further includes a model training module 706, which, before inputting the voiceprint features, first audio features, and second audio features of the bone conduction audio signal into the trained amplitude prediction model to obtain the predicted amplitude, is used to:
[0192] Acquire air conduction audio signal samples and bone conduction audio signal samples;
[0193] Acquire the third audio features of air conduction audio signal samples and the fourth audio features of bone conduction audio signal samples;
[0194] The bone conduction audio signal samples are used as input data for the feature extraction model. The output of the feature extraction model, the third audio feature, and the fourth audio feature are used as input data for the amplitude prediction model. The amplitude of the air conduction audio signal samples is used as the target output of the amplitude prediction model. The feature extraction model and the amplitude prediction model are fused and trained to obtain the trained feature extraction model and the trained amplitude prediction model.
[0195] In one embodiment, when acquiring air conduction audio signal samples and bone conduction audio signal samples, the model training module 706 is further configured to:
[0196] Raw air conduction signal samples were collected using an air conduction microphone and raw bone conduction signal samples were collected using a bone conduction microphone. Specifically, unidirectional raw air conduction signal samples were collected using a unidirectional air conduction microphone and omnidirectional raw air conduction signal samples were collected using an omnidirectional air conduction microphone.
[0197] A short-time Fourier transform is performed on the original air conduction signal sample to obtain the frequency domain signal sample of the original air conduction signal sample. Based on the original air conduction signal sample and the frequency domain signal sample of the original air conduction signal sample, the air conduction audio signal sample is obtained.
[0198] A short-time Fourier transform is performed on the original bone conduction signal sample to obtain the frequency domain signal of the original bone conduction signal sample. Based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample, the bone conduction audio signal sample is obtained.
[0199] The aforementioned audio enhancement device comprises a first acquisition module 701 acquiring air-conduction audio signals and bone-conduction audio signals; a second acquisition module 702 acquiring first audio features of the air-conduction audio signals and second audio features of the bone-conduction audio signals; a feature extraction module 703 extracting features from the bone-conduction audio signals using a trained feature extraction model to obtain voiceprint features of the bone-conduction audio signals; an amplitude prediction module 704 inputting the voiceprint features, first audio features, and second audio features of the bone-conduction audio signals into a trained amplitude prediction model to obtain a predicted amplitude; and an audio enhancement module 705 obtaining the target audio signal based on the predicted amplitude. Audio signals acquired through air conduction have a wide frequency range, while audio signals acquired through bone conduction are almost unaffected by environmental noise. This application utilizes a first-level model to extract voiceprint features from bone-conduction audio signals. The extracted voiceprint features retain the advantage of low noise in bone conduction, thus aiding in noise reduction. The voiceprint features, along with the audio features of the air-conduction and bone-conduction audio signals, are fused in a pre-trained second-level model, and the resulting predicted amplitude can be used to generate the enhanced audio signal. The dual-model mechanism of this application combines the advantages of bone conduction and air conduction, thereby obtaining a target audio signal that does not lose frequency bands and achieves noise reduction, thus improving audio quality.
[0200] In an exemplary embodiment, as shown in FIG7, a fusion training device is provided, including: a fourth acquisition module 801, a fifth acquisition module 802, and a fusion training module 803, wherein:
[0201] The fourth acquisition module 801 is used to acquire air conduction audio signal samples and bone conduction audio signal samples;
[0202] The fifth acquisition module 802 is used to acquire the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample;
[0203] The fusion training module 803 is used to use bone conduction audio signal samples as input data for the feature extraction model, the output of the feature extraction model, the third audio feature and the fourth audio feature as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model. The feature extraction model and the amplitude prediction model are fused and trained to obtain the trained feature extraction model and the trained amplitude prediction model.
[0204] In one embodiment, when acquiring air conduction audio signal samples and bone conduction audio signal samples, the fourth acquisition module 801 is further configured to:
[0205] Raw air conduction signal samples were collected using an air conduction microphone and raw bone conduction signal samples were collected using a bone conduction microphone. Specifically, unidirectional raw air conduction signal samples were collected using a unidirectional air conduction microphone and omnidirectional raw air conduction signal samples were collected using an omnidirectional air conduction microphone.
[0206] A short-time Fourier transform is performed on the original air conduction signal sample to obtain the frequency domain signal sample of the original air conduction signal sample. Based on the original air conduction signal sample and the frequency domain signal sample of the original air conduction signal sample, the air conduction audio signal sample is obtained.
[0207] A short-time Fourier transform is performed on the original bone conduction signal sample to obtain the frequency domain signal of the original bone conduction signal sample. Based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample, the bone conduction audio signal sample is obtained.
[0208] The aforementioned fusion training device comprises a fourth acquisition module 801 acquiring air conduction audio signal samples and bone conduction audio signal samples; a fifth acquisition module 802 acquiring the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; and a fusion training module 803 using the bone conduction audio signal samples as input data for the feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model, thereby fusing and training the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model. Audio signals acquired through air conduction have a wide frequency range, while audio signals acquired through bone conduction are almost unaffected by environmental noise. This application designs and trains a dual-model mechanism. The first-level model can extract voiceprint features from the bone conduction audio signal, and the extracted voiceprint features continue the advantage of low noise in bone conduction, thus assisting in noise reduction. The second-level model can fuse the voiceprint features output by the first-level model with the audio features of the air conduction audio signal and the bone conduction audio signal. The primary and secondary models are trained together to learn the output of the feature extraction model and the relationship between the audio features of air-conducted and bone-conducted audio signal samples and the amplitude of the air-conducted audio signal samples. The trained dual model can be used to predict the desired amplitude of audio in audio enhancement, thereby achieving audio enhancement and improving audio quality.
[0209] The modules in the aforementioned audio enhancement device and fusion training device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0210] Terms such as “component,” “module,” and “system” are intended to refer to computer-related entities, which can be hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, executable code, a thread of execution, a program, and / or a computer. For illustration, a running program on a server and the server itself can both be components. One or more components may reside within a process and / or a thread of execution, and components may be located within a single computer and / or distributed across two or more computers.
[0211] In an exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram is shown in Figure 8. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile computer-readable storage medium and internal memory. The non-volatile computer-readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile computer-readable storage medium. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement an audio enhancement method.
[0212] In an exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as shown in Figure 9. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile computer-readable storage medium and internal memory. The non-volatile computer-readable storage medium stores an operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile computer-readable storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer-readable instructions are executed by the processor, they implement an audio enhancement method. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0213] Those skilled in the art will understand that the structures shown in Figures 8 and 9 are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.
[0214] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps in the above method embodiments.
[0215] In one embodiment, a computer-readable storage medium is provided storing computer-readable instructions that, when executed by a processor, implement the steps in the above method embodiments.
[0216] In one embodiment, a computer program product is provided, the computer program product including computer-readable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer-readable instructions from the computer-readable storage medium, and executes the computer-readable instructions, causing the computer device to perform the steps in the above-described method embodiments.
[0217] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0218] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0219] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An audio enhancement method, characterized in that, include: Acquire air conduction audio signals and bone conduction audio signals; acquire a first audio feature of the air conduction audio signal and a second audio feature of the bone conduction audio signal; The bone conduction audio signal is subjected to feature extraction using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal; the voiceprint features of the bone conduction audio signal, the first audio feature, and the second audio feature are input into a trained amplitude prediction model to obtain the predicted amplitude; the target audio signal is obtained based on the predicted amplitude.
2. The method according to claim 1, characterized in that, The acquisition of air conduction audio signals and bone conduction audio signals includes: acquiring raw air conduction signals based on an air conduction microphone, and acquiring raw bone conduction signals based on a bone conduction microphone; performing a short-time Fourier transform on the raw air conduction signals to obtain the frequency domain signal of the raw air conduction signals, and obtaining air conduction audio signals based on the raw air conduction signals and the frequency domain signal of the raw air conduction signals; performing a short-time Fourier transform on the raw bone conduction signals to obtain the frequency domain signal of the raw bone conduction signals, and obtaining bone conduction audio signals based on the raw bone conduction signals and the frequency domain signal of the raw bone conduction signals.
3. The method according to claim 2, characterized in that, The step of obtaining the target audio signal based on the predicted amplitude includes: multiplying the predicted amplitude by the phase of the air-conducted audio signal, using the result of the multiplication as an enhanced signal in the time-frequency domain; and performing a short-time inverse Fourier transform on the enhanced signal to obtain the target audio signal.
4. The method according to claim 3, characterized in that, The method of acquiring raw air conduction signals based on air conduction microphones includes: acquiring raw air conduction signals from different directions using multiple air conduction microphones to obtain multiple raw air conduction signals.
5. The method according to claim 4, characterized in that, The plurality of air conduction microphones includes unidirectional air conduction microphones and omnidirectional air conduction microphones. The process of acquiring multiple original air conduction signals from different directions based on the plurality of air conduction microphones includes: determining a target direction in which the audio signal is greater than a preset decibel threshold, and acquiring the audio signal in the target direction based on the unidirectional air conduction microphone to obtain a unidirectional original air conduction signal; and acquiring the audio signal in all directions based on the omnidirectional air conduction microphone to obtain an omnidirectional original air conduction signal.
6. The method according to claim 5, characterized in that, The air conduction audio signal includes multiple signals. The step of performing a short-time Fourier transform on the original air conduction signal to obtain its frequency domain signal, and obtaining the air conduction audio signal based on the original air conduction signal and its frequency domain signal, includes: performing a short-time Fourier transform on the multiple original air conduction signals to obtain the frequency domain signal of each original air conduction signal; and obtaining multiple air conduction audio signals based on each original air conduction signal and its corresponding frequency domain signal. The multiple air conduction audio signals include a unidirectional air conduction audio signal converted from the unidirectional original air conduction signal. The step of multiplying the predicted amplitude with the phase of the air conduction audio signal and using the result as the time-frequency domain enhancement signal includes: multiplying the predicted amplitude with the phase of the unidirectional air conduction audio signal and using the result as the time-frequency domain enhancement signal.
7. The method according to any one of claims 1-6, characterized in that, Before inputting the voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal into the trained amplitude prediction model to obtain the predicted amplitude, the method further includes: acquiring air conduction audio signal samples and bone conduction audio signal samples; acquiring the third audio feature of the air conduction audio signal samples and the fourth audio feature of the bone conduction audio signal samples; using the bone conduction audio signal samples as input data for the feature extraction model, using the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for the amplitude prediction model, using the amplitude of the air conduction audio signal samples as the target output of the amplitude prediction model, and performing fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.
8. The method according to claim 7, characterized in that, The acquisition of air conduction audio signal samples and bone conduction audio signal samples includes: acquiring original air conduction signal samples based on an air conduction microphone, and acquiring original bone conduction signal samples based on a bone conduction microphone, wherein unidirectional original air conduction signal samples are acquired directionally based on a unidirectional air conduction microphone, and omnidirectional original air conduction signal samples are acquired based on an omnidirectional air conduction microphone; performing a short-time Fourier transform on the original air conduction signal samples to obtain frequency domain signal samples of the original air conduction signal samples, and obtaining air conduction audio signal samples based on the original air conduction signal samples and the frequency domain signal samples of the original air conduction signal samples; performing a short-time Fourier transform on the original bone conduction signal samples to obtain frequency domain signal samples of the original bone conduction signal samples, and obtaining bone conduction audio signal samples based on the original bone conduction signal samples and the frequency domain signal samples of the original bone conduction signal samples.
9. An audio enhancement device, characterized in that, include: The first acquisition module is used to acquire air conduction audio signals and bone conduction audio signals; The second acquisition module is used to acquire the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal; feature The extraction module is used to extract features from the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal. The amplitude prediction module is used to input the voiceprint features of the bone conduction audio signal, the first audio feature, and the second audio feature into the trained amplitude prediction model to obtain the predicted amplitude. An audio enhancement module is used to obtain a target audio signal based on the predicted amplitude.
10. An earphone, comprising an air conduction microphone, a bone conduction microphone, a memory, and a processor, wherein the air conduction microphone is used to acquire air conduction signals, the bone conduction microphone is used to acquire bone conduction signals, and the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.