Audio enhancement method, apparatus, and earphones

The audio enhancement method uses trained models to process air and bone conduction signals, addressing noise reduction and quality issues in low signal-to-noise environments by reconstructing enhanced audio signals, maintaining frequency integrity.

JP2026079754APending Publication Date: 2026-05-15ANKER INNOVATIONS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ANKER INNOVATIONS TECH CO LTD
Filing Date
2025-10-23
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing audio enhancement technologies struggle to effectively reduce noise and improve audio quality in environments with low signal-to-noise ratios and interfering human voices, especially when using single microphones or conventional methods.

Method used

An audio enhancement method that combines air conduction and bone conduction signals using trained feature extraction and amplitude prediction models to generate a predicted amplitude, which is used to reconstruct an enhanced audio signal, leveraging the low noise characteristics of bone conduction signals and the wide frequency range of air conduction signals.

Benefits of technology

The dual-model mechanism achieves noise reduction without losing frequency bandwidth, resulting in improved audio quality by combining the advantages of both air and bone conduction signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026079754000001_ABST
    Figure 2026079754000001_ABST
Patent Text Reader

Abstract

This invention provides an audio enhancement method, apparatus, and earphones that improve audio quality. [Solution] A method that combines the advantages of bone conduction and air conduction, and in which the obtained target audio signal achieves noise reduction without losing frequency bandwidth and improves audio quality, comprising: acquiring an air conduction audio signal and a bone conduction audio signal; acquiring a first audio feature of the air conduction audio signal and a second audio feature of the bone conduction audio signal; performing feature extraction on the bone conduction audio signal using a trained feature extraction model to acquire the voiceprint feature of the bone conduction audio signal; inputting the voiceprint feature, first audio feature and second audio feature of the bone conduction audio signal into a trained amplitude prediction model to acquire a predicted amplitude; and enhancing the audio based on the predicted amplitude to acquire a target audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to the technical field of audio processing, and more specifically, to an audio enhancement method, apparatus, and earphones. [Background technology]

[0002] This section is intended to provide background or context to the embodiments of the present invention described in the claims and embodiments for carrying out the invention. The descriptions herein are not considered prior art as they are included in this section.

[0003] With the widespread use of mobile communication devices, people can make calls or recordings anytime, anywhere. However, ambient noise and interfering voices during calls or recordings can impair the clarity of the audio signal collected by the device, affecting the quality of the call or recording. Therefore, it is necessary to reduce noise in the audio, eliminate noise and interference during calls or recordings, and improve the audio quality. [Overview of the project]

[0004] In light of the above, there is a need to provide an audio enhancement method, apparatus, and earphones that can improve audio quality in response to the aforementioned technical problems.

[0005] In the first aspect, the present application is: To acquire air conduction audio signals and bone conduction audio signals, To acquire the first audio features of the air conduction audio signal and the second audio features of the bone conduction audio signal, The process involves performing feature extraction on the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal, The voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features are input to a trained amplitude prediction model to obtain a predicted amplitude. The present invention provides an audio enhancement method that includes acquiring a target audio signal based on the predicted amplitude.

[0006] In the second aspect, the present application is: A first acquisition module that acquires air conduction audio signals and bone conduction audio signals, A second acquisition module for acquiring the first audio features of the air conduction audio signal and the second audio features of the bone conduction audio signal, A feature extraction module that performs feature extraction on the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal, An amplitude prediction module that inputs the voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features into a trained amplitude prediction model to obtain a predicted amplitude, The present invention provides an audio enhancement device that includes an audio enhancement module that acquires a target audio signal based on the predicted amplitude.

[0007] In a third aspect, the present invention relates to an earphone comprising an air conduction microphone for collecting air conduction signals, a bone conduction microphone for collecting bone conduction signals, a memory storing a computer program, and a processor, wherein when the processor executes the computer program, Steps to acquire air conduction audio signals and bone conduction audio signals, The steps include acquiring a first audio feature of the air conduction audio signal and a second audio feature of the bone conduction audio signal, The steps include: performing feature extraction on the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal; The steps include inputting the voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features into a trained amplitude prediction model to obtain a predicted amplitude, The present invention provides earphones that enable the steps of acquiring a target audio signal based on the predicted amplitude.

[0008] The above audio enhancement method, apparatus, and earphones acquire an air conduction audio signal and a bone conduction audio signal, acquire a first audio feature of the air conduction audio signal and a second audio feature of the bone conduction audio signal, perform feature extraction on the bone conduction audio signal using a trained feature extraction model to acquire the voiceprint features of the bone conduction audio signal, input the voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features into a trained amplitude prediction model to acquire a predicted amplitude, and acquire a target audio signal based on the predicted amplitude.

[0009] Audio signals collected by air conduction have the characteristic of a wide frequency range, while audio signals collected by bone conduction are hardly affected by ambient noise interference. In this application, a primary model is used to extract voiceprint features from the bone conduction audio signal. The extracted voiceprint features inherit the advantage of low noise in bone conduction and can assist in noise reduction. The voiceprint features and the audio features of the air conduction audio signal and the bone conduction audio signal are fused in a pre-trained secondary model, and the resulting predicted amplitude can be used to generate an enhanced audio signal. The dual-model mechanism of this application combines the advantages of bone conduction and air conduction, and the resulting target audio signal achieves noise reduction without losing frequency bandwidth and improves audio quality.

[0010] To more clearly explain the technical solutions in the embodiments of this application or related technologies, the drawings necessary for describing the embodiments of this application or related technologies are briefly described below. Clearly, the drawings described below are only a few embodiments of this application, and those skilled in the art can obtain other relevant drawings based on these drawings without any creative work. [Brief explanation of the drawing]

[0011] [Figure 1] It is an application environment diagram of an audio enhancement method and a fusion training method in an embodiment. [Figure 2] It is a schematic diagram showing the flow of an audio enhancement method in an embodiment. [Figure 3] It is a schematic diagram showing the flow of a fusion training method in an embodiment. [Figure 4] It is a schematic diagram showing the flow of an audio enhancement method and a fusion training method in another embodiment. [Figure 5] It is a structural block diagram of an audio enhancement device in an embodiment. [Figure 6] It is another structural block diagram of an audio enhancement device in an embodiment. [Figure 7] It is a structural block diagram of a fusion training device in an embodiment. [Figure 8] It is an internal configuration diagram of a computer device in an embodiment. [Figure 9] It is an internal configuration diagram of a computer device in another embodiment.

Embodiments for Carrying Out the Invention

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following will further elaborate on this application while referring to the drawings and embodiments. It should be noted that the specific embodiments described here are merely for explaining this application and are not used in the description of this application. When "first" and "second" are described, they are merely for distinguishing technical features and should not be understood as indicating or implying relative importance, implicitly or explicitly indicating the number of the indicated technical features, or implicitly or explicitly indicating the sequence relationship of the indicated technical features.

[0013] In the description of this application, unless there are clear limitations, words such as installation, attachment, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meaning of these words in this application by combining the specific content of the technical solutions.

[0014] Audio enhancement, also called audio noise reduction, refers to the process of removing the noise components in an audio signal and retaining the desired audio signal. The methods of audio enhancement can be divided into the conventional audio enhancement algorithms and the audio enhancement based on neural networks. The audio enhancement based on neural networks has advantages in performance, especially for unexpected noise types such as burst noise. Due to the complexity of the noise destruction process, the audio enhancement based on neural networks shows obvious advantages in environments such as low signal-to-noise ratio and non-stationary noise.

[0015] In the related field of earphone audio enhancement, the audio enhancement technology based on neural networks combines the conventional beamforming technology and the single-channel audio enhancement technology to improve the quality and intelligibility of calls and can obtain excellent effects in some scenes. However, when the signal-to-noise ratio is extremely low and there is interfering human voice, the effect becomes poor.

[0016] In some cases, the earphone multi-microphone audio enhancement technology can be used to deal with the human voice interference in audio, but the effect is still poor in scenes with a low signal-to-noise ratio. In addition, the multi-microphone audio enhancement technology cannot be applied in the case of only a single microphone.

[0017] With the research and development of VPU (bone conduction microphones), the high signal-to-noise ratio of VPU signals at low frequencies (e.g., below 1 kHz) has played a significant role in the development of audio enhancement technology for wearable devices such as earphones. Audio enhancement based on VPU signals still exhibits good effects in scenes with extremely low signal-to-noise ratios and scenes with interfering human voices, ensuring the quality of voice communication. Taking advantage of the high signal-to-noise ratio of VPU signals at low frequencies, low-frequency VPU signals and high-frequency earphone talk signals can be spliced ​​and then processed by a single-channel audio enhancement network. However, such simple splicing of VPU signals and talk signals still has several problems, such as poor rejection of interfering human voices, the presence of a "frequency band break" phenomenon in the processing effect which affects the perceived sound, and poor recovery effect for high frequencies.

[0018] To solve the above problems, the present invention provides an audio enhancement method and a fusion training method. Referring to Figure 1, Figure 1 is a diagram illustrating the application environment of the audio enhancement method and the fusion training method in one embodiment. The audio enhancement method and the fusion training method according to the embodiment of the present invention may be applied to an application environment as shown in Figure 1. Terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system may be integrated with server 104, or it may be located in the cloud or on another server.

[0019] The terminal and the server may each independently execute the audio enhancement method and the fusion training method according to the embodiment of the present application.

[0020] For example, the terminal acquires air conduction audio signal samples and bone conduction audio signal samples, acquires a third audio feature of the air conduction audio signal sample and a fourth audio feature of the bone conduction audio signal sample, uses the bone conduction audio signal sample as input data for a feature extraction model, uses the output of the feature extraction model, the third audio feature and the fourth audio feature as input data for an amplitude prediction model, uses the amplitude of the air conduction audio signal sample as the target output for the amplitude prediction model, performs fusion training on the feature extraction model and the amplitude prediction model, and obtains a trained feature extraction model and a trained amplitude prediction model.

[0021] The terminal acquires air-conducted audio signals and bone-conducted audio signals, and acquires the first audio feature of the air-conducted audio signal and the second audio feature of the bone-conducted audio signal. The terminal performs feature extraction on the bone-conducted audio signal using a trained feature extraction model to acquire the voiceprint features of the bone-conducted audio signal. The terminal inputs the voiceprint features, first audio features, and second audio features of the bone-conducted audio signal into a trained amplitude prediction model to acquire the predicted amplitude, and acquires the target audio signal based on the predicted amplitude.

[0022] Furthermore, the terminal and server may work together to execute the audio enhancement method and the fusion training method according to the embodiment of the present invention.

[0023] For example, the server acquires air conduction audio signal samples and bone conduction audio signal samples, and acquires the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample. The server uses the bone conduction audio signal sample as input data for a feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for an amplitude prediction model, and the amplitude of the air conduction audio signal sample as the target output of the amplitude prediction model. It then performs fusion training on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model.

[0024] The server provides the terminal with an interface to invoke the model. The terminal acquires air-conducted audio signals and bone-conducted audio signals, and acquires the first audio feature of the air-conducted audio signal and the second audio feature of the bone-conducted audio signal. The terminal performs feature extraction on the bone-conducted audio signal using a trained feature extraction model to acquire the voiceprint feature of the bone-conducted audio signal. The terminal inputs the voiceprint feature, first audio feature, and second audio feature of the bone-conducted audio signal into a trained amplitude prediction model to acquire the predicted amplitude, and then acquires the target audio signal based on the predicted amplitude.

[0025] The terminal and the server may work together to execute the audio enhancement method according to the embodiment of the present invention, and may also work together to execute the fusion training method according to the embodiment of the present invention.

[0026] For example, the terminal acquires air conduction audio signal samples and bone conduction audio signal samples. The terminal acquires the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample, and sends the bone conduction audio signal sample, the third audio feature, and the fourth audio feature to the server. The server uses the bone conduction audio signal sample as input data for a feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for an amplitude prediction model, and the amplitude of the air conduction audio signal sample as the target output for the amplitude prediction model. It then performs fusion training on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.

[0027] The terminal acquires air-conducted audio signals and bone-conducted audio signals. The terminal acquires the first audio feature of the air-conducted audio signal and the second audio feature of the bone-conducted audio signal, and transmits the bone-conducted audio signal, the first audio feature, and the second audio feature to the server. The server performs feature extraction on the bone-conducted audio signal using a trained feature extraction model to acquire the voiceprint feature of the bone-conducted audio signal. The server inputs the voiceprint feature, the first audio feature, and the second audio feature of the bone-conducted audio signal into a trained amplitude prediction model to acquire the predicted amplitude, acquires the target audio signal based on the predicted amplitude, and transmits the target audio signal to the terminal.

[0028] Terminal 102 may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, Internet of Things device, or portable wearable device. The Internet of Things device may be a smart speaker, smart TV, smart air conditioner, smart in-car device, projection device, etc. The portable wearable device may be a smartwatch, smart bracelet, headset, etc. The headset may be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc.

[0029] Server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.

[0030] The terminal 102 and the server 104 may be connected by a communication connection method such as Bluetooth, USB (Universal Serial Bus), or a network, but the present invention is not limited thereto.

[0031] In one exemplary embodiment, an audio enhancement method is provided, as shown in Figure 2, and the method is described as being applied to the terminal in Figure 1, and includes the following steps S202 to S210.

[0032] In step S202, air conduction audio signals and bone conduction audio signals are acquired.

[0033] The air-conducted audio signals and bone-conducted audio signals obtained in this application refer to air-conducted audio signals and bone-conducted audio signals obtained based on the same audio. The air-conducted audio signals and bone-conducted audio signals of the same audio include the same or similar audio content, and may, for example, be audio signals obtained for the same scene at the same time.

[0034] The computer equipment may be a wearable device, and the wearable device may be equipped with an air conduction microphone and a bone conduction microphone, the bone conduction microphone being placed in close contact with the user at the wearing position of the wearable device. When the wearer speaks, the sound is transmitted to the bone conduction microphone by the bone conduction method and collected by the bone conduction microphone as an original bone conduction signal.

[0035] For example, in a voice communication scenario, when a wearer makes a sound while putting on the wearable device, the wearable device collects the wearer's voice signal as an original bone conduction signal using a bone conduction microphone and the wearer's voice signal as an original air conduction signal using an air conduction microphone. Selectively, the air conduction microphone may simultaneously collect the wearer's voice signal and the ambient sound signal of the environment in which the wearer is located, and the simultaneously collected wearer's voice signal and ambient sound signal of the environment in which the wearer is located may be jointly used as the original air conduction signal.

[0036] The air conduction audio signal may be the original air conduction signal collected by an air conduction microphone, or an audio signal obtained by preprocessing the original air conduction signal collected by an air conduction microphone. The bone conduction audio signal may be the original bone conduction signal collected by a bone conduction microphone, or an audio signal obtained by preprocessing the original bone conduction signal collected by a bone conduction microphone. Preprocessing may include, but is not limited to, feature transformation, signal cutting, and signal alignment.

[0037] For example, an original air conduction signal may be collected using an air conduction microphone, a feature transformation may be performed on the original air conduction signal to obtain the frequency domain signal of the original air conduction signal, and an air conduction audio signal may be obtained based on the original air conduction signal and the frequency domain signal of the original air conduction signal. Alternatively, an original bone conduction signal may be collected using a bone conduction microphone, a feature transformation may be performed on the original bone conduction signal to obtain the frequency domain signal of the original bone conduction signal, and a bone conduction audio signal may be obtained based on the original bone conduction signal and the frequency domain signal of the original bone conduction signal. A feature transformation is a transformation that can transform a signal from the time domain to the frequency domain, and includes, but is not limited to, the short-time Fourier transform, wavelet transform, continuous wavelet transform, and S-transform.

[0038] This application does not limit the number of air conduction microphones and bone conduction microphones, and there may be one or more original air conduction signals and original bone conduction signals collected, and correspondingly there may be one or more bone conduction audio signals and one or more air conduction audio signals. For example, a device may be equipped with multiple air conduction microphones and one bone conduction microphone, multiple original air conduction signals may be collected by the multiple air conduction microphones, multiple air conduction audio signals may be obtained by preprocessing each of the multiple original air conduction signals, one original bone conduction signal may be collected by one bone conduction microphone, and multiple bone conduction audio signals may be obtained by preprocessing the one bone conduction signal.

[0039] In step S204, the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal are acquired.

[0040] Audio features are features that contain audio information and can reflect the characteristics of an audio signal. Audio features of an audio signal may include amplitude and phase, and may also include other features that can reflect the characteristics of the audio signal, such as duration and source.

[0041] Note that the terms "first" and "second" in the first and second audio features are used to distinguish the source of the audio features, i.e., they are from air conduction audio signals and bone conduction audio signals, respectively. The first and second audio features may be the same type of audio signal obtained by performing the same processing on air conduction audio signals and bone conduction audio signals.

[0042] In step S206, a trained feature extraction model is used to extract features from the bone conduction audio signal and obtain the voiceprint features of the bone conduction audio signal.

[0043] Since bone conduction audio signals originate directly or indirectly from the original bone conduction signal collected by a bone conduction microphone, the original bone conduction signal contains less noise and the corresponding bone conduction audio signal has less interference due to the bone conduction sound propagation method. In order to achieve the strongest possible noise reduction effect, this invention first performs feature processing on the bone conduction audio signal to obtain voiceprint features of the bone conduction audio signal, and these voiceprint features can play a role in assisting noise reduction during subsequent feature fusion.

[0044] To achieve feature extraction from bone conduction audio signals, this invention pre-trains a feature extraction model. This feature extraction model may be constructed based on a neural network and may include multiple layers, such as an input layer, output layer, hidden layer, and normalization layer, with each layer connected to the others by weights and activation functions. The feature extraction model is trained using bone conduction audio signal samples and corresponding voiceprint feature tags as a training set, allowing the feature extraction model to learn how to extract voiceprint features from bone conduction audio signals. The trained feature extraction model receives bone conduction audio signals as input data and performs feature extraction on the bone conduction audio signals to obtain voiceprint features from the bone conduction audio signals.

[0045] In one embodiment, if the neural network is a convolutional neural network, the feature extraction model may further include a convolutional layer, for example, which performs a convolutional operation on the received data to extract the voiceprint features of the bone conduction audio signal.

[0046] In step S208, the voiceprint features, first audio features, and second audio features of the bone conduction audio signal are input into a trained amplitude prediction model to obtain the predicted amplitude.

[0047] In addition to the feature extraction model, the present invention includes a pre-trained amplitude prediction model, which is used to obtain a predicted amplitude based on the voiceprint features of the input bone conduction audio signal and the audio features of the bone conduction audio signal and air conduction audio signal. The predicted amplitude includes amplitude information of the desired enhancement result of the collected sound.

[0048] In order to obtain predicted amplitude based on the input data of the voiceprint features of the input bone conduction audio signal and the audio features of the bone conduction audio signal and air conduction audio signal, this invention pre-constructs an amplitude prediction model based on a neural network and trains the amplitude prediction model to learn the correspondence between these input data and predicted amplitude values ​​using the learning ability of the neural network.

[0049] The amplitude prediction model may include multiple layers, such as an input layer, an output layer, a hidden layer, and a normalization layer, and each layer is connected to the others by weights and activation functions. The amplitude prediction model is trained using the voiceprint features of bone conduction audio signal samples and the audio features of bone conduction audio signal samples and air conduction audio signal samples as input data, and the amplitude of the audio signal sample or air conduction audio signal sample as target data. When a predetermined training termination condition is reached, the training is terminated and the trained amplitude prediction model is obtained.

[0050] By inputting the voiceprint features, first audio features, and second audio features of a bone conduction audio signal into a trained amplitude prediction model, a predicted amplitude can be obtained. The predicted amplitude value output will differ depending on the target data used during training of the amplitude prediction model. By selecting different target data during training, the amplitude prediction model can appropriately output a predicted amplitude for enhancing either a bone conduction audio signal or an air conduction audio signal.

[0051] For example, if the amplitude of a bone conduction audio signal sample is used as the target data, the amplitude prediction model learns the relationship between the input data and the amplitude of the bone conduction audio signal sample during training. After completing training and obtaining a trained amplitude prediction model, the voiceprint features, first audio features, and second audio features of the bone conduction audio signal are input to the trained amplitude prediction model, and the resulting predicted amplitude value is the predicted amplitude for enhancing the bone conduction audio signal.

[0052] When the amplitude of an air-conducted audio signal sample is used as the target data, the amplitude prediction model learns the relationship between the input data and the amplitude of the air-conducted audio signal sample during training. After completing training and obtaining a trained amplitude prediction model, the voiceprint features, first audio features, and second audio features of the air-conducted audio signal are input to the trained amplitude prediction model, and the resulting predicted amplitude value is the predicted amplitude for enhancing the air-conducted audio signal.

[0053] In step S210, the target audio signal is acquired based on the predicted amplitude.

[0054] In signal processing, amplitude and phase are two fundamental attributes that describe a signal wave. Amplitude refers to the maximum distance the signal wave deviates from a reference value, while phase refers to the position of the signal wave relative to a reference point at a given time. The product of amplitude and phase is a negative representation of the signal wave. After obtaining a predicted amplitude, the audio signal wave can be reconstructed by multiplying the predicted amplitude by a specific phase, and the method of reconstruction can be used to achieve audio enhancement for the corresponding audio signal.

[0055] Based on the audio enhancement demand for different enhancement targets, the predicted amplitude can be multiplied by different phases to reconstruct the signal wave of the enhancement target as the enhancement result for the audio signal of the enhancement target, and the enhancement result is the target audio signal.

[0056] For example, the enhancement target may be a bone conduction-related audio signal, such as an original bone conduction signal or a bone conduction audio signal. In this case, during the training phase, the amplitude prediction model is trained using the phase of a bone conduction audio signal sample as target data. When obtaining the target audio signal based on the predicted amplitude, the predicted amplitude is multiplied by the phase of the bone conduction audio signal, and the resulting target audio signal is the enhancement result of the bone conduction-related audio signal.

[0057] For example, the enhancement target may be an air conduction-related audio signal, such as the original air conduction signal or an air conduction audio signal. In this case, during the training phase, the amplitude prediction model is trained using the phase of an air conduction audio signal sample as the target data. When obtaining the target audio signal based on the predicted amplitude, the predicted amplitude is multiplied by the phase of the air conduction audio signal, and the resulting target audio signal is the enhancement result of the air conduction-related audio signal.

[0058] In the above audio enhancement method, air-conducted audio signals and bone-conducted audio signals are acquired, and a first audio feature of the air-conducted audio signal and a second audio feature of the bone-conducted audio signal are acquired. Feature extraction is performed on the bone-conducted audio signal using a trained feature extraction model to acquire the voiceprint features of the bone-conducted audio signal. The voiceprint features, first audio features, and second audio features of the bone-conducted audio signal are input to a trained amplitude prediction model to acquire the predicted amplitude, and the target audio signal is acquired based on the predicted amplitude. Audio signals collected by air conduction have the characteristic of a wide frequency range, and audio signals collected by bone conduction are hardly affected by ambient noise interference. In this application, voiceprint features are extracted from the bone-conducted audio signal using a primary model, and the extracted voiceprint features can assist in noise reduction, inheriting the advantage of low bone conduction noise. The voiceprint features and the audio features of the air-conducted audio signal and bone-conducted audio signal are fused in a pre-trained secondary model, and the resulting predicted amplitude can be used to generate an enhanced audio signal. The dual-model mechanism of this invention combines the advantages of bone conduction and air conduction, resulting in a target audio signal that achieves noise reduction without losing frequency bandwidth and improves audio quality.

[0059] In one embodiment, air conduction audio signals and bone conduction audio signals are obtained. This includes collecting an original air conduction signal using an air conduction microphone and an original bone conduction signal using a bone conduction microphone; performing a short-time Fourier transform on the original air conduction signal to obtain the frequency domain signal of the original air conduction signal; and obtaining an air conduction audio signal based on the original air conduction signal and the frequency domain signal of the original air conduction signal; and performing a short-time Fourier transform on the original bone conduction signal to obtain the frequency domain signal of the original bone conduction signal; and obtaining a bone conduction audio signal based on the original bone conduction signal and the frequency domain signal of the original bone conduction signal.

[0060] The original air conduction signal and original bone conduction signal are collected synchronously using an air conduction microphone and a bone conduction microphone, and the collected original air conduction signal and original bone conduction signal are signals in the time domain. For example, when a user makes a call, the air conduction microphone collects the voice signal of the user making the call in real time to obtain the original air conduction signal, and the bone conduction microphone collects the voice signal of the user making the call in real time to obtain the original bone conduction signal.

[0061] To introduce a conversion in the frequency domain, the original air conduction signal and the original bone conduction signal are converted to the frequency domain using a short-time Fourier transform, thereby obtaining the frequency domain signals of the original air conduction signal and the original bone conduction signal. The original air conduction signal and the original bone conduction signal are time-domain signals. By combining the original air conduction signal and the original air conduction signal's frequency domain signals, a time-frequency domain air conduction audio signal can be obtained. Similarly, by combining the original bone conduction signal and the original bone conduction signal's frequency domain signals, a time-frequency domain bone conduction audio signal can be obtained. This allows the original air conduction signal and the original bone conduction signal to be converted to the time-frequency domain so that frequency domain information can be introduced during subsequent model processing.

[0062] In the above embodiment, the original bone conduction signal and air conduction signal in the time domain can be converted into bone conduction audio signal and air conduction audio signal in the time-frequency domain, which contain more information, by using the short-time Fourier transform before inputting them into the model. In this way, the model can learn the frequency domain information of the bone conduction audio signal and air conduction audio signal and output more accurate prediction results.

[0063] In one embodiment, acquiring a target audio signal based on predicted amplitude is possible. This method includes multiplying the predicted amplitude by the phase of the air-conducted audio signal, using the result as the time-frequency domain enhancement signal, and performing a short-time inverse Fourier transform on the enhancement signal to obtain the target audio signal.

[0064] When using feature transformation to convert the original air conduction signal and original bone conduction signal to the time-frequency domain, an enhancement signal in the time-frequency domain is obtained, and then the enhancement signal is processed using the inverse transformation of the feature transformation to convert the time-frequency domain enhancement signal back to the time domain to obtain the target audio signal that can be heard directly.

[0065] If the feature transformation method is the Short-Time Fourier Transform, after obtaining the enhancement signal in the time-frequency domain, the Short-Time Inverse Fourier Transform is performed on the enhancement signal to obtain the target audio signal.

[0066] In the above embodiment, bone conduction and air conduction signals are converted back and forth between the time domain and the time-frequency domain by short-time Fourier transform and inverse short-time Fourier transform, and after the model obtains the desired predicted amplitude, an inverse transform from the time-frequency domain to the time domain is performed to obtain a clear target audio signal.

[0067] In one embodiment, there may be multiple air conduction microphones, and when collecting the original air conduction signal using the air conduction microphones, the original air conduction signal may be collected from different directions using the multiple air conduction microphones to obtain multiple original air conduction signals.

[0068] This invention does not limit the number of air conduction microphones. If there are multiple air conduction microphones, each air conduction microphone can be mounted in a different orientation, and the original air conduction signal can be collected from different directions to obtain multiple original air conduction signals from different directions.

[0069] Using earphones as an example, multiple air conduction microphones may be attached to the same earphone in different orientations (e.g., front, back, left, right, up, down). Depending on the orientation at which they are attached, the multiple air conduction microphones can collect audio signals from multiple directions of the earphone, thereby obtaining multiple original air conduction signals. Due to the differences in collection direction, the audio, intensity, and direction contained in the multiple original air conduction signals will differ, enriching the amount of audio information contained in the original air conduction signals. In some cases, it becomes easier to more accurately position each of these audio signals, further improving audio quality and facilitating audio processing such as stereo reduction.

[0070] In one embodiment, the multiple air conduction microphones include unidirectional air conduction microphones and omnidirectional air conduction microphones, and multiple original air conduction signals are obtained by collecting the original air conduction signals from different directions using the multiple air conduction microphones. This includes determining a target direction in which the audio signal is greater than a predetermined decibel threshold, directionally collecting the audio signal in the target direction using a unidirectional air conduction microphone to obtain a unidirectional original air conduction signal, and collecting audio signals from all directions using an omnidirectional air conduction microphone to obtain an omnidirectional original air conduction signal.

[0071] In the embodiments of this invention, the multiple air conduction microphones installed in the device may include unidirectional air conduction microphones and omnidirectional air conduction microphones. The omnidirectional air conduction microphones collect audio signals from all directions to obtain an omnidirectional original air conduction signal. For example, when a user makes a call, the bone conduction microphone collects the user's audio signal to obtain an original bone conduction signal, the unidirectional air conduction microphones directionally collect audio signals from the direction in which the user is located to obtain a unidirectional original air conduction signal, and the omnidirectional original air conduction microphones collect audio signals from all directions in the environment in which the user is located to obtain an omnidirectional original air conduction microphone signal. The audio signals from all directions in the environment in which the user is located may include not only the user's audio signal but also background sound signals from the environment in which the user is located.

[0072] When collecting audio signals from one or more directions using a unidirectional air conduction microphone, the target direction where the speaker (e.g., user) is located can be determined based on the volume of the sound. For example, by pre-setting a decibel threshold and continuously monitoring audio signals from different directions, if an audio signal louder than the predetermined decibel threshold is detected, the target direction where the audio signal louder than the predetermined decibel threshold is located can be determined, and the audio signal from the speaker can be acquired as a unidirectional original air conduction signal by directionally collecting the audio signal in the target direction using a unidirectional air conduction microphone.

[0073] In some cases, it may not be necessary to set a specific target for collection. Audio signals from different directions are continuously monitored, and if an audio signal greater than a predetermined decibel threshold is detected, the audio signal greater than the predetermined decibel threshold is selected as the target for collection. This audio signal greater than the predetermined decibel threshold is then directionally collected using a unidirectional air conduction microphone to obtain the original unidirectional air conduction signal.

[0074] In the above embodiment, the unidirectional air conduction microphone, acting as the primary air conduction microphone, collects a relatively pure audio signal and ensures a relatively high signal-to-noise ratio, making it easier for the model to decouple the target sound during processing. The omnidirectional air conduction microphone, acting as an auxiliary air conduction microphone, utilizes the characteristic of not leaking frequency bands through air conduction to ensure the integrity of the audio signal input to the model. By working together, the two can improve the accuracy of the predicted amplitude values ​​output from the model.

[0075] In one embodiment, there are multiple air-conducted audio signals, and a short-time Fourier transform is performed on the original air-conducted signal to obtain the frequency domain signal of the original air-conducted signal. An air-conducted audio signal is then obtained based on the original air-conducted signal and its frequency domain signal. This includes performing a short-time Fourier transform on multiple original air conduction signals to obtain the frequency domain signal of each original air conduction signal, and obtaining multiple air conduction audio signals based on each original air conduction signal and its corresponding frequency domain signal.

[0076] Multiple air-conducted audio signals include unidirectional air-conducted audio signals converted from the original unidirectional air-conducted signal.

[0077] Unidirectional air conduction microphones and omnidirectional air conduction microphones can synchronously collect multiple original air conduction signals. When performing a short-time Fourier transform on the original air conduction signals to obtain the frequency domain signals of the original air conduction signals, and then acquiring an air conduction audio signal based on the original air conduction signals and their frequency domain signals, each original air conduction signal is processed separately. As an example where one microphone collects one signal, the unidirectional original air conduction signals collected by each unidirectional air conduction microphone are converted into frequency domain signals using a short-time Fourier transform, and a unidirectional air conduction audio signal is acquired based on each unidirectional original air conduction signal and its frequency domain signal. Similarly, the omnidirectional original air conduction signals collected by each omnidirectional air conduction microphone are converted into frequency domain signals using a short-time Fourier transform, and an omnidirectional air conduction audio signal is acquired based on each omnidirectional original air conduction signal and its frequency domain signal.

[0078] To ensure clarity, this application does not limit the number of unidirectional air conduction microphones, omnidirectional air conduction microphones, unidirectional original air conduction signals, omnidirectional original air conduction signals, unidirectional air conduction audio signals, or omnidirectional air conduction audio signals.

[0079] In one embodiment, multiplying the predicted amplitude by the phase of the air conduction audio signal and using the result obtained from the multiplication as an enhancement signal in the time-frequency domain includes multiplying the predicted amplitude by the phase of the unidirectional air conduction audio signal and using the result obtained from the multiplication as an enhancement signal in the time-frequency domain.

[0080] When a device is equipped with both a unidirectional air conduction microphone and an omnidirectional air conduction microphone, obtaining a time-frequency domain enhancement signal based on the predicted amplitude includes multiplying the predicted amplitude by the phase of the unidirectional air conduction audio signal and using the result of this multiplication as the time-frequency domain enhancement signal.

[0081] If multiple unidirectional air conduction microphones are included, one unidirectional air conduction microphone is designated as the primary air conduction microphone. This involves multiplying the predicted amplitude by the phase of the unidirectional air conduction audio signal corresponding to the primary air conduction microphone, and using the result as the time-frequency domain enhancement signal. For example, the distance from each unidirectional air conduction microphone to the sound source may be obtained, and the unidirectional air conduction microphone closest to the sound source may be designated as the primary air conduction microphone. Alternatively, for example, the average value of the original air conduction signals collected by each unidirectional air conduction microphone may be determined, and the unidirectional air conduction microphone with the largest average value of the collected original air conduction signals may be designated as the primary air conduction microphone.

[0082] In the above embodiment, the original air conduction signal is collected using multiple air conduction microphones, and by utilizing the characteristic that the collection direction of the multiple air conduction microphones is different, and combining this with multi-microphone audio enhancement technology, the amount of information that the model can process is further increased, allowing the model to output a more accurate predicted amplitude based on more information.

[0083] In one embodiment, the voiceprint features, first audio features, and second audio features of the bone conduction audio signal are input to a trained amplitude prediction model before obtaining the predicted amplitude. The method further includes obtaining air conduction audio signal samples and bone conduction audio signal samples, obtaining a third audio feature of the air conduction audio signal samples and a fourth audio feature of the bone conduction audio signal samples, using the bone conduction audio signal samples as input data for a feature extraction model, using the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for an amplitude prediction model, using the amplitude of the air conduction audio signal samples as the target output for the amplitude prediction model, performing fusion training on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model.

[0084] In this application, before obtaining predicted amplitude using the amplitude prediction model, the amplitude prediction model is constructed and the constructed amplitude prediction model is trained. When training the amplitude prediction model, the data preprocessing steps are the same as the data preprocessing steps when using the amplitude prediction model; that is, for a specific explanation of obtaining air conduction audio signal samples and bone conduction audio signal samples, refer to the relevant explanation of obtaining air conduction audio signals and bone conduction audio signals in the above-mentioned embodiment, and for a specific explanation of obtaining the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample, refer to the relevant explanation of obtaining the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal in the above-mentioned embodiment.

[0085] For example, when acquiring air conduction audio signal samples and bone conduction audio signal samples, the original air conduction signal sample is collected using an air conduction microphone, the original air conduction signal sample is collected using a bone conduction microphone, the unidirectional original air conduction signal sample is collected directionally using a unidirectional air conduction microphone, the omnidirectional original air conduction signal sample is collected using an omnidirectional air conduction microphone, a short-time Fourier transform is performed on the original air conduction signal sample to obtain a frequency domain signal sample of the original air conduction signal sample, an air conduction audio signal sample is acquired based on the original air conduction signal sample and the frequency domain signal sample of the original air conduction signal sample, a short-time Fourier transform is performed on the original bone conduction signal sample to obtain a frequency domain signal sample of the original bone conduction signal sample, and a bone conduction audio signal sample is acquired based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample.

[0086] The third and fourth audio features are audio features of the air conduction audio signal sample and the bone conduction audio signal sample, respectively. The third and fourth audio features are of the same type as the first and second audio features. For example, they are all amplitude and phase.

[0087] In one embodiment, the dual model of the present invention is trained using a fusion training method. Bone conduction audio signal samples are used as input data for the feature extraction model, and the output of the feature extraction model, the third audio feature, and the fourth audio feature are used as input data for the amplitude prediction model. Fusion training is then performed on the feature extraction model and the amplitude prediction model. The data processing method used when training the feature extraction model and the amplitude prediction model can be found in the relevant explanation in the embodiment of the audio enhancement method described above, and is therefore omitted here.

[0088] The target output of fusion training may be determined according to the actual needs. If bone conduction audio signals need to be enhanced, the amplitude of the bone conduction audio signal sample should be used as the target output of the amplitude prediction model. If air conduction audio signals need to be enhanced, the amplitude of the air conduction audio signal sample should be used as the target output of the amplitude prediction model. If the device has multiple air conduction microphones and multiple air conduction source signal samples are collected, the amplitude of the air conduction audio signal sample corresponding to the air conduction original signal sample of the main air conduction microphone should be used as the target output.

[0089] When performing fusion training on a feature extraction model and an amplitude prediction model, only one set of loss functions is set for both the feature extraction model and the amplitude prediction model. The amplitude prediction model continuously acquires a predicted output based on the output of the feature extraction model, compares the predicted output of the feature extraction model with the target output corresponding to the input data based on the set loss function, and continuously reduces the loss value of the loss function to perform synchronous optimization on both the feature extraction model and the amplitude prediction model, thereby achieving the objective of fusion training.

[0090] In the above embodiment, the feature extraction model and the amplitude prediction model learn the output of the feature extraction model, as well as the relationship between the audio features of the air conduction audio signal sample and the bone conduction audio signal sample and the amplitude of the air conduction audio signal sample, through fused training. The trained dual model may be used to predict the desired amplitude of the audio in audio enhancement, and further to achieve audio enhancement and improve audio quality.

[0091] In one exemplary embodiment, a fusion training method is provided, as shown in Figure 3, and the method is described as being applied to the terminal in Figure 1, and includes the following steps S302 to S306.

[0092] In step S302, air conduction audio signal samples and bone conduction audio signal samples are acquired.

[0093] In step S304, the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample are acquired.

[0094] In step S306, bone conduction audio signal samples are used as input data for the feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature are used as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples is used as the target output for the amplitude prediction model. Fusion training is performed on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model.

[0095] In one embodiment, air conduction audio signal samples and bone conduction audio signal samples are obtained. This method includes: collecting original air conduction signal samples using an air conduction microphone; collecting original air conduction signal samples using a bone conduction microphone; directionally collecting unidirectional original air conduction signal samples using a unidirectional air conduction microphone; collecting omnidirectional original air conduction signal samples using an omnidirectional air conduction microphone; performing a short-time Fourier transform on the original air conduction signal samples to obtain frequency domain signal samples of the original air conduction signal samples; obtaining air conduction audio signal samples based on the original air conduction signal samples and their frequency domain signal samples; and performing a short-time Fourier transform on the original bone conduction signal samples to obtain frequency domain signal samples of the original bone conduction signal samples; and obtaining bone conduction audio signal samples based on the original bone conduction signal samples and their frequency domain signal samples.

[0096] In the above fusion training method, air conduction audio signal samples and bone conduction audio signal samples are acquired, the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample are acquired, the bone conduction audio signal sample is used as input data for the feature extraction model, the output of the feature extraction model, the third audio feature and the fourth audio feature are used as input data for the amplitude prediction model, the amplitude of the air conduction audio signal sample is used as the target output of the amplitude prediction model, and fusion training is performed on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model. Audio signals collected by air conduction have the characteristic of a wide frequency range, and audio signals collected by bone conduction are hardly affected by ambient noise interference. In this application, a dual model mechanism is designed and trained, and voiceprint features are extracted from the bone conduction audio signal using the primary model. The extracted voiceprint features inherit the advantage of low bone conduction noise and can assist in noise reduction. The secondary model can fuse the voiceprint features output from the primary model with the audio features of the air conduction audio signal and the bone conduction audio signal. The primary and secondary models learn the output of the feature extraction model, as well as the relationship between the audio features of air-conducted audio signal samples and bone-conducted audio signal samples and the amplitude of the air-conducted audio signal samples, through fused training. The trained dual model may be used to predict the desired amplitude of audio in audio enhancement, and further to implement audio enhancement and improve audio quality.

[0097] Referring to Figure 4, the present invention further provides specific embodiments. Specific embodiments of the audio enhancement method and the fusion training method in application scenarios are as follows:

[0098] In S1, an air conduction microphone is used to collect the original air conduction signal sample, and a bone conduction microphone is used to collect the original bone conduction signal sample.

[0099] A unidirectional air conduction microphone is used to directionally collect unidirectional original air conduction signal samples, and an omnidirectional air conduction microphone is used to collect omnidirectional original air conduction signal samples.

[0100] In S2, a short-time Fourier transform is performed on the original air conduction signal sample to obtain a frequency domain signal sample of the original air conduction signal sample, and an air conduction audio signal sample is obtained based on the original air conduction signal sample and its frequency domain signal sample.

[0101] In S3, the amplitude and phase of the air-conducted audio signal sample are acquired.

[0102] In S4, a short-time Fourier transform is performed on the original bone conduction signal sample to obtain a frequency domain signal sample of the original bone conduction signal sample, and a bone conduction audio signal sample is obtained based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample.

[0103] In S5, the amplitude and phase of bone conduction audio signal samples are acquired.

[0104] In S6, bone conduction audio signal samples are used as input data for the feature extraction model, the output of the feature extraction model, the amplitude and phase of the air conduction audio signal samples, and the amplitude and phase of the bone conduction audio signal samples are used as input data for the amplitude prediction model, and the amplitude of the air conduction audio signal samples is used as the target output for the amplitude prediction model. Fusion training is performed on the feature extraction model and the amplitude prediction model to obtain the trained feature extraction model and the trained amplitude prediction model.

[0105] In S7, the system determines the target direction in which the audio signal is greater than a predetermined decibel threshold, and then uses a unidirectional air conduction microphone to directionally collect the audio signal in the target direction, thereby obtaining the unidirectional original air conduction signal.

[0106] In S8, all-direction audio signals in all directions are collected by an all-direction air conduction microphone to obtain an all-direction original air conduction signal.

[0107] In S9, a short-time Fourier transform is performed on a plurality of original air conduction signals to obtain the frequency domain signals of each original air conduction signal, and a plurality of air conduction audio signals are obtained based on each original air conduction signal and the corresponding frequency domain signal.

[0108] The plurality of air conduction audio signals include unidirectional air conduction audio signals converted from unidirectional original air conduction signals.

[0109] The original air conduction signal m with a length of T in the time domain q may be represented as m q (t), where t represents time and 0 < t ≤ T. Thus, after converting the original air conduction signal m q (t) in the time domain to the time-frequency domain by short-time Fourier transform, the obtained air conduction audio signal may be represented by Equation (1). M q (n,k)=STFT(m q (t)) (1) Here, n is the frame sequence, 0 < n ≤ N, N is the total number of frames, k is the center frequency sequence, 0 < k ≤ K, and K is the total number of frequency points. Here, q represents an air conduction microphone, 0 < q ≤ Q, and Q is the total number of original air conduction signals. In some cases, Q is also equal to the total number of air conduction microphones.

[0110] When one air-conduction microphone is installed in the device, Q = 1. The original air-conduction signal is collected by one air-conduction microphone, and the original air-conduction signal is converted into the time-frequency domain by short-time Fourier transform to obtain one air-conduction audio signal. When multiple air-conduction microphones are installed in the device, Q > 1. The original air-conduction signal is collected by multiple air-conduction microphones, and each original air-conduction signal is converted into the time-frequency domain by short-time Fourier transform to obtain multiple air-conduction audio signals.

[0111] In S10, the amplitudes and phases of multiple air-conduction audio signals are obtained.

[0112] Multiple air-conduction audio signals M q The acquisition of the amplitude Mag of M(n,k) may be expressed by Equation (2). For multiple air-conduction audio signals M q The acquisition of the phase Pha of M(n,k) may be expressed by Equation (3). MagM q M(n,k)=abs(M q (n,k)) (2) PhaM q M(n,k)=M q (n,k) / abs(M q (n,k)) (3)

[0113] In S11, the original air-conduction signal is collected by the air-conduction microphone, and the original bone-conduction signal is collected by the bone-conduction microphone.

[0114] In S12, short-time Fourier transform is performed on the original bone-conduction signal to obtain the frequency-domain signal of the original bone-conduction signal, and a bone-conduction audio signal is obtained based on the original bone-conduction signal and the frequency-domain signal of the original bone-conduction signal.

[0115] The original bone conduction signal v with a length of T in the time domain may be represented as v(t), where t represents time and 0 < t ≤ T. After converting the original bone conduction signal v(t) in the time domain to the time-frequency domain by short-time Fourier transform, the obtained bone conduction audio signal may be represented by Equation (4). V(n,k)=STFT(v(t)) (4) Here, n is the frame sequence of the bone conduction audio signal, 0 < n ≤ N, where N is the total number of frames, k is the center frequency sequence, 0 < k ≤ K, and K is the total number of frequency points.

[0116] In S13, the amplitude and phase of the bone conduction audio signal are obtained.

[0117] The acquisition of the amplitude Mag of the bone conduction audio signal V(n,k) may be represented by Equation (5), and the acquisition of the phase Pha of the bone conduction audio signal V(n,k) may be represented by Equation (6). MagV(n,k)=abs(V(n,k)) (5) PhaV(n,k)=V(n,k) / abs(V(n,k)) (6)

[0118] In S14, feature extraction is performed on the bone conduction audio signal by a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal.

[0119] The voiceprint features of the extracted bone conduction audio signal V(n,k) may be represented as DNN1(V(n,k)).

[0120] In S15, the voiceprint features, the first audio feature, and the second audio feature of the bone conduction audio signal are input into a trained amplitude prediction model to obtain a predicted amplitude.

[0121] The voiceprint features DNN1(V(n,k)) of the bone conduction audio signal, the amplitude MagV(n,k) and phase PhaV(n,k) of the bone conduction audio signal, and the amplitude MagM of the air conduction audio signal q(n,k), Phase PhaM q The values ​​(n,k) are input to a trained amplitude prediction model to obtain the predicted amplitude Mag(n,k). This process may also be expressed as equation (7). Mag(n,k)=DNN2〔M_q(n,k),V(n,k),DNN1(V(n,k))〕 (7)

[0122] In S16, the predicted amplitude is multiplied by the phase of the unidirectional air conduction audio signal, and the result obtained from this multiplication is used as the enhancement signal in the time-frequency domain.

[0123] In S17, a short-time inverse Fourier transform is performed on the enhancement signal to obtain the target audio signal.

[0124] The predicted amplitude Mag(n,k) is multiplied by the phase PhaM1(n,k) of the unidirectional air conduction audio signal, and this process may be expressed as equation (8). X(t)=ISTFT(Mag(n,k)*PhaM1(n,k)) (8)

[0125] X(t) is the final enhancement result, i.e., the target audio signal.

[0126] In the above embodiment, the audio signal collected by air conduction has the characteristic of a wide frequency range, while the audio signal collected by bone conduction is hardly affected by ambient noise interference. In this application, a dual-model mechanism is designed and trained, and the primary and secondary models learn the output of the feature extraction model, as well as the relationship between the audio features of the air-conducted audio signal sample and the bone-conducted audio signal sample and the amplitude of the air-conducted audio signal sample through fusion training. The trained primary model can extract voiceprint features from the bone-conducted audio signal, and the extracted voiceprint features can assist in noise reduction, inheriting the advantage of low noise in bone conduction. The voiceprint features and the audio features of the air-conducted audio signal and the bone-conducted audio signal are fused in the trained secondary model, and the resulting predicted amplitude can be used to generate an enhanced audio signal. The trained dual-model mechanism fuses the advantages of bone conduction and air conduction to output an appropriate predicted amplitude, and the target audio signal obtained based on this predicted amplitude achieves noise reduction and improves audio quality without losing frequency bandwidth.

[0127] For aspects not described in detail in the above-mentioned integrated training method, please refer to the relevant explanations in the audio enhancement method of this application, and the explanation will be omitted here.

[0128] In the flowcharts relating to each of the above embodiments, the steps are shown sequentially by arrows, but it should be understood that these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise explicitly stated herein, the execution of these steps is not limited to a strict order, and these steps may be performed in other orders. Furthermore, at least some of the steps in the flowchart relating to each of the above embodiments may include multiple steps or stages, and these steps or stages are not necessarily performed at the same time, but may be performed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed sequentially or alternately with other steps or at least some of the steps or stages within other steps.

[0129] Based on a similar inventive concept, embodiments of the present application further provide an audio enhancement device that realizes the above-described audio enhancement method. Since the means for solving the problems related to the device are similar to the means for solving the problems related to the above-described method, specific limitations in the following embodiments of one or more audio enhancement devices can be made by referring to the limitations on the above-described audio enhancement method, and therefore, a detailed explanation is omitted here.

[0130] In one exemplary embodiment, as shown in Figure 5, A first acquisition module 701 that acquires air conduction audio signals and bone conduction audio signals, A second acquisition module 702 acquires a first audio feature of an air conduction audio signal and a second audio feature of a bone conduction audio signal, A feature extraction module 703 performs feature extraction on bone conduction audio signals using a pre-trained feature extraction model to obtain voiceprint features of the bone conduction audio signals. An amplitude prediction module 704 obtains a predicted amplitude by inputting the voiceprint features, first audio features, and second audio features of a bone conduction audio signal into a trained amplitude prediction model. An audio enhancement device is provided, which includes an audio enhancement module 705 that acquires a target audio signal based on a predicted amplitude.

[0131] In one embodiment, when acquiring air conduction audio signals and bone conduction audio signals, the first acquisition module 701 further... Original air conduction signals are collected using an air conduction microphone, and original bone conduction signals are collected using a bone conduction microphone. A short-time Fourier transform is performed on the original air conduction signal to obtain the frequency domain signal of the original air conduction signal, and an air conduction audio signal is obtained based on the original air conduction signal and the frequency domain signal of the original air conduction signal. A short-time Fourier transform is performed on the original bone conduction signal to obtain the frequency domain signal of the original bone conduction signal, and a bone conduction audio signal is obtained based on the original bone conduction signal and the frequency domain signal of the original bone conduction signal.

[0132] In one embodiment, when enhancing audio based on predicted amplitude to obtain a target audio signal, the audio enhancement module 705 further... The predicted amplitude is multiplied by the phase of the air-conducted audio signal, and the result obtained from this multiplication is used as the enhancement signal in the time-frequency domain. The enhancement signal is subjected to a short-time inverse Fourier transform to obtain the target audio signal.

[0133] In one embodiment, when collecting the original air conduction signal using an air conduction microphone, the first acquisition module 701 further... Multiple air conduction microphones are used to collect the original air conduction signal from different directions, thereby obtaining multiple original air conduction signals.

[0134] In one embodiment, the multiple air conduction microphones include unidirectional air conduction microphones and omnidirectional air conduction microphones, and when the multiple air conduction microphones collect original air conduction signals from different directions to acquire multiple original air conduction signals, the first acquisition module 701 further... The system determines the target direction in which the audio signal is greater than a predetermined decibel threshold, and uses a unidirectional air conduction microphone to directionally collect the audio signal in the target direction, thereby obtaining the unidirectional original air conduction signal. An omnidirectional air conduction microphone collects audio signals from all directions, obtaining the original omnidirectional air conduction signal.

[0135] In one embodiment, there are multiple air-conducted audio signals, and a short-time Fourier transform is performed on the original air-conducted signal to obtain the frequency domain signal of the original air-conducted signal. When acquiring an air-conducted audio signal based on the original air-conducted signal and the frequency domain signal of the original air-conducted signal, the first acquisition module 701 further... A short-time Fourier transform is performed on multiple original air conduction signals to obtain the frequency domain signal of each original air conduction signal, and based on each original air conduction signal and its corresponding frequency domain signal, multiple air conduction audio signals are obtained, including a unidirectional air conduction audio signal converted from the unidirectional original air conduction signal. The predicted amplitude is multiplied by the phase of the air-conducted audio signal, and the result obtained from this multiplication is used as the enhancement signal in the time-frequency domain. The predicted amplitude is multiplied by the phase of the unidirectional air conduction audio signal, and the result obtained from this multiplication is used as the enhancement signal in the time-frequency domain.

[0136] In one embodiment, referring to Figure 6, the audio enhancement device further includes a model training module 706, which inputs the voiceprint features, first audio features, and second audio features of the bone conduction audio signal into a trained amplitude prediction model to obtain the predicted amplitude. We obtained samples of air conduction audio signals and bone conduction audio signals. The third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample were obtained. Bone conduction audio signal samples are used as input data for a feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature are used as input data for an amplitude prediction model, and the amplitude of an air conduction audio signal sample is used as the target output for the amplitude prediction model. Fusion training is performed on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model.

[0137] In one embodiment, when acquiring air conduction audio signal samples and bone conduction audio signal samples, the model training module 706 further... Original air conduction signal samples are collected using an air conduction microphone, original air conduction signal samples are collected using a bone conduction microphone, unidirectional original air conduction signal samples are collected directionally using a unidirectional air conduction microphone, and omnidirectional original air conduction signal samples are collected using an omnidirectional air conduction microphone. A short-time Fourier transform is performed on the original air conduction signal sample to obtain a frequency domain signal sample of the original air conduction signal sample, and an air conduction audio signal sample is obtained based on the original air conduction signal sample and the frequency domain signal sample of the original air conduction signal sample. A short-time Fourier transform is performed on the original bone conduction signal sample to obtain a frequency domain signal sample of the original bone conduction signal sample, and a bone conduction audio signal sample is obtained based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample.

[0138] In the above audio enhancement device, the first acquisition module 701 acquires air conduction audio signals and bone conduction audio signals, the second acquisition module 702 acquires the first audio feature of the air conduction audio signal and the second audio feature of the bone conduction audio signal, the feature extraction module 703 performs feature extraction on the bone conduction audio signal using a trained feature extraction model to acquire the voiceprint features of the bone conduction audio signal, the amplitude prediction module 704 inputs the voiceprint features, first audio features, and second audio features of the bone conduction audio signal into a trained amplitude prediction model to acquire the predicted amplitude, and the audio enhancement module 705 acquires the target audio signal based on the predicted amplitude. Audio signals collected by air conduction have the characteristic of having a wide frequency range, and audio signals collected by bone conduction are hardly affected by ambient noise interference. In this application, a primary model is used to extract voiceprint features from the bone conduction audio signal, and the extracted voiceprint features can assist in noise reduction, inheriting the advantage of low bone conduction noise. Voiceprint features and audio features of air conduction audio signals and bone conduction audio signals are fused in a pre-trained secondary model, and the resulting predicted amplitude can be used to generate an enhanced audio signal. The dual-model mechanism of this application combines the advantages of bone conduction and air conduction, and the resulting target audio signal achieves noise reduction and improves audio quality without losing frequency bandwidth.

[0139] In one exemplary embodiment, as shown in Figure 7, A fourth acquisition module 801 for acquiring air conduction audio signal samples and bone conduction audio signal samples, A fifth acquisition module 802 acquires the third audio feature of an air conduction audio signal sample and the fourth audio feature of a bone conduction audio signal sample, A fusion training device is provided, which includes a fusion training module 803 that takes bone conduction audio signal samples as input data for a feature extraction model, uses the output of the feature extraction model, a third audio feature, and a fourth audio feature as input data for an amplitude prediction model, uses the amplitude of an air conduction audio signal sample as the target output of the amplitude prediction model, performs fusion training on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model.

[0140] In one embodiment, when acquiring air conduction audio signal samples and bone conduction audio signal samples, the fourth acquisition module 801 further... Original air conduction signal samples are collected using an air conduction microphone, original air conduction signal samples are collected using a bone conduction microphone, unidirectional original air conduction signal samples are collected directionally using a unidirectional air conduction microphone, and omnidirectional original air conduction signal samples are collected using an omnidirectional air conduction microphone. A short-time Fourier transform is performed on the original air conduction signal sample to obtain a frequency domain signal sample of the original air conduction signal sample, and an air conduction audio signal sample is obtained based on the original air conduction signal sample and the frequency domain signal sample of the original air conduction signal sample. A short-time Fourier transform is performed on the original bone conduction signal sample to obtain a frequency domain signal sample of the original bone conduction signal sample, and a bone conduction audio signal sample is obtained based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample.

[0141] In the above-described fusion training device, the fourth acquisition module 801 acquires air conduction audio signal samples and bone conduction audio signal samples, the fifth acquisition module 802 acquires the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample, and the fusion training module 803 uses the bone conduction audio signal sample as input data for a feature extraction model, the output of the feature extraction model, the third audio feature, and the fourth audio feature as input data for an amplitude prediction model, and the amplitude of the air conduction audio signal sample as the target output of the amplitude prediction model. Fusion training is performed on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model. Audio signals collected by air conduction have the characteristic of a wide frequency range, and audio signals collected by bone conduction are hardly affected by ambient noise interference. In this invention, a dual-model mechanism is designed and trained, and voiceprint features are extracted from the bone conduction audio signal using the primary model. The extracted voiceprint features inherit the advantage of low bone conduction noise and can assist in noise reduction. The secondary model can fuse voiceprint features output from the primary model with audio features from air-conducted and bone-conducted audio signals. The primary and secondary models learn the output of the feature extraction model, as well as the relationship between the audio features of air-conducted and bone-conducted audio signal samples and the amplitude of the air-conducted audio signal samples, through fusion training. The trained dual model may be used to predict desired audio amplitudes in audio enhancement, further implement audio enhancement, and improve audio quality.

[0142] All or part of the modules within the audio enhancement device and the fusion training device described above may be implemented by software, hardware, or a combination thereof. Each of the above modules may be embedded in the processor of the computer device in hardware form, or it may be independent of this processor, or it may be stored in the memory of the computer device in software form, in order to facilitate the processor calling and executing the operations corresponding to each of the above modules.

[0143] For example, terms such as “assembly,” “module,” and “system” are intended to refer to entities related to a computer, which may be hardware, a combination of hardware and software, software, or running software. For example, an assembly may be, but is not limited to, a process running on a processor, a processor, an object, executable code, an execution thread, a program, and / or a computer. For illustrative purposes, an execution program running on a server and a server may both be assemblies. One or more assemblies may reside in a process and / or an execution thread, and assemblies may reside in one computer and / or be distributed across two or more computers.

[0144] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal configuration diagram may be as shown in Figure 8. The computer device includes a processor, memory, an input / output interface (abbreviated as I / O), and a communication interface. The processor, memory, and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device provides computation and control functions. The memory of the computer device includes a non-volatile computer-readable storage medium and internal memory. The non-volatile computer-readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the execution of the operating system and computer-readable instructions in the non-volatile computer-readable storage medium. The input / output interface of the computer device exchanges information between the processor and external devices. The communication interface of the computer device is connected to an external terminal via a network for communication. When the computer-readable instructions are executed by the processor, an audio enhancement method is realized.

[0145] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal configuration diagram may be as shown in Figure 9. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computation and control functions. The memory of the computer device includes a non-volatile computer-readable storage medium and internal memory. The non-volatile computer-readable storage medium stores an operating system and computer-readable instructions. The internal memory provides an environment for the execution of the operating system and computer-readable instructions in the non-volatile computer-readable storage medium. The input / output interface of the computer device exchanges information between the processor and external devices. The communication interface of the computer device communicates with external terminals by wired or wireless means, and the wireless means may be implemented by Wi-Fi, a mobile cellular network, Near Field Communication (NFC), or other technology. When the computer-readable instruction is executed by the processor, an audio enhancement method is realized. The display unit of the computer device forms a visually visible screen and may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered by the display screen, a key, trackball, or touchpad installed on the casing of the computer device, or an external keyboard, touchpad, or mouse.

[0146] As those skilled in the art will understand, the structures shown in Figures 8 and 9 are merely block diagrams of some structures related to the solution of the present invention and do not limit the computer equipment to which the solution of the present invention is applied. Specific computer equipment may include more or fewer components than those shown, may combine several components, or may have different component arrangements.

[0147] In one embodiment, a computer device is further provided, which includes a memory storing computer-readable instructions and a processor, and when the processor executes the computer-readable instructions, the steps in each embodiment of the above method are realized.

[0148] In one embodiment, a computer-readable storage medium is provided in which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the steps in each embodiment of the above method are realized.

[0149] In one embodiment, a computer program product is provided which includes computer-readable instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer-readable instructions from the computer-readable storage medium and executes the computer-readable instructions, thereby causing the computer device to perform the steps in each of the above-described embodiment of the method.

[0150] As those skilled in the art will understand, the flow of all or part of the above-described method embodiments may be completed by a computer-readable instruction commanding the relevant hardware, the computer-readable instruction being stored in a non-volatile computer-readable storage medium and, when executed, including the flow of each of the above-described method embodiments. Any reference to memory, database, or other medium used in each embodiment of the present application may include at least one of non-volatile memory and / or volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. Rather than being limited, RAM may take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database in the embodiments of this application may include at least one of relational databases and non-relational databases. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processor in each embodiment of this application may be, but is not limited to, a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, programmable logic, quantum computing-based data processing logic, an artificial intelligence (AI) processor, and the like.

[0151] The technical features of the above embodiments can be combined in any way, and for the sake of explanation, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in these combinations of technical features, they should fall within the scope described in this application.

[0152] The above embodiments merely illustrate some embodiments of the present application, and although the descriptions are specific and detailed, they should not be understood as limiting the scope of the claims of the present application. Furthermore, a person skilled in the art could make several modifications and improvements without departing from the concept of the present application, and these also fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the claims.

Claims

1. Acquiring air conduction audio signals using an air conduction microphone and acquiring bone conduction audio signals using a bone conduction microphone, To acquire the first audio characteristics of the air conduction audio signal and the second audio characteristics of the bone conduction audio signal, The process involves performing feature extraction on the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal, The voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features are input to a trained amplitude prediction model to obtain a predicted amplitude. An audio enhancement method characterized by comprising acquiring a target audio signal based on the predicted amplitude.

2. Acquiring the aforementioned air conduction audio signal and bone conduction audio signal is, The air conduction microphone collects the original air conduction signal, and the bone conduction microphone collects the original bone conduction signal. A short-time Fourier transform is performed on the original air conduction signal to obtain the frequency domain signal of the original air conduction signal, and an air conduction audio signal is obtained based on the original air conduction signal and the frequency domain signal of the original air conduction signal. The method according to claim 1, characterized by comprising: performing a short-time Fourier transform on the original bone conduction signal to obtain a frequency domain signal of the original bone conduction signal; and obtaining a bone conduction audio signal based on the original bone conduction signal and the frequency domain signal of the original bone conduction signal.

3. Acquiring the target audio signal based on the predicted amplitude is, The predicted amplitude and the phase of the air-conducted audio signal are multiplied, and the result obtained from this multiplication is used as the enhancement signal in the time-frequency domain. The method according to the second invention, characterized in that it includes performing a short-time inverse Fourier transform on the enhancement signal to obtain a target audio signal.

4. Collecting the original air conduction signal using the aforementioned air conduction microphone is, The method according to claim 3, characterized in that it includes acquiring multiple original air conduction signals by collecting original air conduction signals from different directions using multiple air conduction microphones.

5. The aforementioned plurality of air conduction microphones include unidirectional air conduction microphones and omnidirectional air conduction microphones, and the plurality of air conduction microphones collect original air conduction signals from different directions to obtain multiple original air conduction signals. The process involves determining a target direction in which the audio signal is greater than a predetermined decibel threshold, and then directionally collecting the audio signal in that target direction using the unidirectional air conduction microphone to obtain the unidirectional original air conduction signal. The method according to claim 4, characterized in that it includes collecting audio signals from all directions using the omnidirectional air conduction microphone to obtain an omnidirectional original air conduction signal.

6. There are multiple air-conducted audio signals, and a short-time Fourier transform is performed on the original air-conducted signal to obtain the frequency domain signal of the original air-conducted signal, and an air-conducted audio signal is obtained based on the original air-conducted signal and the frequency domain signal of the original air-conducted signal. The process includes performing a short-time Fourier transform on the plurality of original air conduction signals to obtain the frequency domain signal of each original air conduction signal, and obtaining a plurality of air conduction audio signals, including the unidirectional air conduction audio signal converted from the unidirectional original air conduction signal, based on each original air conduction signal and the corresponding frequency domain signal. Multiplying the predicted amplitude by the phase of the air-conducted audio signal and using the result obtained from this multiplication as an enhancement signal in the time-frequency domain is: The method according to claim 5, characterized in that it includes multiplying the predicted amplitude by the phase of the unidirectional air conduction audio signal, and using the result obtained from the multiplication as an enhancement signal in the time-frequency domain.

7. Before obtaining the predicted amplitude, input the voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features into a trained amplitude prediction model. To obtain air conduction audio signal samples and bone conduction audio signal samples, To obtain the third audio feature of the air conduction audio signal sample and the fourth audio feature of the bone conduction audio signal sample, The method according to any one of claims 1 to 6, further comprising: using the bone conduction audio signal sample as input data for the feature extraction model; using the output of the feature extraction model, the third audio feature and the fourth audio feature as input data for the amplitude prediction model; using the amplitude of the air conduction audio signal sample as the target output for the amplitude prediction model; and performing fusion training on the feature extraction model and the amplitude prediction model to obtain a trained feature extraction model and a trained amplitude prediction model.

8. Obtaining the aforementioned air conduction audio signal samples and bone conduction audio signal samples is possible. This involves collecting original air conduction signal samples using an air conduction microphone, collecting original bone conduction signal samples using a bone conduction microphone, directionally collecting unidirectional original air conduction signal samples using a unidirectional air conduction microphone, and collecting omnidirectional original air conduction signal samples using an omnidirectional air conduction microphone. A short-time Fourier transform is performed on the original air conduction signal sample to obtain a frequency domain signal sample of the original air conduction signal sample, and an air conduction audio signal sample is obtained based on the original air conduction signal sample and the frequency domain signal sample of the original air conduction signal sample. The method according to the present invention, characterized in that it includes performing a short-time Fourier transform on the original bone conduction signal sample to obtain a frequency domain signal sample of the original bone conduction signal sample, and obtaining a bone conduction audio signal sample based on the original bone conduction signal sample and the frequency domain signal sample of the original bone conduction signal sample.

9. A first acquisition module for acquiring air conduction audio signals and bone conduction audio signals, A second acquisition module for acquiring a first audio feature of the air conduction audio signal and a second audio feature of the bone conduction audio signal, A feature extraction module that performs feature extraction on the bone conduction audio signal using a trained feature extraction model to obtain the voiceprint features of the bone conduction audio signal, An amplitude prediction module that inputs the voiceprint features of the bone conduction audio signal, the first audio features, and the second audio features into a trained amplitude prediction model to obtain a predicted amplitude, An audio enhancement device characterized by including an audio enhancement module that acquires a target audio signal based on the predicted amplitude.

10. An earphone comprising an air conduction microphone for collecting air conduction signals, a bone conduction microphone for collecting bone conduction signals, a memory storing a computer program, and a processor, wherein when the processor executes the computer program, it realizes a step of the method according to any one of claims 1 to 8.