Speech enhancement method and apparatus, electronic device, storage medium, and program

By detecting the wearing status in the headphones, collecting audio data inside and outside the ear canal, calculating the frequency response and occlusion effect curves, and performing filtering and noise signal enhancement, the problem of excessive resource consumption in voice enhancement technology is solved, and the headphone battery life is improved.

CN116193313BActive Publication Date: 2026-05-05XIAOMI TECH (WUHAN) CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAOMI TECH (WUHAN) CO LTD
Filing Date
2022-12-23
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing voice enhancement technologies consume a lot of computing resources, resulting in short battery life for headphone products.

Method used

By detecting when a user is wearing headphones, the system emits prompt audio data through the speaker, collects audio data inside and outside the ear canal, calculates the frequency response curve and the target occlusion effect curve, performs filtering and noise signal enhancement, and obtains an enhanced sound signal.

Benefits of technology

It reduces the high computational resource consumption of voice enhancement, improves voice enhancement efficiency, and extends the battery life of headphone products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116193313B_ABST
    Figure CN116193313B_ABST
Patent Text Reader

Abstract

This disclosure provides a speech enhancement method, apparatus, electronic device, storage medium, and program, relating to the field of audio processing technology. The specific steps are as follows: in response to detecting that a user is wearing headphones, a prompt audio data is emitted through a speaker in the headphones, and audio data within the ear canal is collected; a target occlusion effect curve is determined based on the prompt audio data and the audio data within the ear canal; a second sound signal is filtered to obtain a third sound signal; and a noise signal is calculated based on the first sound signal to enhance the third sound signal. This disclosure determines the frequency response curve within the ear canal and the occlusion effect curve by using the prompt audio data and the audio data within the ear canal, thereby enhancing the third sound signal, avoiding high computational resource consumption during speech enhancement, and improving the efficiency of speech enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing technology, and in particular to a speech enhancement method, apparatus, electronic device, storage medium, and program. Background Technology

[0002] Among related technologies, speech enhancement technology was one of the earliest applications in phone calls, reducing ambient noise around the speaker. Initially, it was mainly based on traditional signal processing methods, estimating noise through Gaussian assumptions, calculating gain values ​​based on the estimated noise, and then modulating the noisy signal to obtain a clean speech signal. In recent years, with the improvement of computing power, neural networks have been used to improve algorithm performance, especially in scenarios with multiple speakers and non-stationary noise. However, current speech enhancement algorithms consume a lot of computing resources and power, resulting in short battery life for headphone products. Summary of the Invention

[0003] This disclosure provides a speech enhancement method, apparatus, electronic device, storage medium, and program to at least solve the problem of excessive computational resources consumed by speech enhancement in related technologies. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a speech enhancement method is provided, comprising:

[0005] In response to detecting that a user is wearing headphones, a prompt audio data is emitted through the speaker in the headphones, and audio data inside the ear canal is collected;

[0006] The frequency response curve is calculated based on the prompt audio data and the audio data inside the ear canal, and the target occlusion effect curve is determined based on the frequency response curve.

[0007] The first sound signal outside the ear canal is collected, and the second sound signal inside the ear canal is collected.

[0008] The second sound signal is filtered according to the target occlusion effect curve to obtain the third sound signal;

[0009] A noise signal is calculated based on the first sound signal, and the third sound signal is enhanced based on the noise signal to obtain an enhanced sound signal.

[0010] Optionally, the step of calculating the frequency response curve based on the prompt audio data and the audio data inside the ear canal specifically includes:

[0011] The time-domain signal corresponding to the prompt audio data is calculated and Fourier transformed to generate the prompt audio domain signal. The time-domain signal of the audio data in the ear canal is then Fourier transformed to generate the frequency domain signal in the ear canal.

[0012] The cross power spectrum of the prompt audio domain signal and the ear canal frequency domain signal is calculated based on the prompt audio domain signal and the ear canal frequency domain signal, and the auto power spectrum of the prompt audio domain signal is also calculated.

[0013] Divide the cross power spectrum by the self power spectrum to obtain the frequency response curve.

[0014] Optionally, after the step of calculating the noise signal based on the first sound signal, the method further includes:

[0015] The preset passive noise reduction curve is multiplied point-to-point with the noise signal to obtain the leakage noise signal.

[0016] Optionally, the step of determining the target suffocation effect curve based on the frequency response curve specifically includes:

[0017] Obtain a preset mapping table, wherein the mapping table includes the correspondence between a preset frequency response curve and a preset occlusion effect curve;

[0018] The preset occlusion effect curve corresponding to the frequency response curve in the mapping table is determined as the target occlusion effect curve.

[0019] Optionally, the step of filtering the second sound signal according to the target occlusion effect curve to obtain the third sound signal specifically includes:

[0020] The target occlusion effect curve is convolved with the second sound signal to obtain the third sound signal.

[0021] Optionally, the step of enhancing the third audio signal based on the noise signal specifically includes:

[0022] Obtain a preset reference power spectrum, wherein the reference power spectrum is the power spectrum corresponding to a signal that contains only human voice;

[0023] Divide the reference power spectrum by the power spectrum of the noise signal to obtain the posterior signal-to-noise ratio;

[0024] The predicted prior signal-to-noise ratio (SNR) of the current signal frame is calculated by combining the posterior SNR with the prior SNR of the previous signal frame.

[0025] A filter function is determined based on the prior signal-to-noise ratio, and the third audio signal is enhanced based on the filter function.

[0026] Optionally, the formulaic expression for the posterior signal-to-noise ratio is as follows:

[0027] Where, γ k(n) represents the posterior signal-to-noise ratio, where n is the signal frame, k is the frequency point, and P yy (,n) represents the power spectrum of the third sound signal, P dd (,n) represents the power spectrum of the noise signal;

[0028] The formulaic expression for the prior signal-to-noise ratio is as follows:

[0029] Where, ξ k (n) represents the prior signal-to-noise ratio, P xx (,n) represents the reference power spectrum.

[0030] Optionally, the formula for calculating the predicted prior signal-to-noise ratio (SNR) of the current signal frame from the posterior SNR and the prior SNR of the previous signal frame is expressed as follows:

[0031] Where α is the smoothing factor, This is the prior signal-to-noise ratio prediction value. P is the predicted reference power spectrum value of the previous frame of the current signal frame. dd (,n-1) represents the power spectrum of the noise signal in the previous frame, max( k (n)-1,0) is the selection of γ k The maximum value between (n)-1 and 0.

[0032] Optionally, the step of determining a filter function based on the prior signal-to-noise ratio and enhancing the third audio signal based on the filter function specifically includes:

[0033] The filter function is obtained based on the prior signal-to-noise ratio prediction value of the current signal frame, and can be expressed as follows:

[0034] Among them, H k (n) is the filter function;

[0035] Gain control is applied to the third audio signal based on the filter function and the posterior signal-to-noise ratio to enhance the third audio signal and obtain the enhanced audio signal.

[0036] According to a second aspect of the present disclosure, a voice enhancement device is provided, comprising:

[0037] The first acquisition module is used to respond to the detection that the user is wearing headphones by emitting prompt audio data through the speaker in the headphones and acquiring audio data in the ear canal;

[0038] The occlusion effect curve acquisition module is used to calculate the frequency response curve based on the prompt audio data and the audio data in the ear canal, and to determine the target occlusion effect curve based on the frequency response curve.

[0039] The second acquisition module is used to acquire the first sound signal outside the ear canal and the second sound signal inside the ear canal;

[0040] The filtering module is used to filter the second sound signal according to the target occlusion effect curve to obtain the third sound signal;

[0041] The voice enhancement module is used to calculate a noise signal based on the first sound signal, and enhance the third sound signal based on the noise signal to obtain an enhanced sound signal.

[0042] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0043] processor;

[0044] Memory used to store the processor's executable instructions;

[0045] The processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.

[0046] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method as described in any one of the first aspects above.

[0047] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method according to any one of the first aspects described above.

[0048] According to a sixth aspect of the present disclosure, an earphone is provided, comprising:

[0049] The headset body and the call microphone, feedforward microphone, and feedback microphone mounted on the headset body, as well as bone conduction sensors or gyroscopes, are included.

[0050] The microphone is used to collect user voice data;

[0051] The feedforward microphone is used to collect audio data outside the user's ear canal;

[0052] The feedback microphone is used to collect audio data from inside the user's ear canal;

[0053] The bone conduction sensor or gyroscope is used to detect the user's speaking state;

[0054] processor;

[0055] Memory used to store processor-executable instructions;

[0056] The processor is configured to execute the method described in any one of the embodiments of the first aspect.

[0057] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0058] This disclosure determines the frequency response curve within the ear canal and the occlusion effect curve by using prompt audio data and ear canal audio data, thereby enhancing the third sound signal. This avoids the high computational resource consumption of voice enhancement, improves the efficiency of voice enhancement, and helps reduce the power consumption of headphone products and increase battery life.

[0059] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0061] Figure 1 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment.

[0062] Figure 2 This is a schematic diagram of a Bluetooth headset structure according to an exemplary embodiment.

[0063] Figure 3 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment.

[0064] Figure 4 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment.

[0065] Figure 5 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment.

[0066] Figure 6 This is a block diagram illustrating a speech enhancement device according to an exemplary embodiment.

[0067] Figure 7 This is a block diagram illustrating an apparatus according to an exemplary embodiment.

[0068] Figure 8 This is a block diagram illustrating an apparatus according to an exemplary embodiment. Detailed Implementation

[0069] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0070] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0071] With social progress and the improvement of people's living standards, headphones have become an indispensable part of daily life. People have raised new requirements for the sound quality, practicality, and call quality of headphones. However, all types of headphones, especially in-ear headphones, will obstruct or block the wearer's ear canal to some extent during use. This can lead to the blockage effect, causing a decrease in the user's hearing threshold and making their own voice sound muffled. The blockage effect occurs when sound is conducted through the bone to the ear canal, but due to the obstruction of the earpiece, it cannot diffuse outwards, resulting in a significant sound amplification at low frequencies. When a person speaks, sound travels not only through the mouth but also through the bone to the ear canal. However, when wearing in-ear headphones, a large portion of external noise is blocked from entering the ear canal due to the obstruction of the headphones. Therefore, the external microphone picks up more noise, while the feedback microphone inside the ear canal picks up much less noise. During a call using headphones, the other party hears the sound picked up by the external microphone, so they hear more noise.

[0072] There are two main methods for enhancing speech captured by a call microphone in related technologies:

[0073] 1. Based on signal processing methods, noise is estimated, then the signal-to-noise ratio is calculated, a Wiener filter is designed, and the final gain value is calculated to correct the noisy signal and obtain the final clean speech.

[0074] 2. Based on neural network methods, features of signals such as noise and speech are extracted, and a noise reduction model is trained. Through this model, a noisy speech signal can be transformed into a clean speech signal.

[0075] Figure 1 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment, such as... Figure 1As shown, the method includes the following steps.

[0076] Step 101: In response to detecting that the user is wearing headphones, prompt audio data is emitted through the speaker in the headphones, and audio data in the ear canal is collected.

[0077] In this embodiment, the earphone is equipped with a sensor that can detect whether the user is wearing the earphone. When the user is detected wearing the earphone, it indicates that the user may intend to use the earphone for a call. In order to ensure timely voice enhancement during the user's call, it is necessary to emit prompt audio data through the speaker in the earphone to test the microphone in the earphone and determine the target occlusion effect curve.

[0078] Figure 2 This is a schematic diagram illustrating the structure of a Bluetooth headset according to an exemplary embodiment. Figure 2 As shown, the Bluetooth headset includes a call microphone, a feedforward microphone, and a feedback microphone. The call microphone and feedforward microphone are located outside the user's ear, while the feedback microphone is located inside the user's ear canal. The call microphone is used to collect the user's voice data, the feedforward microphone is used to collect audio data outside the user's ear canal, and the feedback microphone is used to collect audio data inside the user's ear canal. This embodiment utilizes the feedback microphone to collect the audio data inside the ear canal to enhance the human voice voice of the audio data collected by the call microphone.

[0079] Step 102: Calculate the frequency response curve based on the prompt audio data and the audio data inside the ear canal, and determine the target occlusion effect curve based on the frequency response curve.

[0080] In this embodiment, the frequency response curve is a shorthand for the frequency response curve. The frequency response curve reflects the processing capability of the feedback microphone for audio prompts when a user wears the headphones. In a test circuit, the output signal frequency of the signal generator is continuously varied while maintaining a constant amplitude. The amplifier's corresponding output level for this continuous variation is recorded at the output terminal using an oscilloscope or other recorder, thus plotting a curve on a coordinate system representing the frequency corresponding to a given level. After determining the frequency response curve, the target blocking effect curve within the headphones can be determined based on it. The target blocking effect curve reflects the characteristics of the blocking effect within the headphones, and filters can be designed based on this curve to reduce the impact of the blocking effect.

[0081] Step 103: Collect the first sound signal outside the ear canal and collect the second sound signal inside the ear canal.

[0082] During a user's call or recording using the headset, the first sound signal is collected by the call microphone, and the second sound signal is collected by the feedback microphone.

[0083] Step 104: Filter the second sound signal according to the target occlusion effect curve to obtain the third sound signal.

[0084] In this embodiment, the target occlusion effect curve reflects the characteristics of the occlusion effect in the headphones. A filter can be designed based on this curve to filter the second sound signal to remove the low-frequency enhancement caused by the occlusion effect and obtain the third sound signal.

[0085] Step 105: Calculate the noise signal based on the first sound signal, and enhance the third sound signal based on the noise signal to obtain an enhanced sound signal.

[0086] In this embodiment, due to the passive noise reduction of the earphone shell, the first sound signal contains more external noise. By obtaining the noise signal in the first sound signal, the third sound signal is enhanced based on the noise signal to obtain an enhanced sound signal. This allows only the user's voice to be enhanced while suppressing the voices of other speakers. This enhances the user's voice, ensuring the other end hears a clear human voice.

[0087] This disclosure determines the frequency response curve within the ear canal and the occlusion effect curve by using prompt audio data and ear canal audio data, thereby enhancing the third sound signal. This avoids the high computational resource consumption of voice enhancement, improves the efficiency of voice enhancement, and helps reduce the power consumption of headphone products and increase battery life.

[0088] Figure 3 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment, such as... Figure 3 As shown, Figure 1 Step 102 includes the following steps.

[0089] Step 201: Calculate the time-domain signal corresponding to the prompt audio data, perform a Fourier transform to generate the prompt audio domain signal, and perform a Fourier transform on the time-domain signal of the audio data in the ear canal to generate the frequency domain signal in the ear canal.

[0090] The prompt audio data and the audio data inside the ear canal are time-domain signals. They are first converted into frequency-domain signals to analyze the components at each frequency. This embodiment uses Fourier transform. This is used to convert time-domain signals to frequency-domain signals.

[0091] Step 202: Calculate the cross power spectrum of the prompt audio domain signal and the ear canal frequency domain signal based on the prompt audio domain signal and the ear canal frequency domain signal, and calculate the auto-power spectrum of the prompt audio domain signal.

[0092] In this embodiment, the cross power spectrum and the self power spectrum are calculated using the following formulas:

[0093]

[0094]

[0095] Wherein, A is the prompt audio domain signal, B is the ear canal frequency domain signal, and G... AB Let G be the cross-power spectrum of A and B. AA Let A be the auto-power spectrum, and N be the number of points in the Fourier transform. The auto-power spectrum reflects the waveform similarity between signal samples in the same cue audio domain at the same moment. The cross-power spectrum, also known as the cross spectrum, is used to describe the statistical correlation between the cue audio domain signal and the frequency domain signal in the ear canal in the frequency domain.

[0096] Step 203: Divide the cross power spectrum by the self power spectrum to obtain the frequency response curve.

[0097] In this embodiment, the frequency response curve is obtained using the following formula: Where H AB This refers to the frequency response curve.

[0098] Optionally, after the step of calculating the noise signal based on the first sound signal, the method further includes:

[0099] The preset passive noise reduction curve is multiplied point-to-point with the noise signal to obtain the leakage noise signal.

[0100] In this embodiment, the method for calculating the noise signal based on the first sound signal is based on existing noise estimation methods, such as the recursive averaging method with minimum value control. The obtained noise signal is M, and the noise leaking into the ear canal, N = M × H1, can be calculated based on the preset passive noise reduction curve H1.

[0101] Figure 4 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment, such as... Figure 4 As shown, Figure 1 Step 102 includes the following steps.

[0102] Step 301: Obtain a preset mapping table, wherein the mapping table includes the correspondence between a preset frequency response curve and a preset occlusion effect curve.

[0103] Step 302: Determine the preset suffocation effect curve corresponding to the frequency response curve in the mapping table as the target suffocation effect curve.

[0104] In this embodiment, the mapping relationship between different frequency response curves and blockage curves is preset and stored in a mapping table. After determining the frequency response curve, the preset blockage effect curve corresponding to the frequency response curve can be determined as the target blockage effect curve by looking up the table.

[0105] Optionally, the step of filtering the second sound signal according to the target occlusion effect curve to obtain the third sound signal specifically includes:

[0106] The target occlusion effect curve is convolved with the second sound signal to obtain the third sound signal.

[0107] Based on the target occlusion effect curve, a filter F is designed to act on the frequency domain signal (denoted as P) within the ear canal. This filter can remove the low-frequency enhancement caused by the occlusion effect, obtaining a third sound signal (denoted as Q) with a high signal-to-noise ratio. The calculation formula is as follows:

[0108] Q = P * F, where * represents convolution operation.

[0109] It should be noted that the filter can be designed or stored in a mapping table, and the corresponding filter F can be selected directly based on the target blocking effect curve.

[0110] Figure 5 This is a flowchart illustrating a speech enhancement method according to an exemplary embodiment, such as... Figure 5 As shown, Figure 1 Step 105 includes the following steps.

[0111] Step 401: Obtain a preset reference power spectrum, wherein the reference power spectrum is the power spectrum corresponding to a signal that contains only human voice.

[0112] In this embodiment, the estimated noise signal is applied to calculate the gain value. First, the reference power spectrum is obtained, which contains only the power spectrum of the pure human voice and needs to be estimated.

[0113] Step 402: Divide the reference power spectrum by the power spectrum of the noise signal to obtain the posterior signal-to-noise ratio.

[0114] Optionally, the formulaic expression for the posterior signal-to-noise ratio is as follows:

[0115] Where, γ k (n) represents the posterior signal-to-noise ratio, where n is the signal frame, k is the frequency point, and P yy (,n) represents the power spectrum of the third sound signal, P dd (,n) represents the power spectrum of the noise signal;

[0116] The formulaic expression for the prior signal-to-noise ratio is as follows:

[0117] Where, ξ k (n) represents the prior signal-to-noise ratio, P xx (,n) represents the reference power spectrum.

[0118] Step 403: Calculate the predicted a priori signal-to-noise ratio (SNR) of the current signal frame by combining the posterior SNR and the prior SNR of the previous signal frame.

[0119] The formula for calculating the predicted prior signal-to-noise ratio (SNR) of the current signal frame from the posterior SNR and the prior SNR of the previous signal frame is expressed as follows:

[0120] Where α is the smoothing factor, This is the prior signal-to-noise ratio prediction value. P is the predicted reference power spectrum value of the previous frame of the current signal frame. dd (,n-1) represents the power spectrum of the noise signal in the frame preceding the current signal frame. max( k (n)-1,0) is the selection of γ k The maximum value between (n)-1 and 0 can be used to filter out frequency points in the posterior signal-to-noise ratio where the spectral component is less than zero. This embodiment implements the estimation of the prior signal-to-noise ratio using the posterior signal-to-noise ratio with the dcision-directed formula.

[0121] Step 404: Determine the filter function based on the prior signal-to-noise ratio, and enhance the third audio signal based on the filter function.

[0122] The filter function is obtained based on the prior signal-to-noise ratio prediction value of the current signal frame, and can be expressed as follows:

[0123] Among them, H k (n) is the filter function;

[0124] Then, gain control is applied to the third audio signal based on the filter function and the posterior signal-to-noise ratio. The formulaic expression for obtaining the gain based on the filter function and the posterior signal-to-noise ratio is as follows:

[0125]

[0126] pass Gain control of the third audio signal yields the enhanced audio signal described above. This gain control may include adjusting the enhanced audio signal to achieve a preset volume value, and amplifying the human voice signal to make the human voice clearer when the enhanced audio signal is played. The second volume value may be 45 dB, 50 dB, or 60 dB, etc.

[0127] Figure 6 This is a block diagram illustrating a speech enhancement device according to an exemplary embodiment. (Refer to...) Figure 6 The device 600 includes:

[0128] The first acquisition module 610 is used to, in response to detecting that a user is wearing headphones, emit prompt audio data through the speaker in the headphones and acquire audio data in the ear canal;

[0129] The occlusion effect curve acquisition module 620 is used to calculate the frequency response curve based on the prompt audio data and the audio data in the ear canal, and to determine the target occlusion effect curve based on the frequency response curve.

[0130] The second acquisition module 630 is used to acquire the first sound signal outside the ear canal and the second sound signal inside the ear canal;

[0131] The filtering module 640 is used to filter the second sound signal according to the target occlusion effect curve to obtain the third sound signal;

[0132] The voice enhancement module is used to calculate a noise signal based on the first sound signal, and enhance the third sound signal based on the noise signal to obtain an enhanced sound signal.

[0133] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0134] Figure 7 This is a block diagram illustrating a device 800 according to an exemplary embodiment. For example, device 800 may be a true wireless stereo (TWS) Bluetooth headset, a semi-in-ear Bluetooth headset, an in-ear Bluetooth headset, a headband Bluetooth headset, a bone conduction Bluetooth headset, etc.

[0135] Reference Figure 7 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0136] Processing component 802 typically controls the overall operation of device 800, such as audio acquisition, audio playback, data communication (e.g., audio transmission, mode switching instructions, on / off instructions, volume control instructions, etc.), audio signal processing to obtain frequency response curves, filter functions, and related operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components.

[0137] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0138] Power supply component 806 provides power to various components of device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 800.

[0139] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes three microphones (a call microphone, a feedforward microphone, and a feedback microphone), which are configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes one or more speakers for outputting audio signals.

[0140] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, volume buttons, power buttons, and lock buttons.

[0141] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in position of device 800 or a component of device 800, and the presence or absence of user contact with device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0142] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0143] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0144] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to perform the above method. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, and optical data storage device.

[0145] Figure 8 This is a block diagram illustrating an apparatus 900 according to an exemplary embodiment. For example, apparatus 900 may be provided as a server. (Refer to...) Figure 8The apparatus 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by memory 932 for storing instructions, such as application programs, that can be executed by the processing component 922. The application programs stored in memory 932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 922 is configured to execute instructions to perform the methods described above.

[0146] The device 900 may also include a power supply component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output (I / O) interface 958. The device 900 can operate on an operating system stored in memory 932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0147] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0148] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A speech enhancement method, characterized in that, include: In response to detecting that a user is wearing headphones, a prompt audio data is emitted through the speaker in the headphones, and audio data inside the ear canal is collected; The frequency response curve is calculated based on the prompt audio data and the audio data inside the ear canal, and the target occlusion effect curve is determined based on the frequency response curve. The system collects a first sound signal from outside the ear canal and a second sound signal from inside the ear canal. The first sound signal is collected by the microphone. The second sound signal is filtered according to the target occlusion effect curve to obtain the third sound signal; A noise signal is calculated based on the first sound signal, and the third sound signal is enhanced based on the noise signal to obtain an enhanced sound signal.

2. The method according to claim 1, characterized in that, The step of calculating the frequency response curve based on the prompt audio data and the ear canal audio data specifically includes: The time-domain signal corresponding to the prompt audio data is calculated and Fourier transformed to generate the prompt audio domain signal. The time-domain signal of the audio data in the ear canal is then Fourier transformed to generate the frequency domain signal in the ear canal. The cross power spectrum of the prompt audio domain signal and the ear canal frequency domain signal is calculated based on the prompt audio domain signal and the ear canal frequency domain signal, and the auto power spectrum of the prompt audio domain signal is also calculated. Divide the cross power spectrum by the self power spectrum to obtain the frequency response curve.

3. The method according to claim 1, characterized in that, After the step of calculating the noise signal based on the first sound signal, the method further includes: The preset passive noise reduction curve is multiplied point-to-point with the noise signal to obtain the leakage noise signal.

4. The method according to claim 1, characterized in that, The step of determining the target blocking effect curve based on the frequency response curve specifically includes: Obtain a preset mapping table, wherein the mapping table includes the correspondence between a preset frequency response curve and a preset occlusion effect curve; The preset occlusion effect curve corresponding to the frequency response curve in the mapping table is determined as the target occlusion effect curve.

5. The method according to claim 1, characterized in that, The step of filtering the second sound signal according to the target occlusion effect curve to obtain the third sound signal specifically includes: The target occlusion effect curve is convolved with the second sound signal to obtain the third sound signal.

6. The method according to claim 1, characterized in that, The step of enhancing the third audio signal based on the noise signal specifically includes: Obtain a preset reference power spectrum, wherein the reference power spectrum is the power spectrum corresponding to a signal that contains only human voice; Divide the reference power spectrum by the power spectrum of the noise signal to obtain the posterior signal-to-noise ratio; The a priori signal-to-noise ratio prediction value of the current signal frame is obtained by calculating the posterior signal-to-noise ratio and the prior signal-to-noise ratio of the previous signal frame. A filter function is determined based on the prior signal-to-noise ratio, and the third audio signal is enhanced based on the filter function.

7. The method according to claim 6, characterized in that, The formulaic expression for the posterior signal-to-noise ratio is as follows: ,in, Let n be the a posteriori signal-to-noise ratio, n be the signal frame, and k be the frequency point. The power spectrum of the third sound signal. The power spectrum of the noise signal; The formulaic expression for the prior signal-to-noise ratio is as follows: ,in, Let be the prior signal-to-noise ratio. The reference power spectrum is given.

8. The method according to claim 7, characterized in that, The formula for calculating the predicted prior signal-to-noise ratio (SNR) of the current signal frame from the posterior SNR and the prior SNR of the previous signal frame is expressed as follows: ,in, As a smoothing factor, This is the prior signal-to-noise ratio prediction value. The reference power spectrum prediction value of the previous frame is given by the current signal frame. This is the power spectrum of the noise signal in the frame preceding the current signal frame. To select The maximum value between 0 and 0.

9. The method according to claim 8, characterized in that, The step of determining the filter function based on the prior signal-to-noise ratio and enhancing the third audio signal based on the filter function specifically includes: The filter function is obtained based on the prior signal-to-noise ratio prediction value of the current signal frame, and can be expressed as follows: ,in, The filter function is described above. Gain control is applied to the third audio signal based on the filter function and the posterior signal-to-noise ratio to enhance the third audio signal and obtain the enhanced audio signal.

10. A speech enhancement device, characterized in that, include: The first acquisition module is used to respond to the detection that the user is wearing headphones by emitting prompt audio data through the speaker in the headphones and acquiring audio data in the ear canal; The occlusion effect curve acquisition module is used to calculate the frequency response curve based on the prompt audio data and the audio data in the ear canal, and to determine the target occlusion effect curve based on the frequency response curve. The second acquisition module is used to acquire the first sound signal outside the ear canal and the second sound signal inside the ear canal. The first sound signal is acquired by the call microphone. The filtering module is used to filter the second sound signal according to the target occlusion effect curve to obtain the third sound signal; The voice enhancement module is used to calculate a noise signal based on the first sound signal, and enhance the third sound signal based on the noise signal to obtain an enhanced sound signal.

11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of an electronic device, enable the electronic device to perform the method as described in any one of claims 1 to 9.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Signal processing device, signal processing method, and program

    CN107431852A

  • Data processing method and related device

    CN114157945A