Audio signal processing method, device, readable storage medium and earphone
By filtering ambient sound signal and extracting vocal signals, the problem of noise amplification when vocal signals are enhanced in the prior art is solved, and clear vocal perception and ambient sound perception in transparent mode are achieved, which improves the user experience.
Patent Information
- Application Number
- CN202111093716.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-09-17
AI Technical Summary
When the prior art enhances the vocal signal, it is difficult to effectively distinguish and enhance the vocal signal without amplifying the noise signal, resulting in the noise and vocal enhancement in the user experience and may damage the voice signal.
By obtaining the ambient sound signal, using a preset permeable filter for filtering, extracting the voice signal in the ambient sound signal, and sending it to the speaker to play with the filtered ambient sound signal.
Without losing voice, the vocal part of the transparent ambient sound signal is enhanced to provide clear vocal perception, improve the user's environmental transparency experience, and be able to pay attention to important ambient sounds to ensure personal safety.
Smart Images

Figure CN113810828B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of audio processing, and in particular, to an audio signal processing method, apparatus, readable storage medium, and earphone. Background Art
[0002] To adapt to different scenarios, many existing earphones are provided with a noise reduction mode and a transparency mode. The noise reduction mode is used to block external sound signals, and the transparency mode is used to allow external sound signals to enter the human ear. When a user wears earphones and wants to talk to others, they can switch to the transparency mode without removing the earphones, which is equivalent to the effect of removing the earphones, thus achieving clear conversation with the other party. However, there is usually noise in the environment. We expect to hear more human voices and less noise during conversation, so a human voice enhancement function that makes the human voice clearer is added.
[0003] Since the voice frequency band ranges from 300 Hz to 3,400 Hz, the current human voice enhancement methods usually process sounds in different frequency bands separately: for low-frequency noise below 300 Hz, reverse sound waves are applied to cancel it; for the voice frequency band of 300 Hz to 3,400 Hz, a transparency filter is used for filtering, and then a band-pass filter is used to filter and amplify the energy of the noise-reduced voice; finally, the reverse sound wave is superimposed on the band-pass amplified sound wave and played by the speaker. However, in the actual environment, the noise is distributed across the entire frequency band, and there will also be noise in the 300 Hz to 3,400 Hz frequency band. The noise will be amplified while the voice is amplified, and the actual user experience is that only the noise and the human voice are enhanced together. Moreover, there may be voices below 300 Hz, and the voice may be damaged by sound wave cancellation. Summary of the Invention
[0004] To overcome the problems in the related art, the present disclosure provides an audio signal processing method, apparatus, readable storage medium, and earphone.
[0005] According to a first aspect of an embodiment of the present disclosure, an audio signal processing method is provided, which is applied to an earphone and includes:
[0006] Obtain an ambient sound signal, where the ambient sound signal is a sound signal in the environment around the earphone;
[0007] Perform filtering processing on the ambient sound signal according to a preset transparency filter to obtain a first audio signal;
[0008] Extract a human voice signal from the ambient sound signal to obtain a second audio signal;
[0009] Send the first audio signal and the second audio signal to a speaker, and control the speaker to synchronously play the first audio signal and the second audio signal.
[0010] Optionally, extracting the human voice signal from the ambient sound signal includes:
[0011] Extracting the human voice signal from the ambient sound signal through Wiener filtering.
[0012] Optionally, extracting the human voice signal from the ambient sound signal through Wiener filtering includes:
[0013] Transforming the ambient sound signal from the time domain to the frequency domain through Fourier transform to obtain the frequency-domain signal corresponding to the ambient sound signal;
[0014] For each audio frame in the frequency-domain signal, determining the Wiener filtering coefficient corresponding to the audio frame;
[0015] Filtering the audio frame using the Wiener filtering coefficient corresponding to the audio frame to obtain the frequency-domain human voice signal in the audio frame;
[0016] Performing inverse Fourier transform on the frequency-domain human voice signal to obtain the human voice signal in the time-domain signal corresponding to the audio frame.
[0017] Optionally, determining the Wiener filtering coefficient corresponding to the audio frame includes:
[0018] Determining the power spectrum corresponding to the audio frame and performing noise estimation on the audio frame to obtain the power spectrum corresponding to the noise signal in the audio frame;
[0019] Determining the Wiener filtering coefficient corresponding to the audio frame according to the power spectrum corresponding to the audio frame, the power spectrum corresponding to the noise signal in the audio frame, and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame.
[0020] Optionally, determining the Wiener filtering coefficient corresponding to the audio frame according to the power spectrum corresponding to the audio frame, the power spectrum corresponding to the noise signal in the audio frame, and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame includes:
[0021] Determining the posterior signal-to-noise ratio corresponding to the audio frame according to the power spectrum corresponding to the audio frame and the power spectrum corresponding to the noise signal in the audio frame;
[0022] Determining the prior signal-to-noise ratio estimation value corresponding to the audio frame according to the posterior signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame;
[0023] Generating the Wiener filtering coefficient corresponding to the audio frame according to the prior signal-to-noise ratio estimation value.
[0024] Optionally, determining the estimated value of the a priori signal-to-noise ratio corresponding to the audio frame according to the a posteriori signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame includes:
[0025] Determining the estimated value of the a priori signal-to-noise ratio corresponding to the audio frame according to the a posteriori signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame through the following formula:
[0026]
[0027] Wherein, is the estimated value of the a priori signal-to-noise ratio corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal, n>1; P dd (n - 1) is the power spectrum corresponding to the noise signal in the (n - 1)th audio frame in the frequency-domain signal corresponding to the ambient sound signal; α is a weight coefficient; is the power spectrum corresponding to the frequency-domain human voice signal in the (n - 1)th audio frame in the frequency-domain signal corresponding to the ambient sound signal; γ(n) is the a posteriori signal-to-noise ratio corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal.
[0028] Optionally, extracting the human voice signal in the ambient sound signal to obtain a second audio signal includes:
[0029] Inputting the ambient sound signal into a pre-trained human voice extraction model to obtain a second audio signal.
[0030] Optionally, the preset transparent filter is a plurality of cascaded second-order IIR filters.
[0031] According to the second aspect of the embodiments of the present disclosure, there is provided an audio signal processing device applied to a headset, including:
[0032] An acquisition module configured to acquire an ambient sound signal, wherein the ambient sound signal is a sound signal in the environment around the headset;
[0033] A filtering module configured to perform filtering processing on the ambient sound signal acquired by the acquisition module according to a preset transparent filter to obtain a first audio signal;
[0034] An extraction module configured to extract the human voice signal in the ambient sound signal acquired by the acquisition module to obtain a second audio signal;
[0035] A playback control module configured to send the first audio signal and the second audio signal to a speaker and control the speaker to synchronously play the first audio signal and the second audio signal.
[0036] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the audio signal processing method provided in the first aspect of the present disclosure are implemented.
[0037] According to a fourth aspect of the embodiments of the present disclosure, there is provided a headset, comprising:
[0038] a controller, a feedforward microphone and a speaker communicatively connected to the controller;
[0039] the feedforward microphone is configured to collect an ambient sound signal and send the ambient sound signal to the controller;
[0040] the speaker is configured to play an audio signal according to a control instruction of the controller;
[0041] the controller includes a memory and a processor, and a computer program capable of running on the processor is stored on the memory; when the processor is used to run the computer program, the audio signal processing method provided in the first aspect of the present disclosure is executed.
[0042] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: filtering the ambient sound signal according to a preset transparency filter to obtain a first audio signal, and simultaneously extracting a human voice signal from the ambient sound signal to obtain a second audio signal; then, sending the first audio signal and the second audio signal to the speaker, and controlling the speaker to play the first audio signal and the second audio signal synchronously. In this way, by extracting the human voice signal in the ambient sound signal, superimposing the filtered ambient sound signal, and through synchronous playback, the human voice part in the transparent ambient sound signal can be enhanced without losing the voice, providing a clear human voice perception in the transparency mode and providing a good ambient transparency experience for the user. Thus, the user does not need to take off the headset and can clearly hear the other person speaking in the transparency mode, and can also notice important ambient sounds including car whistles and traffic sounds, ensuring personal safety. In addition, since the superimposed is the human voice signal, the noise will not be enhanced, making the human voice heard by the user clearer.
[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0045] Figure 1 is a flowchart of an audio signal processing method shown according to an exemplary embodiment.
[0046] Figure 2 It is a schematic diagram showing the frequency response curve A during bone conduction and the frequency response curve B after wearing the earphone with passive noise reduction according to an exemplary embodiment.
[0047] Figure 3 It is a schematic diagram showing a target curve C according to an exemplary embodiment.
[0048] Figure 4 It is a schematic diagram showing a target curve C and a frequency response curve D according to an exemplary embodiment.
[0049] Figure 5 It is a flowchart of a method for extracting human voice signals from environmental sound signals through Wiener filtering according to an exemplary embodiment.
[0050] Figure 6 It is a block diagram of an audio signal processing device according to an exemplary embodiment.
[0051] Figure 7 It is a block diagram of an audio signal processing device according to an exemplary embodiment. Detailed implementation manners
[0052] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0053] Figure 1 It is a flowchart of an audio signal processing method according to an exemplary embodiment, where the method can be applied to an earphone. As Figure 1 shown, the method includes the following S101 to S104.
[0054] In S101, an environmental sound signal is acquired.
[0055] In the present disclosure, the environmental sound signal is a sound signal in the environment around the earphone. The above earphone includes a noise reduction mode and a transparency mode, where the noise reduction mode is used to block external sound signals, and the transparency mode is used to let external sound signals enter the human ear. Among them, the above audio signal processing method can be applied to the scenario where the earphone is in the transparency mode.
[0056] Exemplarily, the environmental sound signal can be acquired through a feedforward microphone on the earphone.
[0057] In S102, the ambient sound signal is filtered according to a preset transparent filter to obtain a first audio signal.
[0058] In the present disclosure, when the user wears the earphone, the passive sound insulation of the earphone will bring about the internal and external sound pressure level difference of the ambient sound signal. When the earphone is in the transparent mode, the ambient sound signal can be filtered by a preset transparent filter arranged in the earphone after entering the earphone, so as to make up for the internal and external sound pressure level difference of the ambient sound signal caused by the passive sound insulation of the earphone, enabling the user to experience the input experience of the sound signal in a state close to not wearing the earphone. That is, after the transparent mode is turned on, the human ear's response to the outside world is an open response.
[0059] Exemplarily, the above preset transparent filter can be multiple (for example, 6) cascaded second-order IIR filters, and this transparent filter filters the ambient sound signal using the frequency response curve D to make up for the internal and external sound pressure level difference of the ambient sound signal caused by the passive sound insulation of the earphone.
[0060] In S103, the human voice signal in the ambient sound signal is extracted to obtain a second audio signal.
[0061] In S104, the first audio signal and the second audio signal are sent to the speaker, and the speaker is controlled to play the first audio signal and the second audio signal synchronously.
[0062] The technical solution provided by the embodiment of the present disclosure may include the following beneficial effects: The ambient sound signal is filtered according to a preset transparent filter to obtain a first audio signal, and at the same time, the human voice signal in the ambient sound signal is extracted to obtain a second audio signal; then, the first audio signal and the second audio signal are sent to the speaker, and the speaker is controlled to play the first audio signal and the second audio signal synchronously. In this way, by extracting the human voice signal in the ambient sound signal, superimposing the filtered ambient sound signal, and through synchronous playback, the human voice part in the transparent ambient sound signal can be enhanced without losing the voice, providing a clear human voice perception in the transparent mode and providing a good ambient transparency experience for the user. Thus, the user does not need to take off the earphone and can clearly hear the other person speaking in the transparent mode. Moreover, the user can also notice important ambient sounds including the whistle sound and the driving sound, ensuring personal safety. In addition, since the superimposed is the human voice signal, the noise will not be enhanced, making the human voice heard by the user clearer.
[0063] The following will detail the determination method of the frequency response curve D used when the above transparent filter (i.e., multiple cascaded second-order IIR filters) filters the ambient sound signal.
[0064] Specifically, the frequency response curve A when the artificial head collects sound without wearing headphones (i.e., without headphones) and the frequency response curve B after passive noise reduction when wearing headphones can be collected (as Figure 2 shown). By comparing the frequency response curve A and the frequency response curve B, the passive noise reduction curve, that is, the target curve C that we need to compensate, is obtained (as Figure 3 shown).
[0065] The target curve C is approximated by designing multiple cascaded second-order IIR filters. The specific design steps are as follows: First, the coefficients of each second-order IIR filter (including frequency, gain, Q value) are randomly initialized; then the coefficients of each second-order IIR filter are randomly updated, and the compensation curve E is calculated, and the difference between the compensation curve E and the target curve C is compared (for example, calculating the cosine distance between the two curves); if the difference between the compensation curve E and the target curve C is smaller than the previous update, then based on the current IIR coefficients, the coefficients of each second-order IIR filter are continued to be updated; if the difference between the compensation curve E and the target curve C is larger than the previous update, then based on the IIR coefficients obtained from the previous update, the coefficients of each second-order IIR filter are continued to be updated. And so on, multiple iterations are performed until the difference between the compensation curve E and the target curve C stabilizes (for example, the difference between the maximum value and the minimum value in the last 10 times is less than the preset threshold); after that, the compensation curve E obtained from the most recent update is used as the above-mentioned frequency response curve D. Exemplarily, the frequency response curve D and the target curve C are as Figure 4 shown.
[0066] The following will specifically describe the specific implementation manner of extracting the human voice signal from the ambient sound signal in S102 above.
[0067] In one implementation manner, the ambient sound signal can be input into a pre-trained human voice extraction model to obtain a second audio signal.
[0068] In the present disclosure, the above-mentioned human voice extraction model can be trained in the following manner:
[0069] First, a reference ambient sound signal and the human voice signal in the reference ambient sound signal are obtained;
[0070] Then, the model is trained by taking the reference ambient sound signal as the input of the human voice extraction model and the human voice signal in the reference ambient sound signal as the target output of the human voice extraction model, so as to obtain the above-mentioned human voice extraction model.
[0071] Exemplarily, the above-mentioned voice extraction model can be a Convolutional Neural Networks (CNN) + Long Short-Term Memory (LSTM), or a Recurrent Neural Network (RNN), etc.
[0072] In another implementation, the voice signal in the ambient sound signal can be extracted through Wiener filtering. Specifically, it can be implemented through Figure 5 S1031 to S1034 shown in
[0073] In S1031, the ambient sound signal is transformed from the time domain to the frequency domain through Fourier transform to obtain the frequency domain signal corresponding to the ambient sound signal.
[0074] In S1032, for each audio frame in the frequency domain signal, the Wiener filter coefficient corresponding to the audio frame is determined.
[0075] In S1033, the audio frame is filtered using the Wiener filter coefficient corresponding to the audio frame to obtain the frequency domain voice signal in the audio frame.
[0076] Exemplarily, the audio frame can be filtered using the Wiener filter coefficient corresponding to the audio frame through the following equation (1) to obtain the frequency domain voice signal in the audio frame:
[0077]
[0078] Where Y(n) is the nth audio frame in the frequency domain signal corresponding to the above-mentioned ambient sound signal, n = 1, 2, …, m, and m is the total number of audio frames included in the frequency domain signal corresponding to the above-mentioned ambient sound signal; is the frequency domain voice signal in the nth audio frame Y(n) in the frequency domain signal corresponding to the above-mentioned ambient sound signal; H(n) is the Wiener filter coefficient corresponding to the nth audio frame Y(n) in the frequency domain signal corresponding to the above-mentioned ambient sound signal.
[0079] In S1034, the frequency domain voice signal is subjected to inverse Fourier transform to obtain the voice signal in the time domain signal corresponding to the audio frame.
[0080] The following details the specific implementation of determining the Wiener filter coefficient corresponding to the audio frame in S1032 above. Specifically, it can be implemented through the following steps (1) and (2):
[0081] (1) Determine the power spectrum corresponding to the audio frame, and perform noise estimation on the audio frame to obtain the power spectrum corresponding to the noise signal in the audio frame.
[0082] Exemplarily, methods such as Voice Active Detection (VAD) and Minimum Controlled Regressive Averaging (MCRA) can be used to estimate the noise of the audio frame, so as to obtain the power spectrum corresponding to the noise signal in the audio frame.
[0083] (2) Determine the Wiener filter coefficient corresponding to the audio frame according to the power spectrum corresponding to the audio frame, the power spectrum corresponding to the noise signal in the audio frame, and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame.
[0084] Specifically, the Wiener filter coefficient corresponding to the audio frame can be determined through the following steps 1) to 3):
[0085] 1) Determine the posterior signal-to-noise ratio corresponding to the audio frame according to the power spectrum corresponding to the audio frame and the power spectrum corresponding to the noise signal in the audio frame.
[0086] Exemplarily, the posterior signal-to-noise ratio corresponding to the audio frame can be determined according to the power spectrum corresponding to the audio frame and the power spectrum corresponding to the noise signal in the audio frame through the following equation (2):
[0087]
[0088] where γ(n) is the posterior signal-to-noise ratio corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal; P yy () is the power spectrum corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal; P dd () is the power spectrum corresponding to the noise signal in the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal.
[0089] 2) Determine the estimated value of the prior signal-to-noise ratio corresponding to the audio frame according to the posterior signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame.
[0090] Exemplarily, the estimated value of the prior signal-to-noise ratio corresponding to the audio frame can be determined according to the posterior signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame through the following equation (3):
[0091]
[0092] where, is the estimated value of the prior signal-to-noise ratio corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal, n > 1; P dd(n - 1) is the power spectrum corresponding to the noise signal in the (n - 1)-th audio frame of the frequency-domain signal corresponding to the above ambient sound signal; α is a weight coefficient; is the power spectrum corresponding to the frequency-domain human voice signal in the (n - 1)-th audio frame of the frequency-domain signal corresponding to the above ambient sound signal, And,
[0093] 3) Generate the Wiener filter coefficient corresponding to this audio frame according to the prior SNR estimate value.
[0094] Exemplarily, the Wiener filter coefficient corresponding to this audio frame can be generated according to the prior SNR estimate value through the following equation (4):
[0095]
[0096] Based on the same inventive concept, the present disclosure also provides an audio signal processing device. As Figure 6 shown, the device 600 includes:
[0097] An acquisition module 601, configured to acquire an ambient sound signal, where the ambient sound signal is a sound signal in the environment around the earphone;
[0098] A filtering module 602, configured to perform filtering processing on the ambient sound signal acquired by the acquisition module according to a preset transparent filter to obtain a first audio signal;
[0099] An extraction module 603, configured to extract the human voice signal in the ambient sound signal acquired by the acquisition module 601 to obtain a second audio signal;
[0100] A playback control module 604, configured to send the first audio signal and the second audio signal to a speaker, and control the speaker to synchronously play the first audio signal and the second audio signal.
[0101] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: Filter the ambient sound signal according to a preset transparency filter to obtain a first audio signal, and at the same time extract the human voice signal in the ambient sound signal to obtain a second audio signal; then, send the first audio signal and the second audio signal to the speaker, and control the speaker to play the first audio signal and the second audio signal synchronously. In this way, by extracting the human voice signal in the ambient sound signal, superimposing the filtered ambient sound signal, and through synchronous playback, the human voice part in the transparent ambient sound signal can be enhanced without losing the voice, providing a clear human voice perception in the transparency mode and providing a good ambient transparency experience for the user. Therefore, the user does not need to take off the headphones and can clearly hear the other person speaking in the transparency mode. Moreover, the user can also notice important ambient sounds including the sound of a siren and the sound of a moving vehicle, ensuring personal safety. In addition, since the superimposed is the human voice signal, the noise will not be enhanced, making the human voice heard by the user clearer.
[0102] Optionally, the extraction module 603 is configured to extract the human voice signal in the ambient sound signal through Wiener filtering.
[0103] Optionally, the extraction module 603 includes:
[0104] A first transformation sub-module, configured to transform the ambient sound signal from the time domain to the frequency domain through Fourier transform to obtain the frequency domain signal corresponding to the ambient sound signal;
[0105] A first determination sub-module, configured to determine the Wiener filtering coefficient corresponding to each audio frame in the frequency domain signal;
[0106] A filtering sub-module, configured to filter each audio frame by using the Wiener filtering coefficient corresponding to the audio frame to obtain the frequency domain human voice signal in the audio frame;
[0107] A second transformation sub-module, configured to perform inverse Fourier transform on the frequency domain human voice signal to obtain the human voice signal in the time domain signal corresponding to the audio frame.
[0108] Optionally, the first determination sub-module includes:
[0109] A second determination sub-module, configured to determine the power spectrum corresponding to the audio frame and perform noise estimation on the audio frame to obtain the power spectrum corresponding to the noise signal in the audio frame;
[0110] A third determination sub-module, configured to determine the Wiener filtering coefficient corresponding to the audio frame according to the power spectrum corresponding to the audio frame, the power spectrum corresponding to the noise signal in the audio frame, and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame.
[0111] Optionally, the third determination sub-module includes:
[0112] A fourth determination sub-module, configured to determine a posteriori signal-to-noise ratio corresponding to the audio frame according to the power spectrum corresponding to the audio frame and the power spectrum corresponding to the noise signal in the audio frame;
[0113] A fifth determination sub-module, configured to determine an a priori signal-to-noise ratio estimation value corresponding to the audio frame according to the a posteriori signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame;
[0114] A generation sub-module, configured to generate a Wiener filter coefficient corresponding to the audio frame according to the a priori signal-to-noise ratio estimation value.
[0115] Optionally, the fifth determination sub-module is configured to determine an a priori signal-to-noise ratio estimation value corresponding to the audio frame according to the a posteriori signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame through the following formula:
[0116]
[0117] Wherein, is the a priori signal-to-noise ratio estimation value corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal, n>1; P dd (n - 1) is the power spectrum corresponding to the noise signal in the (n - 1)th audio frame in the frequency-domain signal corresponding to the ambient sound signal; α is a weight coefficient; is the power spectrum corresponding to the frequency-domain human voice signal in the (n - 1)th audio frame in the frequency-domain signal corresponding to the ambient sound signal; γ(n) is the a posteriori signal-to-noise ratio corresponding to the nth audio frame in the frequency-domain signal corresponding to the ambient sound signal.
[0118] Optionally, the extraction module 603 is configured to input the ambient sound signal into a pre-trained human voice extraction model to obtain a second audio signal.
[0119] Optionally, the preset transparent filter is a plurality of cascaded second-order IIR filters.
[0120] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0121] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the above audio signal processing method provided by the present disclosure are implemented.
[0122] The present disclosure also provides a headset, including:
[0123] A controller, a feedforward microphone, and a speaker communicatively connected to the controller;
[0124] A feedforward microphone for collecting ambient sound signals and sending the ambient sound signals to the controller;
[0125] A speaker for playing an audio signal according to a control instruction of the controller;
[0126] The controller includes a memory and a processor. A computer program capable of running on the processor is stored on the memory; when the processor is used to run the computer program, it executes the audio signal processing method provided in the first aspect of the present disclosure.
[0127] Figure 7 It is a block diagram of an audio signal processing apparatus 800 shown according to an exemplary embodiment. For example, the apparatus 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0128] Referring to Figure 7 , the apparatus 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0129] The processing component 802 generally controls the overall operation of the apparatus 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned audio signal processing method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0130] The memory 804 is configured to store various types of data to support the operation of the apparatus 800. Examples of these data include instructions for any application program or method operating on the apparatus 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0131] The power component 806 provides power for various components of the device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 800.
[0132] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0133] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0134] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0135] The sensor assembly 814 includes one or more sensors for providing an assessment of the status of the device 800 in various aspects. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0136] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0137] In an exemplary embodiment, the device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above audio signal processing method.
[0138] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by the processor 820 of the device 800 to complete the above audio signal processing method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0139] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a programmable device. The computer program has a code portion for performing the above audio signal processing method when executed by the programmable device.
[0140] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0141] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio signal processing method, characterized in that, applied to headphones, includes: Obtain an ambient sound signal, where the ambient sound signal is a sound signal in the environment around the headphones; Filter the ambient sound signal according to a preset transparent filter to obtain a first audio signal, where the preset transparent filter is a plurality of cascaded second-order IIR filters; Extract the human voice signal in the ambient sound signal to obtain a second audio signal; Send the first audio signal and the second audio signal to a speaker, and control the speaker to synchronously play the first audio signal and the second audio signal; Among them, the plurality of cascaded second-order IIR filters use a frequency response curve D to filter the ambient sound signal, and the frequency response curve D is determined by the following method: Randomly initialize the coefficients of each second-order IIR filter; Randomly update the coefficients of each second-order IIR filter; Calculate a compensation curve E based on the updated coefficients; Calculate a first difference between the compensation curve E and a target curve C, where the target curve C is determined based on the frequency response curve A when listening without earphones and the frequency response curve B after passive noise reduction when wearing the headphones; If the first difference is less than or equal to a second difference, then based on the updated coefficients, return to the step of randomly updating the coefficients of each second-order IIR filter until the difference between the compensation curve E and the target curve C reaches a preset stable condition, where the second difference is calculated based on the coefficients after the previous update; If the first difference is greater than the second difference, then based on the coefficients after the previous update, return to the step of randomly updating the coefficients of each second-order IIR filter until the difference between the compensation curve E and the target curve C reaches the stable condition; If the difference between the compensation curve E and the target curve C reaches the stable condition, then determine the compensation curve E calculated most recently as the frequency response curve D.
2. The method according to claim 1, characterized in that, The extracting the human voice signal in the ambient sound signal includes: Extract the human voice signal in the ambient sound signal through Wiener filtering.
3. The method according to claim 2, characterized in that, The extracting the human voice signal in the ambient sound signal through Wiener filtering includes: Transform the ambient sound signal from the time domain to the frequency domain through Fourier transform to obtain a frequency domain signal corresponding to the ambient sound signal; For each audio frame in the frequency domain signal, determine the Wiener filtering coefficient corresponding to the audio frame; Filter the audio frame using the Wiener filtering coefficient corresponding to the audio frame to obtain a frequency domain human voice signal in the audio frame; Perform an inverse Fourier transform on the frequency domain human voice signal to obtain the human voice signal in the time domain signal corresponding to the audio frame.
4. The method according to claim 3, characterized in that, The determining the Wiener filtering coefficient corresponding to the audio frame includes: Determine the power spectrum corresponding to the audio frame, and perform noise estimation on the audio frame to obtain the power spectrum corresponding to the noise signal in the audio frame; Determine the Wiener filter coefficient corresponding to the audio frame according to the power spectrum corresponding to the audio frame, the power spectrum corresponding to the noise signal in the audio frame, and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame.
5. The method according to claim 4, characterized in that, The determining the Wiener filter coefficient corresponding to the audio frame according to the power spectrum corresponding to the audio frame, the power spectrum corresponding to the noise signal in the audio frame, and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame includes: Determine the posterior signal-to-noise ratio corresponding to the audio frame according to the power spectrum corresponding to the audio frame and the power spectrum corresponding to the noise signal in the audio frame; Determine the estimated value of the prior signal-to-noise ratio corresponding to the audio frame according to the posterior signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame; Generate the Wiener filter coefficient corresponding to the audio frame according to the estimated value of the prior signal-to-noise ratio.
6. The method according to claim 5, characterized in that, The determining the estimated value of the prior signal-to-noise ratio corresponding to the audio frame according to the posterior signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame includes: Determine the estimated value of the prior signal-to-noise ratio corresponding to the audio frame according to the posterior signal-to-noise ratio and the power spectrum corresponding to the noise signal in the previous audio frame of the audio frame through the following formula: wherein, is the a priori signal-to-noise ratio estimate corresponding to the n th audio frame in the frequency-domain signal corresponding to the ambient sound signal, n > 1; is the power spectrum of the noise signal in the n -1 audio frames in the frequency-domain signal corresponding to the ambient sound signal; is the weight coefficient; is the power spectrum of the frequency-domain human voice signal in the n -1 audio frames in the frequency-domain signal corresponding to the ambient sound signal; is the a posteriori signal-to-noise ratio corresponding to the n th audio frame in the frequency-domain signal corresponding to the ambient sound signal.
7. The method according to claim 1, characterized in that, The extracting the human voice signal from the environmental sound signal to obtain a second audio signal includes: Input the environmental sound signal into a pre-trained human voice extraction model to obtain a second audio signal.
8. An audio signal processing device, characterized in that, Applied to a headset, including: An acquisition module configured to acquire an environmental sound signal, where the environmental sound signal is a sound signal in the environment around the headset; A filtering module configured to perform filtering processing on the environmental sound signal acquired by the acquisition module according to a preset transparent filter to obtain a first audio signal, where the preset transparent filter is a plurality of cascaded second-order IIR filters; An extraction module configured to extract the human voice signal from the environmental sound signal acquired by the acquisition module to obtain a second audio signal; A playback control module configured to send the first audio signal and the second audio signal to a speaker, and control the speaker to synchronously play the first audio signal and the second audio signal; Wherein, the plurality of cascaded second-order IIR filters perform filtering processing on the environmental sound signal using a frequency response curve D, and the frequency response curve D is determined by the following method: Randomly initialize the coefficients of each of the second-order IIR filters; Randomly update the coefficients of each of the second-order IIR filters; Calculate a compensation curve E based on the updated coefficients; Calculate a first gap between the compensation curve E and the target curve C, where the target curve C is determined based on the frequency response curve A during binaural beats and the frequency response curve B after passive noise reduction with the headphones on; If the first gap is less than or equal to a second gap, then based on the updated coefficient, return to the step of randomly updating the coefficients of each of the second-order IIR filters until the gap between the compensation curve E and the target curve C reaches a preset stable condition, where the second gap is calculated based on the coefficients after the previous update; If the first gap is greater than the second gap, then based on the coefficients after the previous update, return to the step of randomly updating the coefficients of each of the second-order IIR filters until the gap between the compensation curve E and the target curve C reaches the stable condition; If the gap between the compensation curve E and the target curve C reaches the stable condition, then determine the most recently calculated compensation curve E as the frequency response curve D.
9. A computer-readable storage medium, on which computer program instructions are stored, characterized in that, when the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A pair of headphones, characterized in that, comprising: a controller, a feedforward microphone and a speaker communicatively connected to the controller; the feedforward microphone, configured to collect an ambient sound signal and send the ambient sound signal to the controller; the speaker, configured to play an audio signal according to a control instruction of the controller; the controller, including a memory and a processor, where a computer program capable of running on the processor is stored on the memory; when the processor is used to run the computer program, the audio signal processing method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Hearing-aid multichannel voice enhancing algorithm based on iterative Wiener filtering
CN109961799A
Sound processing method, device and equipment
CN110970057A
Transparent mode adjusting device and method of wireless earphone and wireless earphone
CN111885460A