A wide dynamic range compression method and apparatus based on adaptive noise detection
By using adaptive noise detection and spectral flatness assessment, and setting a fixed gain factor for the noise channel, the problem that noise signal amplification affects hearing aid speech recognition in existing technologies is solved, resulting in better speech recognition performance and hearing protection.
Patent Information
- Application Number
- CN202411949052.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing wide dynamic range compression algorithms, when the signal is noise, will amplify the channel signal when the sound pressure level in the channel exceeds the hearing threshold, thus greatly reducing the speech recognition performance of the hearing aid.
An adaptive noise detection method is used to convert the input audio time-domain signal into a frequency-domain frame signal, and the signal attributes are determined based on the spectral flatness. A fixed gain factor is set for the noise channel to adjust the gain and avoid noise amplification, while the voice channel is processed by calculating the gain factor based on the user's hearing curve.
It improves the speech recognition performance of hearing aids, avoids the damage to hearing caused by unnecessary amplification of noise signals, and enhances the clarity of speech signals.
Smart Images

Figure CN119943073B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic equipment technology, and in particular to a wide dynamic range compression method and apparatus based on adaptive noise detection. Background Technology
[0002] Wide Dynamic Range Compression (WDRC) is one of the core algorithms in digital hearing aids. Its main function is to compensate for sounds of different frequencies and intensities to varying degrees, thereby converting ambient sounds into the range of hearing loss.
[0003] The wide dynamic range compression algorithm first divides the speech into multiple independent frequency channels in the frequency domain. Within each channel, based on the audiometry threshold diagram of the hearing-impaired patient, it processes the data independently, mapping the dynamic range of normal sound to the hearing range of the patient, thereby determining the gain of that channel. Then, the gains calculated for different channels are applied to the input signal in the frequency domain, and finally, the sound signals are synthesized and output. One parameter of the patient's audiometry threshold information is the hearing threshold value; when the sound pressure level of the signal in a channel exceeds this value, the signal in that channel will be amplified accordingly.
[0004] However, wide dynamic range compression algorithms process speech and noise indiscriminately. When the signal is noise, if the sound pressure level of the signal in the channel is greater than the hearing threshold, the channel signal will also be amplified, which greatly reduces the speech recognition effect of the hearing aid.
[0005] Therefore, existing technologies have shortcomings and need to be improved and developed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a wide dynamic range compression method and device based on adaptive noise detection, which addresses the above-mentioned defects of the prior art. The aim is to solve the problem that in the prior art, when the signal is noise, the channel signal will also be amplified when the signal sound pressure value in the channel is greater than the hearing threshold, thereby greatly reducing the speech recognition effect of the hearing aid.
[0007] The technical solution adopted by this invention to solve the technical problem is as follows:
[0008] A wide dynamic range compression method based on adaptive noise detection, wherein the method includes:
[0009] The input audio time-domain signal is converted into several frame signals in the frequency domain, and each frame signal is divided into several channel signals;
[0010] If the current frame signal is a speech frame, then determine the attribute information of several channel signals corresponding to the current frame signal;
[0011] When the attribute information of the channel signal is a noise channel, a preset fixed gain factor is used as the gain factor of the channel signal for gain adjustment to obtain the adjusted spectrum sequence of the channel signal.
[0012] Based on the adjusted spectral sequences, channel synthesis is performed on the signals of each channel, and the synthesized time-domain speech signal is output.
[0013] In one embodiment of this application, after converting the input audio time-domain signal into several frame signals in the frequency domain, and dividing each frame signal into several channel signals, the method further includes:
[0014] If the current frame signal is a noisy frame, then a preset fixed gain factor is used as the gain factor of the noisy frame for gain adjustment, so as to obtain the adjusted spectrum sequence corresponding to each channel signal of the noisy frame.
[0015] In one embodiment of this application, if the current frame signal is a speech frame, before determining the attribute information of the several channel signals corresponding to the current frame signal, the method further includes:
[0016] Calculate the first spectral flatness of the current frame signal, and determine the attribute information of the current frame signal based on the first spectral flatness. The attribute information of the current frame signal is either a speech frame or a noise frame.
[0017] In one embodiment of this application, if the current frame signal is a speech frame, then the attribute information of several channel signals corresponding to the current frame signal is determined, including:
[0018] If the current frame signal is a speech frame, then calculate the second spectral flatness corresponding to each channel signal in the current frame signal;
[0019] Obtain the noise frame flatness corresponding to the initial noise frame in the currently input audio time domain signal, and calculate the difference between the second spectral flatness and the noise frame flatness;
[0020] If the difference between the second spectral flatness and the noise frame flatness is less than a preset threshold, then the attribute information of the channel signal is a noise channel.
[0021] In one embodiment of this application, after obtaining the noise frame flatness corresponding to the initial noise frame in the currently input audio time-domain signal and calculating the difference between the second spectral flatness and the noise frame flatness, the method further includes:
[0022] If the difference between the second spectral flatness and the noise frame flatness is greater than or equal to a preset threshold, then the attribute information of the channel signal is a voice channel.
[0023] In one embodiment of this application, if the current frame signal is a speech frame, after determining the attribute information of several channel signals corresponding to the current frame signal, the method further includes:
[0024] When the attribute information of the channel signal is a voice channel, the gain factor of the voice channel is calculated based on the preset user hearing curve function and the input sound level of the voice channel.
[0025] The gain of the voice channel is adjusted according to the gain factor of the voice channel to obtain the adjusted spectrum sequence.
[0026] In one embodiment of this application, the preset fixed gain factor is less than or equal to 1, and the preset fixed gain factor can be customized.
[0027] This application also provides a wide dynamic range compression device based on adaptive noise detection, wherein the device includes:
[0028] The channel division module is used to convert the input audio time-domain signal into several frame signals in the frequency domain, and divide each frame signal into several channel signals;
[0029] The attribute determination module is used to determine the attribute information of several channel signals corresponding to the current frame signal if the current frame signal is a speech frame.
[0030] The gain adjustment module is used to adjust the gain of the channel signal by using a preset fixed gain factor as the gain factor of the channel signal when the attribute information of the channel signal is a noise channel, so as to obtain the adjusted spectrum sequence of the channel signal.
[0031] The channel synthesis module is used to synthesize the channel signals based on the adjusted spectral sequences and output the synthesized time-domain speech signal.
[0032] This application also provides a terminal, comprising: a memory, a processor, and a wide dynamic range compression program based on adaptive noise detection stored in the memory and executable on the processor, wherein the wide dynamic range compression program based on adaptive noise detection implements the steps of the wide dynamic range compression method based on adaptive noise detection as described above when executed by the processor.
[0033] This application also provides a computer-readable storage medium storing a computer program that can be executed to implement the steps of the wide dynamic range compression method based on adaptive noise detection as described above.
[0034] This invention provides a wide dynamic range compression method and apparatus based on adaptive noise detection. The wide dynamic range compression method based on adaptive noise detection includes: converting an input audio time-domain signal into several frame signals in the frequency domain, and dividing each frame signal into several channel signals; if the current frame signal is a speech frame, determining the attribute information of the several channel signals corresponding to the current frame signal; when the attribute information of the channel signal is a noise channel, adjusting the gain by using a preset fixed gain factor as the gain factor of the channel signal to obtain the adjusted spectrum sequence corresponding to the channel signal; based on each adjusted spectrum sequence, performing channel synthesis on each channel signal, and outputting the channel-synthesized time-domain speech signal. This embodiment sets a fixed gain factor for the noise channel, avoiding the defect in wide dynamic range compression algorithms where noise is amplified along with speech, thus improving the speech recognition effect of hearing aids. Attached Figure Description
[0035] Figure 1 This is a flowchart of a preferred embodiment of the wide dynamic range compression method based on adaptive noise detection in this invention;
[0036] Figure 2 This is a logical schematic diagram of a preferred embodiment of the wide dynamic range compression method based on adaptive noise detection in this invention;
[0037] Figure 3 This is a functional principle block diagram of a preferred embodiment of the wide dynamic range compression device based on adaptive noise detection in this invention;
[0038] Figure 4 This is a functional principle block diagram of a preferred embodiment of the terminal in this invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0040] Please see Figure 1 , Figure 1 This is a flowchart of the wide dynamic range compression method based on adaptive noise detection in this invention. Figure 1 As shown in the embodiments of the present invention, the wide dynamic range compression method based on adaptive noise detection includes:
[0041] Step S100: Convert the input audio time-domain signal into several frame signals in the frequency domain, and divide each frame signal into several channel signals.
[0042] Specifically, the continuous audio time-domain signal is divided into multiple shorter frames to facilitate subsequent processing. Each frame is windowed to reduce spectral leakage. The Fast Fourier Transform (FFT) is used to convert the time-domain signal into a frequency-domain signal, resulting in several frame signals in the frequency domain. Channels are divided according to WDRC, with each channel representing a different frequency range.
[0043] In this embodiment of the application, after step S100, the method further includes: if the current frame signal is a noise frame, then a preset fixed gain factor is used as the gain factor of the noise frame for gain adjustment, so as to obtain the adjusted spectrum sequence corresponding to each channel signal of the noise frame.
[0044] Specifically, the attribute information of each frame signal is determined. If the current frame signal is a noise frame, a preset fixed gain factor is applied to the corresponding frequency band to adjust the dynamic range of the frame signal. Existing technology calculates an appropriate gain for the signal in each channel based on the input-output curve of WDRC. However, when the sound pressure level of the signal in a channel exceeds the hearing threshold, the noise signal will also be amplified. This application, on the other hand, performs separate gain adjustment on noise frames, meaning that the noise signal will not be amplified.
[0045] This application combines WDRC channel segmentation and gain to calculate the spectral flatness of the speech frame and each channel signal separately; based on the spectral flatness value, it distinguishes between speech and noise, thereby enabling different processing of speech and noise, such as... Figure 2 As shown.
[0046] like Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in this embodiment further includes:
[0047] Step S200: If the current frame signal is a voice frame, then determine the attribute information of several channel signals corresponding to the current frame signal.
[0048] In this embodiment of the application, before step S200, the method further includes: calculating the first spectral flatness of the current frame signal, and determining the attribute information of the current frame signal based on the first spectral flatness, wherein the attribute information of the current frame signal is a speech frame or a noise frame.
[0049] In one specific embodiment, if the value of the first spectral flatness of the current frame signal is large (close to 1), and the difference between the first spectral flatness of the current frame signal and the first spectral flatness of the adjacent frame signal is small, generally less than or equal to 3% of the first spectral flatness of the current frame signal, then the current frame signal is a noise frame. If the value of the first spectral flatness of the current frame signal is small (close to 0), and the difference between the first spectral flatness of the current frame signal and the first spectral flatness of the adjacent frame signal is large, generally greater than or equal to 20% of the first spectral flatness of the current frame signal, then the current frame signal is a speech frame.
[0050] In another specific embodiment, if the difference between the first spectral flatness of all channel signals is small, the current frame signal is a noise frame; if the difference between the first spectral flatness is large, the current frame signal is a speech frame.
[0051] Specifically, this application embodiment uses the spectral flatness of speech to determine whether the signal is a speech frame or a noise frame, and combines the difference in speech spectral flatness to identify the attribute information of the channel signal, and then performs differential processing according to the different attribute information, thereby realizing adaptive differential processing of speech and noise.
[0052] The embodiments of this application utilize spectral flatness to detect noise frames, thereby amplifying the difference between noise frames and speech frames.
[0053] In one embodiment of this application, step S200 specifically includes:
[0054] Step S210: If the current frame signal is a speech frame, calculate the second spectral flatness corresponding to each channel signal in the current frame signal;
[0055] Step S220: Obtain the noise frame flatness corresponding to the initial noise frame in the currently input audio time domain signal, and calculate the difference between the second spectral flatness and the noise frame flatness;
[0056] Step S230a: If the difference between the second spectral flatness and the noise frame flatness is less than a preset threshold, then the attribute information of the channel signal is a noise channel.
[0057] The formula for calculating the second spectral flatness for each channel signal is as follows:
[0058]
[0059] Where k represents the channel bandwidth, m represents the frequency point, Y(m) represents the energy of the frequency point, and M is the number of channels.
[0060] The flatness of the noise frame corresponding to the initial noise frame is obtained by adaptive tracking of noise during the wide dynamic range compression algorithm. The flatness of the noise frame corresponding to the initial noise frame is used as a reference value for judging subsequent noise channels.
[0061] The embodiments of this application determine the noise channel based on the value of the spectral flatness of each channel signal, thereby realizing different processing of speech and noise.
[0062] In this embodiment of the application, after step S220, the method further includes:
[0063] Step S230b: If the difference between the second spectral flatness and the noise frame flatness is greater than or equal to a preset threshold, then the attribute information of the channel signal is a voice channel.
[0064] The embodiments of this application determine the speech channel based on the value of the spectral flatness of each channel signal, thereby realizing different processing of speech and noise.
[0065] like Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in this embodiment further includes:
[0066] Step S300: When the attribute information of the channel signal is a noise channel, a preset fixed gain factor is used as the gain factor of the channel signal to adjust the gain, thereby obtaining the adjusted spectrum sequence of the channel signal.
[0067] Gain adjustment refers to applying gain to a corresponding frequency band to adjust the dynamic range of a signal. Specifically, a new spectral sequence is obtained based on the gain factor sequence of each channel. If the sequence of channel m is X(m), and the obtained gain factor of that channel is G(m), then the adjusted spectral sequence Y(m) of that channel is: Y(m) = X(m) * G(m).
[0068] In this embodiment, a fixed gain factor G is set for the noise channel. By combining WDRC and spectral flatness, differential amplification of speech and noise signals is achieved. Furthermore, the noise channel is detected using spectral flatness, thus solving the problem of noise instability caused by differential amplification.
[0069] In this embodiment of the application, after step S200, the method further includes:
[0070] When the attribute information of the channel signal is a voice channel, the gain factor of the voice channel is calculated based on the preset user hearing curve function and the input sound level of the voice channel.
[0071] The gain of the voice channel is adjusted according to the gain factor of the voice channel to obtain the adjusted spectrum sequence.
[0072] Specifically, if the input sound level of user channel m is x m If the user's hearing curve function is y = F(x), then the gain factor for this channel is:
[0073] In this embodiment, the speech channel is calculated based on the user's hearing curve function and the input sound level of the speech channel. An appropriate gain is calculated for the signal in each speech channel, which is different from the gain processing of the noise channel. This avoids the defect in wide dynamic range compression algorithms where noise is amplified at the same time as speech.
[0074] In this embodiment, the preset fixed gain factor is less than or equal to 1, and the preset fixed gain factor can be customized. Specifically, the preset fixed gain factor in this embodiment can be customized according to different needs. Since the preset fixed gain factor is set to less than or equal to 1, this means that the noise signal will not be amplified when passing through the system, or will remain unchanged, which helps to reduce unnecessary background noise. Furthermore, this application allows users to customize the preset fixed gain factor, so that the system can be adjusted according to different environmental needs and personal preferences. For example, in a particularly noisy environment, the user can choose a lower gain factor to further suppress noise; while in a relatively quiet environment, a gain factor close to 1 can be selected to maintain the detail of the background sound. In addition, since the gain of the speech frame is calculated based on the user's hearing curve, it can ensure that the amplified speech signal is within the user's hearing comfort range. At the same time, limiting the gain of the noise signal also helps to avoid the potential hearing damage caused by prolonged exposure to high-intensity noise.
[0075] This application can significantly improve the clarity of calls or audio content by processing noise and speech signals separately and optimizing the gain of the speech signal.
[0076] like Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in this embodiment further includes:
[0077] Step S400: Based on the adjusted spectrum sequences, perform channel synthesis on the signals of each channel and output the synthesized time-domain speech signal.
[0078] Specifically, converting the gain-adjusted frequency-domain signal back to a time-domain signal typically involves an inverse fast Fourier transform (IFFT) and windowing. The processed time-domain signal is then sent to a digital-to-analog converter (DAC) to be converted into an analog signal, which is subsequently sent to the hearing aid's speaker for the user to hear.
[0079] In this application, when applying gain to both the noise frame and the noise channel, a preset fixed gain factor is used. The speech frame, however, is calculated using the user's hearing curve. This ensures that even if the sound pressure level of the signal within the channel exceeds the hearing threshold, the noise signal will not be amplified along with the speech frame. In other words, this application utilizes spectral flatness to detect the noise frame and noise channel, applying different gain processing to speech and noise, thus avoiding the drawback of the WDRC algorithm where noise is amplified along with speech.
[0080] In one embodiment, such as Figure 3 As shown, based on the above-described wide dynamic range compression method based on adaptive noise detection, the present invention also provides a wide dynamic range compression device based on adaptive noise detection, comprising:
[0081] The channel division module 100 is used to convert the input audio time-domain signal into several frame signals in the frequency domain, and divide each frame signal into several channel signals;
[0082] The attribute determination module 200 is used to determine the attribute information of several channel signals corresponding to the current frame signal if the current frame signal is a speech frame.
[0083] The gain adjustment module 300 is used to adjust the gain of the channel signal by using a preset fixed gain factor as the gain factor of the channel signal when the attribute information of the channel signal is a noise channel, so as to obtain the adjusted spectrum sequence of the channel signal.
[0084] The channel synthesis module 400 is used to synthesize the channel signals based on the adjusted spectrum sequences and output the synthesized time-domain speech signal.
[0085] Figure 4 A schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal may include:
[0086] The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.
[0087] When the processor 502 executes the program, it implements the wide dynamic range compression method based on adaptive noise detection provided in the above embodiments.
[0088] Furthermore, the terminal also includes:
[0089] Communication interface 503 is used for communication between memory 501 and processor 502.
[0090] The memory 501 is used to store computer programs that can run on the processor 502.
[0091] The memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0092] If the memory 501, processor 502, and communication interface 503 are implemented independently, they can be interconnected via a bus to communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in the diagram, but this does not imply that there is only one bus or one type of bus.
[0093] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.
[0094] Processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0095] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the wide dynamic range compression method based on adaptive noise detection as described above.
[0096] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0097] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0098] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0099] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can read and execute instructions from or in conjunction with such an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). In addition, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically by optically scanning paper or other media, then editing, interpreting or otherwise processing them as necessary, and then storing them in computer memory.
[0100] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0101] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.
[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
[0103] In summary, this invention discloses a wide dynamic range compression method and apparatus based on adaptive noise detection. The wide dynamic range compression method based on adaptive noise detection includes: converting an input audio time-domain signal into several frame signals in the frequency domain, and dividing each frame signal into several channel signals; if the current frame signal is a speech frame, determining the attribute information of the several channel signals corresponding to the current frame signal; when the attribute information of the channel signal is a noise channel, adjusting the gain by using a preset fixed gain factor as the gain factor of the channel signal to obtain the adjusted spectrum sequence corresponding to the channel signal; based on each adjusted spectrum sequence, performing channel synthesis on each channel signal, and outputting the synthesized time-domain speech signal. This embodiment sets a fixed gain factor for the noise channel, avoiding the defect in wide dynamic range compression algorithms where noise is amplified along with speech, thus improving the speech recognition effect of hearing aids.
[0104] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A wide dynamic range compression method based on adaptive noise detection, characterized in that, The method comprises: Converting an input audio time domain signal into a plurality of frame signals in a frequency domain, and dividing each of the frame signals into a plurality of channel signals; Determining attribute information of a current frame signal, the attribute information of the current frame signal being a speech frame or a noise frame; If the current frame signal is a speech frame, determining attribute information of a plurality of channel signals corresponding to the current frame signal; When the attribute information of a channel signal is a noise channel, a preset fixed gain factor is used as a gain factor of the channel signal for gain adjustment to obtain an adjusted spectrum sequence corresponding to the channel signal; Based on each of the adjusted spectrum sequences, channel synthesis is performed on each of the channel signals, and a time domain speech signal after channel synthesis is outputted; If the current frame signal is a speech frame, determining attribute information of a plurality of channel signals corresponding to the current frame signal comprises: If the current frame signal is a speech frame, calculating a second spectral flatness corresponding to each of the channel signals in the current frame signal; Obtaining a noise frame flatness corresponding to an initial noise frame in a current input audio time domain signal, and calculating a difference between the second spectral flatness and the noise frame flatness; If the difference between the second spectral flatness and the noise frame flatness is less than a preset threshold, the attribute information of the channel signal is a noise channel; After obtaining the noise frame flatness corresponding to the initial noise frame in the current input audio time domain signal and calculating the difference between the second spectral flatness and the noise frame flatness, the method further comprises: If the difference between the second spectral flatness and the noise frame flatness is greater than or equal to the preset threshold, the attribute information of the channel signal is a speech channel; After determining the attribute information of the plurality of channel signals corresponding to the current frame signal if the current frame signal is a speech frame, the method further comprises: When the attribute information of a channel signal is a speech channel, a gain factor of the speech channel is calculated based on a preset user hearing curve function and an input sound level of the speech channel; The speech channel is gain-adjusted according to the gain factor of the speech channel to obtain the adjusted spectrum sequence.
2. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that, After converting an input audio time domain signal into a plurality of frame signals in a frequency domain, and dividing each of the frame signals into a plurality of channel signals, the method further comprises: If the current frame signal is a noise frame, a preset fixed gain factor is used as a gain factor of the noise frame for gain adjustment to obtain an adjusted spectrum sequence corresponding to each of the channel signals of the noise frame.
3. The adaptive noise detection based wide dynamic range compression method of claim 1, wherein, The method for determining attribute information of a current frame signal comprises: Calculating a first spectral flatness of the current frame signal, and determining attribute information of the current frame signal based on the first spectral flatness.
4. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that, The preset fixed gain factor is less than or equal to 1, and the preset fixed gain factor can be set by a user.
5. A wide dynamic range compression apparatus based on adaptive noise detection, characterized in that, The device comprises: A channel division module configured to convert an input audio time domain signal into a plurality of frame signals in a frequency domain, and divide each of the frame signals into a plurality of channel signals; An attribute determination module configured to determine attribute information of a plurality of channel signals corresponding to a current frame signal if the current frame signal is a speech frame. The gain adjustment module is configured to, when the attribute information of the channel signal is noise channel, perform gain adjustment on the channel signal by taking a preset fixed gain factor as a gain factor of the channel signal to obtain an adjusted spectrum sequence corresponding to the channel signal. The channel synthesis module is configured to perform channel synthesis on each channel signal based on the adjusted spectrum sequence corresponding to the channel signal, and output a time-domain speech signal after channel synthesis. Attribute information of a current frame signal is determined, and the attribute information of the current frame signal is a speech frame or a noise frame. If the current frame signal is a speech frame, attribute information of a plurality of channel signals corresponding to the current frame signal is determined, including: If the current frame signal is a speech frame, a second spectrum flatness corresponding to each channel signal in the current frame signal is calculated. A noise frame flatness corresponding to an initial noise frame in the input audio time-domain signal is obtained, and a difference between the second spectrum flatness and the noise frame flatness is calculated. If the difference between the second spectrum flatness and the noise frame flatness is less than a preset threshold, the attribute information of the channel signal is noise channel. After the noise frame flatness corresponding to the initial noise frame in the input audio time-domain signal is obtained, and the difference between the second spectrum flatness and the noise frame flatness is calculated, the method further includes: If the difference between the second spectrum flatness and the noise frame flatness is greater than or equal to the preset threshold, the attribute information of the channel signal is speech channel. After the attribute information of the plurality of channel signals corresponding to the current frame signal is determined if the current frame signal is a speech frame, the method further includes: When the attribute information of the channel signal is speech channel, a gain factor of the speech channel is calculated based on a preset user hearing curve function and an input sound level of the speech channel. The gain factor of the speech channel is used to perform gain adjustment on the speech channel to obtain the adjusted spectrum sequence.
6. A terminal, characterized by comprising: The method includes: A memory, a processor, and a wide dynamic range compression program based on adaptive noise detection stored in the memory and executable on the processor, wherein the wide dynamic range compression program based on adaptive noise detection, when executed by the processor, implements the steps of the wide dynamic range compression method based on adaptive noise detection in any one of claims 1-4.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, which can be executed to implement the steps of the wide dynamic range compression method based on adaptive noise detection in any one of claims 1-4.
Citation Information
Patent Citations
Mobile communication terminal and voice enhancement method and module thereof
CN104867498A
Voice endpoint determination method and device, storage medium and electronic device
CN110706693A