Wide dynamic range compression method and device based on adaptive noise detection

By introducing an adaptive noise detection mechanism into the wide dynamic range compression algorithm, the gain of the noise channel is judged and adjusted, and the problem that noise amplification in the prior art will reduce the speech recognition effect, which significantly improves the speech recognition effect of the hearing aid.

CN119943073AActive Publication Date: 2025-05-06ZHUZHOU SHENGYU MEDICAL DEVICE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202411949052.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

When the signal is noisy, when the signal is louder than the listening threshold, the channel signal is also amplified, resulting in a greatly reduced voice recognition effect of the hearing aid.

Method used

A wide dynamic range compression method based on adaptive noise detection is adopted. By converting the input audio time domain signal into a frequency domain frame signal and dividing it into a channel signal, it is determined that the current frame signal is a speech frame or a noise frame. If it is a noise frame, a preset fixed gain factor is used for gain adjustment to avoid the reduction of speech recognition effect caused by noise amplification.

Benefits of technology

It effectively avoids the problem that noise is also amplified in the wide dynamic range compression algorithm, and improves the speech recognition effect of the hearing aid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943073A_ABST
    Figure CN119943073A_ABST
Patent Text Reader

Abstract

The invention provides a wide dynamic range compression method and device based on adaptive noise detection, and the method comprises the steps: converting an input audio time domain signal into a plurality of frame signals in a frequency domain, and dividing each frame signal into a plurality of channel signals; if the current frame signal is a voice frame, determining attribute information of a plurality of channel signals corresponding to the current frame signal; when the attribute information of the channel signal is a noise channel, performing gain adjustment by taking a preset fixed gain factor as a gain factor of the channel signal to obtain an adjusted spectrum sequence corresponding to the channel signal; and performing channel synthesis on each channel signal based on each adjusted frequency spectrum sequence, and outputting a time domain voice signal after channel synthesis. According to the embodiment of the invention, the fixed gain factor is set for the noise channel, so that the defect that the noise is also amplified while the voice is amplified in a wide dynamic range compression algorithm is avoided, and the voice recognition effect of the hearing aid is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic equipment, and in particular to a wide dynamic range compression method and device based on adaptive noise detection. Background Art

[0002] The wide dynamic range compression (WDRC) algorithm is one of the core algorithms of digital hearing aids. Its main function is to compensate for sounds of different frequencies and intensities to varying degrees, thereby converting ambient sounds into the hearing range after damage.

[0003] The wide dynamic range compression algorithm first divides the speech into multiple independent frequency domain channels in the frequency domain. In each channel, according to the audiometric threshold diagram of the hearing-impaired person, independent processing is performed to map the dynamic range of normal sound into the hearing range of the hearing-impaired person, thereby determining the gain of the channel. The gains calculated in different channels are then applied to the input signal in the frequency domain, and finally the sound signal is synthesized and output. One of the parameters of the patient's hearing threshold information is the hearing threshold value. When the signal sound pressure value in the channel is greater than this value, the signal of the channel will be amplified accordingly.

[0004] However, the wide dynamic range compression algorithm processes speech and noise indiscriminately. When the signal is noise, when the signal sound pressure value in the channel is greater than the hearing threshold, the channel signal will also be amplified, thereby greatly reducing the speech recognition effect of the hearing aid.

[0005] Therefore, the prior art has defects and needs to be improved and developed. Summary of the invention

[0006] The technical problem to be solved by the present invention is that, in view of the above-mentioned defects of the prior art, a wide dynamic range compression method and device based on adaptive noise detection is provided, aiming to solve the problem in the prior art that when the signal is noise, when the signal sound pressure value in the channel is greater than the hearing threshold, the channel signal will also be amplified, thereby greatly reducing the speech recognition effect of the hearing aid.

[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0008] A wide dynamic range compression method based on adaptive noise detection, wherein the method comprises:

[0009] Converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals;

[0010] If the current frame signal is a speech frame, determining attribute information of a plurality of channel signals corresponding to the current frame signal;

[0011] When the attribute information of the channel signal is a noise channel, a preset fixed gain factor is used as the gain factor of the channel signal to perform gain adjustment to obtain an adjusted spectrum sequence corresponding to the channel signal;

[0012] Based on each of the adjusted spectrum sequences, channel synthesis is performed on each channel signal, and a time-domain speech signal after channel synthesis is output.

[0013] In one embodiment of the present application, after converting the input audio time domain signal into a plurality of frame signals in the frequency domain and dividing each of the frame signals into a plurality of channel signals, the method further includes:

[0014] If the current frame signal is a noise frame, a preset fixed gain factor is used as a gain factor of the noise frame to perform gain adjustment to obtain an adjusted spectrum sequence corresponding to each channel signal of the noise frame.

[0015] In one embodiment of the present application, if the current frame signal is a speech frame, before determining the attribute information of the plurality of channel signals corresponding to the current frame signal, the method further includes:

[0016] A first spectrum flatness of a current frame signal is calculated, and attribute information of the current frame signal is determined based on the first spectrum flatness, where the attribute information of the current frame signal is a speech frame or a noise frame.

[0017] In one embodiment of the present application, if the current frame signal is a speech frame, determining the attribute information of several channel signals corresponding to the current frame signal includes:

[0018] If the current frame signal is a speech frame, calculating the second spectrum flatness corresponding to each channel signal in the current frame signal;

[0019] Obtaining a noise frame flatness corresponding to an initial noise frame in a currently input audio time domain signal, and calculating a difference between the second spectrum flatness and the noise frame flatness;

[0020] If the difference between the second spectrum flatness and the noise frame flatness is smaller than a preset threshold, the attribute information of the channel signal is a noise channel.

[0021] In one embodiment of the present application, after obtaining the noise frame flatness corresponding to the initial noise frame in the currently input audio time domain signal and calculating the difference between the second spectrum flatness and the noise frame flatness, the method further includes:

[0022] If the difference between the second spectrum flatness and the noise frame flatness is greater than or equal to a preset threshold, the attribute information of the channel signal is a speech channel.

[0023] In one embodiment of the present application, if the current frame signal is a speech frame, after determining the attribute information of several channel signals corresponding to the current frame signal, the method further includes:

[0024] When the attribute information of the channel signal is a speech channel, a gain factor of the speech channel is calculated based on a preset user hearing curve function and an input sound level of the speech channel;

[0025] The gain of the speech channel is adjusted according to the gain factor of the speech channel to obtain the adjusted spectrum sequence.

[0026] In one embodiment of the present application, the preset fixed gain factor is less than or equal to 1, and the preset fixed gain factor can be customized.

[0027] The present application also provides a wide dynamic range compression device based on adaptive noise detection, wherein the device comprises:

[0028] A channel division module, used for converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals;

[0029] An attribute determination module, used for determining attribute information of a plurality of channel signals corresponding to the current frame signal if the current frame signal is a speech frame;

[0030] A gain adjustment module, used for, when the attribute information of the channel signal is a noise channel, using a preset fixed gain factor as the gain factor of the channel signal to perform gain adjustment, so as to obtain an adjusted spectrum sequence corresponding to the channel signal;

[0031] The channel synthesis module is used to perform channel synthesis on each channel signal based on each adjusted spectrum sequence, and output a time domain speech signal after channel synthesis.

[0032] The present application also provides a terminal, which includes: a memory, a processor, and a wide dynamic range compression program based on adaptive noise detection stored in the memory and executable on the processor, wherein the wide dynamic range compression program based on adaptive noise detection, when executed by the processor, implements the steps of the wide dynamic range compression method based on adaptive noise detection as described above.

[0033] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the wide dynamic range compression method based on adaptive noise detection as described above.

[0034] The present invention provides a wide dynamic range compression method and device based on adaptive noise detection, the wide dynamic range compression method based on adaptive noise detection includes: converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals; if the current frame signal is a speech frame, determining the attribute information of the plurality of channel signals corresponding to the current frame signal; when the attribute information of the channel signal is a noise channel, using a preset fixed gain factor as the gain factor of the channel signal to perform gain adjustment, and obtaining an adjusted spectrum sequence corresponding to the channel signal; based on each of the adjusted spectrum sequences, performing channel synthesis on each channel signal, and outputting the time domain speech signal after channel synthesis. The embodiment of the present application sets a fixed gain factor for the noise channel, thereby avoiding the defect that the noise is also amplified while the speech is amplified in the wide dynamic range compression algorithm, and improving the speech recognition effect of the hearing aid. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flow chart of a preferred embodiment of a wide dynamic range compression method based on adaptive noise detection in the present invention;

[0036] Figure 2 It is a logic schematic diagram of a preferred embodiment of a wide dynamic range compression method based on adaptive noise detection in the present invention;

[0037] Figure 3 It is a functional principle block diagram of a preferred embodiment of a wide dynamic range compression device based on adaptive noise detection in the present invention;

[0038] Figure 4 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0040] See also Figure 1 , Figure 1 Flowchart of the wide dynamic range compression method based on adaptive noise detection in the present invention. Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in the embodiment of the present invention includes:

[0041] Step S100: converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals.

[0042] Specifically, the continuous audio time domain signal is divided into multiple shorter frames for subsequent processing, each frame is windowed to reduce spectrum leakage, and the time domain signal is converted into a frequency domain signal using Fast Fourier Transform (FFT) to obtain several frame signals in the frequency domain. According to the WDRC channel division, the frequency domain signal is divided into multiple channels, each channel represents a different frequency range.

[0043] In the embodiment of the present application, after step S100, the following step is further performed: if the current frame signal is a noise frame, a preset fixed gain factor is used as a gain factor of the noise frame to perform gain adjustment to obtain an adjusted spectrum sequence corresponding to each channel signal of the noise frame.

[0044] Specifically, the attribute information of each frame signal is judged. If the current frame signal is a noise frame, a preset fixed gain factor is applied to the corresponding frequency band to adjust the dynamic range of the frame signal. The existing technology calculates the appropriate gain for each signal in the channel according to the input and output curve of WDRC. However, when the sound pressure value of the signal in the channel is greater than the listening threshold, the noise signal will also be amplified. The present application performs a separate gain adjustment on the noise frame, that is, the noise signal will not be amplified.

[0045] This application combines the channel division and gain of WDRC to calculate the spectral flatness of the speech frame and each channel signal respectively; the speech and noise are judged according to the value of the spectral flatness, so as to achieve different processing of speech and noise, such as Figure 2 shown.

[0046] like Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in this embodiment also includes:

[0047] Step S200: If the current frame signal is a speech frame, determine the attribute information of several channel signals corresponding to the current frame signal.

[0048] In the embodiment of the present application, before step S200, the method further includes: calculating a first spectrum flatness of a current frame signal, and determining attribute information of the current frame signal based on the first spectrum flatness, wherein the attribute information of the current frame signal is a speech frame or a noise frame.

[0049] In a specific embodiment, if the value of the first spectrum flatness of the current frame signal is relatively large (close to 1), and the difference between the first spectrum flatness of the current frame signal and the first spectrum flatness of the adjacent frame signal is small, and the difference is generally less than or equal to 3% of the first spectrum flatness of the current frame signal, then the current frame signal is a noise frame. If the value of the first spectrum flatness of the current frame signal is relatively small (close to 0), and the difference between the first spectrum flatness of the current frame signal and the first spectrum flatness of the adjacent frame signal is large, and the difference is generally greater than or equal to 20% of the first spectrum flatness of the current frame signal, then the current frame signal is a speech frame.

[0050] In another specific embodiment, by comparing the first spectrum flatnesses of all channel signals, if the difference between the first spectrum flatnesses is small, the current frame signal is a noise frame; if the difference between the first spectrum flatnesses is large, the current frame signal is a speech frame.

[0051] Specifically, the embodiment of the present application utilizes the spectral flatness of speech to determine whether the signal is a speech frame or a noise frame, and combines the difference in speech spectrum to identify the attribute information of the channel signal, and then performs differential processing according to the difference in attribute information, thereby realizing adaptive differential processing of speech and noise.

[0052] The embodiment of the present application utilizes spectrum flatness to detect noise frames, thereby achieving difference amplification between noise frames and speech frames.

[0053] In one embodiment of the present application, step S200 specifically includes:

[0054] Step S210: If the current frame signal is a speech frame, calculate the second spectrum flatness corresponding to each channel signal in the current frame signal;

[0055] Step S220, obtaining a noise frame flatness corresponding to an initial noise frame in a currently input audio time domain signal, and calculating a difference between the second spectrum flatness and the noise frame flatness;

[0056] Step S230a: If the difference between the second spectrum flatness and the noise frame flatness is less than a preset threshold, the attribute information of the channel signal is a noise channel.

[0057] Calculate the second spectrum flatness corresponding to each channel signal, the formula is as follows:

[0058]

[0059] Where k represents the channel band, m represents the frequency point, Y(m) represents the energy of the frequency point, and M is the number of channels.

[0060] The noise frame flatness corresponding to the initial noise frame is obtained by adaptive tracking processing based on the noise during the wide dynamic range compression algorithm. The noise frame flatness corresponding to the initial noise frame is used as a reference value for judging subsequent noise channels.

[0061] The embodiment of the present application determines the noise channel based on the value of the flatness of the signal spectrum of each channel, thereby achieving different processing of speech and noise.

[0062] In the embodiment of the present application, after step S220, the following steps are further included:

[0063] Step S230b: If the difference between the second spectrum flatness and the noise frame flatness is greater than or equal to a preset threshold, the attribute information of the channel signal is a speech channel.

[0064] The embodiment of the present application determines the speech channel based on the value of the flatness of the spectrum of each channel signal, thereby achieving different processing of speech and noise.

[0065] like Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in this embodiment also includes:

[0066] Step S300: When the attribute information of the channel signal is a noise channel, a preset fixed gain factor is used as the gain factor of the channel signal to perform gain adjustment to obtain an adjusted spectrum sequence corresponding to the channel signal.

[0067] Gain adjustment refers to applying gain to the corresponding frequency band to adjust the dynamic range of the signal. Specifically, a new spectrum sequence is obtained based on the obtained gain factor sequence of each channel. If the sequence of channel m is X(m), and the obtained gain factor of the channel is G(m), then the adjusted spectrum sequence Y(m) of the channel is: Y(m) = X(m)*G(m).

[0068] The embodiment of the present application sets a fixed gain factor G for the noise channel, combines WDRC with spectrum flatness to achieve differential amplification of speech and noise signals, and uses spectrum flatness to detect the noise channel, thereby solving the problem of noise instability caused by differential amplification.

[0069] In the embodiment of the present application, after step S200, the following steps are further included:

[0070] When the attribute information of the channel signal is a speech channel, a gain factor of the speech channel is calculated based on a preset user hearing curve function and an input sound level of the speech channel;

[0071] The gain of the speech channel is adjusted according to the gain factor of the speech channel to obtain the adjusted spectrum sequence.

[0072] Specifically, if the input sound level of user channel m is x m , the user's hearing curve function is y = F (x), then the gain factor of the channel is:

[0073] The voice channel in the embodiment of the present application is calculated based on the user's hearing curve function and the input sound level of the voice channel, and an appropriate gain is calculated for the signal in each voice channel, which is different from the gain processing of the noise channel, thereby avoiding the defect that the noise is also amplified while the voice is amplified in the wide dynamic range compression algorithm.

[0074] In an embodiment of the present application, the preset fixed gain factor is less than or equal to 1, and the preset fixed gain factor can be customized. Specifically, the preset fixed gain factor of the embodiment of the present application can be customized according to different needs. Since the preset fixed gain factor is set to be less than or equal to 1, this means that the noise signal will not be amplified when passing through the system, or remain as it is, which helps to reduce unnecessary background noise. In addition, the present application allows users to customize the preset fixed gain factor so that the system can be adjusted according to different environmental requirements and personal preferences. For example, in an environment with particularly loud noise, the user can choose a lower gain factor to further suppress the noise; while in a relatively quiet environment, a gain factor close to 1 can be selected to maintain the details of the background sound. In addition, since the gain of the speech frame is calculated based on the user's hearing curve, it can ensure that the amplified speech signal is within the user's hearing comfort range. At the same time, limiting the gain of the noise signal also helps to avoid the damage to hearing that may be caused by long-term exposure to high-intensity noise.

[0075] This application can significantly improve the clarity of calls or audio content by processing noise and voice signals separately and optimizing the gain of voice signals.

[0076] like Figure 1 As shown, the wide dynamic range compression method based on adaptive noise detection described in this embodiment also includes:

[0077] Step S400: Based on the adjusted spectrum sequences, perform channel synthesis on each channel signal, and output a time-domain speech signal after channel synthesis.

[0078] Specifically, the gain-adjusted frequency domain signal is converted back to a time domain signal, which usually involves an inverse fast Fourier transform (IFFT) and windowing. The processed time domain signal is then sent to a digital-to-analog converter (DAC) to be converted into an analog signal, which is then sent to the hearing aid's speaker for the user to hear.

[0079] In the embodiment of the present application, when performing gain on noise frames and noise channels, a preset fixed gain factor is used, while the speech frame is calculated through the user's hearing curve, so that even if the signal sound pressure value in the channel is greater than the hearing threshold, the noise signal will not be amplified together with the speech frame. That is, the present application uses spectrum flatness to detect noise frames and noise channels, and performs different gain processing on speech and noise, avoiding the defect that the noise is also amplified while the speech is amplified in the WDRC algorithm.

[0080] In one embodiment, if Figure 3 As shown, based on the above-mentioned wide dynamic range compression method based on adaptive noise detection, the present invention also provides a wide dynamic range compression device based on adaptive noise detection, including:

[0081] The channel division module 100 is used to convert the input audio time domain signal into a plurality of frame signals in the frequency domain, and divide each of the frame signals into a plurality of channel signals;

[0082] The attribute determination module 200 is used to determine the attribute information of a plurality of channel signals corresponding to the current frame signal if the current frame signal is a speech frame;

[0083] The gain adjustment module 300 is used to adjust the gain by using a preset fixed gain factor as the gain factor of the channel signal when the attribute information of the channel signal is a noise channel, so as to obtain an adjusted spectrum sequence corresponding to the channel signal;

[0084] The channel synthesis module 400 is used to perform channel synthesis on each channel signal based on each adjusted spectrum sequence, and output a time-domain speech signal after channel synthesis.

[0085] Figure 4 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:

[0086] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .

[0087] When the processor 502 executes the program, the wide dynamic range compression method based on adaptive noise detection provided in the above embodiment is implemented.

[0088] Furthermore, the terminal further includes:

[0089] The communication interface 503 is used for communication between the memory 501 and the processor 502 .

[0090] The memory 501 is used to store computer programs that can be executed on the processor 502 .

[0091] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0092] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0093] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.

[0094] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0095] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the wide dynamic range compression method based on adaptive noise detection as described above is implemented.

[0096] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0097] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0098] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, in which the order shown or discussed may not be followed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0099] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can read instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary and then storing it in a computer memory.

[0100] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0101] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0102] In addition, each functional unit in each embodiment of the present application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above-mentioned embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above-mentioned embodiments within the scope of the present application.

[0103] In summary, the present invention discloses a wide dynamic range compression method and device based on adaptive noise detection, and the wide dynamic range compression method based on adaptive noise detection includes: converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals; if the current frame signal is a speech frame, determining the attribute information of the plurality of channel signals corresponding to the current frame signal; when the attribute information of the channel signal is a noise channel, using a preset fixed gain factor as the gain factor of the channel signal for gain adjustment to obtain an adjusted spectrum sequence corresponding to the channel signal; based on each of the adjusted spectrum sequences, performing channel synthesis on each channel signal, and outputting the time domain speech signal after channel synthesis. The embodiment of the present application sets a fixed gain factor for the noise channel, thereby avoiding the defect that the noise is also amplified while the speech is amplified in the wide dynamic range compression algorithm, and improving the speech recognition effect of the hearing aid.

[0104] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A wide dynamic range compression method based on adaptive noise detection, characterized in that: The method comprises: Converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals; If the current frame signal is a speech frame, determining attribute information of a plurality of channel signals corresponding to the current frame signal; When the attribute information of the channel signal is a noise channel, a preset fixed gain factor is used as the gain factor of the channel signal to perform gain adjustment to obtain an adjusted spectrum sequence corresponding to the channel signal; Based on each of the adjusted spectrum sequences, channel synthesis is performed on each channel signal, and a time-domain speech signal after channel synthesis is output.

2. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that: After converting the input audio time domain signal into a plurality of frame signals in the frequency domain and dividing each of the frame signals into a plurality of channel signals, the method further includes: If the current frame signal is a noise frame, a preset fixed gain factor is used as a gain factor of the noise frame to perform gain adjustment to obtain an adjusted spectrum sequence corresponding to each channel signal of the noise frame.

3. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that: If the current frame signal is a speech frame, before determining the attribute information of the plurality of channel signals corresponding to the current frame signal, the method further includes: A first spectrum flatness of a current frame signal is calculated, and attribute information of the current frame signal is determined based on the first spectrum flatness, where the attribute information of the current frame signal is a speech frame or a noise frame.

4. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that: If the current frame signal is a speech frame, determining attribute information of a plurality of channel signals corresponding to the current frame signal includes: If the current frame signal is a speech frame, calculating the second spectrum flatness corresponding to each channel signal in the current frame signal; Obtaining a noise frame flatness corresponding to an initial noise frame in a currently input audio time domain signal, and calculating a difference between the second spectrum flatness and the noise frame flatness; If the difference between the second spectrum flatness and the noise frame flatness is smaller than a preset threshold, the attribute information of the channel signal is a noise channel.

5. The wide dynamic range compression method based on adaptive noise detection according to claim 4, characterized in that: After obtaining the noise frame flatness corresponding to the initial noise frame in the currently input audio time domain signal and calculating the difference between the second spectrum flatness and the noise frame flatness, the method further includes: If the difference between the second spectrum flatness and the noise frame flatness is greater than or equal to a preset threshold, the attribute information of the channel signal is a speech channel.

6. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that: If the current frame signal is a speech frame, after determining the attribute information of a plurality of channel signals corresponding to the current frame signal, the method further includes: When the attribute information of the channel signal is a speech channel, a gain factor of the speech channel is calculated based on a preset user hearing curve function and an input sound level of the speech channel; The gain of the speech channel is adjusted according to the gain factor of the speech channel to obtain the adjusted spectrum sequence.

7. The wide dynamic range compression method based on adaptive noise detection according to claim 1, characterized in that: The preset fixed gain factor is less than or equal to 1, and the preset fixed gain factor can be customized.

8. A wide dynamic range compression device based on adaptive noise detection, characterized in that: The device comprises: A channel division module, used for converting the input audio time domain signal into a plurality of frame signals in the frequency domain, and dividing each of the frame signals into a plurality of channel signals; An attribute determination module, used for determining attribute information of a plurality of channel signals corresponding to the current frame signal if the current frame signal is a speech frame; A gain adjustment module, used for, when the attribute information of the channel signal is a noise channel, using a preset fixed gain factor as the gain factor of the channel signal to perform gain adjustment, so as to obtain an adjusted spectrum sequence corresponding to the channel signal; The channel synthesis module is used to perform channel synthesis on each channel signal based on each adjusted spectrum sequence, and output a time domain speech signal after channel synthesis.

9. A terminal, characterized in that: include: A memory, a processor, and a wide dynamic range compression program based on adaptive noise detection stored in the memory and executable on the processor, wherein the wide dynamic range compression program based on adaptive noise detection, when executed by the processor, implements the steps of the wide dynamic range compression method based on adaptive noise detection as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the wide dynamic range compression method based on adaptive noise detection according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Mobile communication terminal and voice enhancement method and module thereof

    CN104867498A

  • Voice endpoint determination method and device, storage medium and electronic device

    CN110706693A

  • Method and device for detecting and eliminating noise, coding method, medium and equipment

    CN115762547A

  • Voice wide dynamic range compression method and device, equipment and storage medium

    CN117789735A

  • Voice section detecting device

    JP2004272052A

Cited By

  • Audio signal processing equipment and method

    CN120730221A