Noise processing method and device for audio signal, frequency band division method and device, and electronic device

By analyzing the frequency domain characteristics and frequency band division of audio signals, the noise of prompts generated by real-time communication devices is identified and processed, solving the problem of unsatisfactory noise processing in existing technologies and achieving more efficient noise identification and audio playback effects.

CN116312593BActive Publication Date: 2026-03-31ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing noise reduction solutions cannot effectively handle the noise from real-time communication devices, resulting in unsatisfactory noise reduction effects and affecting users' audio listening experience.

Method used

By analyzing the frequency domain characteristics of audio signals, multiple preset frequency bands are divided, target frequency bands are identified and targeted noise reduction is performed. Combined with a speech classification model, it is determined whether noise reduction should be performed, providing a frequency band division method and device to optimize the noise processing effect.

Benefits of technology

It improves the accuracy of noise processing and audio playback quality, avoids masking low-energy speech signals and preserving high-energy noise, and optimizes the user's auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312593B_ABST
    Figure CN116312593B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a noise processing method and device of an audio signal, a frequency band division method and device, an audio noise reduction processing method and device of an online conference, a computer readable storage medium and an electronic device. The method comprises: obtaining an audio signal generated in a real-time communication process, and extracting a frequency domain feature of the audio signal; obtaining energy values corresponding to the distribution of the audio signal in a plurality of preset frequency bands according to the frequency domain feature; determining a target frequency band from the plurality of preset frequency bands, the energy value corresponding to the target frequency band being not lower than a preset energy value; and if the target frequency band is included in the plurality of preset frequency bands, and it is determined that the audio signal includes noise according to the number of the target frequency bands, performing noise reduction processing on the audio signal corresponding to the target frequency band. This scheme helps to improve the noise processing effect and thus improve the audio playback effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to a noise processing method and apparatus for audio signals, a frequency band division method and apparatus, an audio noise reduction processing method and apparatus for online meetings, a computer-readable storage medium, and an electronic device. Background Technology

[0002] With the continuous development of real-time communication technology, there are more and more application scenarios involving online transmission of audio data, such as online meetings, online courses, online live broadcasts, and video calls, all of which involve the real-time transmission of audio data.

[0003] To improve the quality of audio data transmission, it is usually necessary to perform noise reduction processing on the audio signals captured by the microphone to filter out some environmental noise, such as the sound of air conditioners running and fans running.

[0004] Current noise reduction solutions primarily target environmental noise and cannot effectively address noise generated by devices engaged in real-time communication. This results in unsatisfactory noise reduction performance and negatively impacts the user's audio listening experience. Summary of the Invention

[0005] This application provides a method and apparatus for noise processing of audio signals, a method and apparatus for frequency band division, a method and apparatus for audio noise reduction processing in online meetings, a computer-readable storage medium, and an electronic device, which can effectively identify the prompt noise generated by the device itself, improve the noise processing effect, and thus improve the audio playback effect.

[0006] This application provides the following solution:

[0007] A noise processing method for an audio signal, comprising:

[0008] Obtain the audio signal generated during real-time communication and extract the frequency domain features of the audio signal;

[0009] Based on the frequency domain characteristics, the energy values ​​of the audio signal distributed in multiple preset frequency bands are obtained;

[0010] A target frequency band is determined from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than the preset energy value;

[0011] If the target frequency band is included among the multiple preset frequency bands, and noise is determined to be present in the audio signal based on the number of target frequency bands, noise reduction processing is performed on the audio signal corresponding to the target frequency band.

[0012] The method further includes:

[0013] The sum of the energy values ​​corresponding to the energy values ​​of each of the multiple preset frequency bands is obtained. The energy values ​​corresponding to the multiple preset frequency bands and the sum of the energy values ​​are used as input to the speech classification model to obtain the output result of the speech classification model.

[0014] If the output indicates that the audio signal includes a speech signal, then the step of performing noise reduction processing on the audio signal corresponding to the target frequency band is executed.

[0015] The method further includes:

[0016] If the output indicates that the audio signal does not include a speech signal, then no noise processing is performed.

[0017] The method further includes:

[0018] Identify the target noise that can be noise-reduced, and display the identification information of the target noise through a preset page. The preset page provides operation options for selecting the noise to be processed from the target noise.

[0019] After obtaining the noise to be processed through the operation options, if it is determined that the audio signal contains noise and the noise contained in the audio signal belongs to the noise to be processed, then noise reduction processing is performed on the audio signal corresponding to the target frequency band.

[0020] Wherein, the preset energy value is E 0i And / or E0, the preset energy value is obtained in the following manner:

[0021] The preset energy value E corresponding to the i-th preset frequency band 0i For: except E i The sum of the maximum energy value among all other energy values, and the first preset value A; E i The value is the energy value corresponding to the i-th preset frequency band; if the audio signal is the original audio signal collected by the microphone, 100≤A≤150; if the audio signal is an audio signal that has been noise-processed, 200≤A≤300.

[0022] The preset energy value E0 is the product of the sum of the energy values ​​corresponding to all preset frequency bands and the second preset value B, where 40% ≤ B ≤ 60%.

[0023] The step of determining whether the audio signal contains noise based on the number of target frequency bands includes:

[0024] If one of the multiple preset frequency bands is a target frequency band, then the audio signal is determined to contain noise.

[0025] The step of determining whether the audio signal contains noise based on the number of target frequency bands includes:

[0026] If all of the preset frequency bands are target frequency bands, then the audio signal is determined to contain noise.

[0027] A frequency band allocation method, comprising:

[0028] The distribution frequency bands corresponding to multiple target noises are obtained, wherein the target noises are application prompt sounds;

[0029] A preset frequency band to be analyzed is determined, and the preset frequency band to be analyzed is divided into multiple preset frequency bands according to the distribution frequency bands corresponding to the multiple target noises, wherein a single distribution frequency band is located within a preset frequency band.

[0030] The number N of the preset frequency bands is: 6≤N≤9.

[0031] The range of the preset frequency band to be analyzed is 80Hz to 4KHz.

[0032] Wherein, the single distributed frequency band is located within a preset frequency band, including:

[0033] The upper limit frequency value of the distributed frequency band is less than the upper limit frequency value of the corresponding preset frequency band.

[0034] An audio noise reduction method for online meetings includes:

[0035] When a first user is having an online meeting with at least one second user, the meeting client associated with the first user obtains the audio signal generated in the space where the first user is located.

[0036] Extract the frequency domain features of the audio signal, and obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands based on the frequency domain features;

[0037] A target frequency band is determined from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than the preset energy value;

[0038] If the target frequency band is included among the multiple preset frequency bands, and noise is determined to be present in the audio signal based on the number of target frequency bands, noise reduction processing is performed on the audio signal corresponding to the target frequency band.

[0039] The noise-reduced audio signal is played back, and the noise-reduced audio signal is sent to the conference client associated with the at least one second user for audio playback.

[0040] An audio signal noise processing device, comprising:

[0041] A frequency domain feature extraction unit is used to obtain audio signals generated during real-time communication and extract the frequency domain features of the audio signals.

[0042] An energy value acquisition unit is used to obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands based on the frequency domain characteristics.

[0043] A target frequency band determination unit is used to determine a target frequency band from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than a preset energy value;

[0044] A noise processing unit is configured to include the target frequency band among the plurality of preset frequency bands, and to perform noise reduction processing on the audio signal corresponding to the target frequency band when it is determined that the audio signal contains noise based on the number of target frequency bands.

[0045] A frequency band allocation device, comprising:

[0046] A frequency band acquisition unit is used to acquire the frequency bands corresponding to multiple target noises, wherein the target noises are application prompt sounds.

[0047] The preset frequency band division unit is used to determine the preset frequency band to be analyzed, and divide the preset frequency band to be analyzed into multiple preset frequency bands according to the distribution frequency bands corresponding to the multiple target noises, and each distribution frequency band is located within a preset frequency band.

[0048] An audio noise reduction processing device for online meetings, applied to a meeting client associated with a first user, the device comprising:

[0049] An audio signal acquisition unit is used to acquire audio signals generated in the space where the first user is located when the first user is having an online meeting with at least one second user.

[0050] A frequency domain feature extraction unit is used to extract the frequency domain features of the audio signal;

[0051] An energy value acquisition unit is used to obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands based on the frequency domain characteristics.

[0052] A target frequency band determination unit is used to determine a target frequency band from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than a preset energy value;

[0053] A noise processing unit is configured to include the target frequency band among the plurality of preset frequency bands, and to perform noise reduction processing on the audio signal corresponding to the target frequency band when it is determined that the audio signal contains noise based on the number of target frequency bands;

[0054] An audio playback unit is used to play the noise-reduced audio signal and to send the noise-reduced audio signal to the conference client associated with the at least one second user for audio playback.

[0055] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of any of the preceding methods.

[0056] An electronic device, comprising:

[0057] One or more processors; and

[0058] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of any of the preceding methods.

[0059] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0060] This application provides a scheme to identify whether an audio signal contains noise based on the energy value distribution of the audio signal in multiple preset frequency bands by analyzing the characteristics of the frequency band distribution of the voice signal and the prompt noise. When it is determined that the audio signal contains noise, noise reduction processing is performed on the audio signal in the target frequency band where the noise is located.

[0061] Specifically, multiple preset frequency bands can be divided first. After obtaining the audio signal, the energy value of the audio signal distributed in each preset frequency band can be calculated. The preset frequency bands with energy values ​​not lower than the preset energy value can be determined as target frequency bands. Then, based on the number of target frequency bands, it can be identified whether the audio signal includes noise.

[0062] This approach, unlike existing technologies that directly perform overall noise reduction on the audio signal, avoids the situation where low-energy speech signals are eliminated while high-energy noise is retained, thus improving the noise processing effect. Furthermore, the solution in this application can identify the target frequency band where the noise is located and perform targeted noise reduction on the audio signal within that target frequency band. This ensures that the noise-reduced audio signal no longer includes the originally acquired noise, further improving the noise processing effect and optimizing the user's auditory experience.

[0063] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0064] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0065] Figure 1 This is a schematic diagram of the real-time communication system provided in the embodiments of this application;

[0066] Figure 2 This is a flowchart of the noise processing method provided in the embodiments of this application;

[0067] Figure 3 This is a schematic diagram of the spectrogram provided in the embodiments of this application;

[0068] Figure 4 This is a schematic diagram of another spectrogram provided in an embodiment of this application;

[0069] Figure 5 This is a schematic diagram of the noise processing device provided in the embodiments of this application;

[0070] Figure 6 This is a schematic diagram of the frequency band division device provided in the embodiments of this application;

[0071] Figure 7 This is a schematic diagram of the audio noise reduction processing device for online meetings provided in the embodiments of this application;

[0072] Figure 8 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0073] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0074] When users engage in real-time communication through terminal devices (such as computers, mobile phones, etc.), the device's microphone can capture the voice signals generated by the user's speech and transmit them to the other device in real time. Simultaneously, the microphone also captures and transmits background noise, such as:

[0075] 1. Environmental noise generated in the environment where the equipment is located, such as the sound of air conditioner running, fan running, fingers tapping on the table, coughing, etc.;

[0076] 2. Noise generated by the terminal device itself, such as system operation prompts of real-time communication applications (e.g., the prompt when a speaker turns on their microphone), message prompts from applications other than real-time communication applications, etc.

[0077] To optimize audio playback, noise reduction is currently primarily achieved through overall analysis and overall sound reduction. Specifically, the audio signal is first subjected to speech detection to determine if human speech is present. If speech is detected, the overall audio signal is reduced in decibels (dB).

[0078] For example, if a person's voice is 60dB and the noise is 30dB, reducing the dB can compress the overall sound by 15dB, thus reducing noise and enhancing speech. This utilizes the masking effect of the human ear (the phenomenon where the perception of a weaker sound is affected by a stronger sound), so even if the noise-reduced audio signal contains both speech and noise signals, the noise is eliminated from an auditory perspective.

[0079] This approach has the following problems: the transmitted audio signal still contains noise signals; noise cancellation is only achieved from the perspective of auditory perception. Furthermore, overall sound reduction of the acquired audio signal may mask some low-decibel speech signals while retaining some high-decibel noise, resulting in unsatisfactory audio playback after noise reduction. In practical use, users need to manually set the noise reduction level before using real-time communication applications. Generally, a higher noise reduction level results in greater sound reduction, and the selection of the noise reduction level directly affects the noise reduction effect, placing high demands on the user.

[0080] Correspondingly, embodiments of this application can provide a noise processing client for audio signals. This client can be deployed as a standalone application on a user-associated terminal device, performing noise reduction processing when the user conducts audio or audio-video communication through a real-time communication application installed on the device. Alternatively, it can be as follows: Figure 1 As shown, the system is integrated into a real-time communication application. It acquires audio signals generated during real-time communication by the microphone, performs noise reduction processing, and then sends the processed audio signals to the peer device of the real-time communication through a cloud server for audio or video playback.

[0081] The following is a detailed description of the specific implementation process of the audio signal noise processing method provided in the embodiments of this application. (See also...) Figure 2 The flowchart shown may include:

[0082] S101: Obtain the audio signal generated during real-time communication and extract the frequency domain features of the audio signal.

[0083] As an example, the audio signal generated during real-time communication is emitted by the speaker, then collected by the microphone and transmitted to the noise processing client provided in this application embodiment. Time-domain sampling can be performed first. For example, the effective part of speech during normal human speech usually does not exceed 4kHz. In order to save resources consumed by noise processing, the time-domain sampling rate can be set to 8kHz. That is, the audio signal collected by the microphone is downsampled to 8kHz and then converted into a spectrogram in the frequency domain to extract the frequency domain features of the audio signal.

[0084] A spectrogram is a spectral analysis view of a speech signal. The horizontal axis of a spectrogram represents time, the vertical axis represents frequency, and the values ​​at each coordinate point are the energy values ​​of the speech signal. Because it uses a two-dimensional plane to represent three-dimensional information, the energy value is represented by color; the darker the color, the stronger the speech energy at that point. For an example, see [link to example]. Figure 3 The spectrum obtained by converting the audio signal shown is a spectrogram.

[0085] S102: Based on the frequency domain characteristics, obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands.

[0086] As an example, the energy value of the audio signal distributed across each preset frequency band can be obtained by calculating the root mean square (RMS); alternatively, the corresponding energy value can be estimated based on the dB value of the audio signal within each preset frequency band. This application does not limit the implementation process for calculating the energy value.

[0087] The inventors analyzed audio signals and found that the speech signal of a user speaking normally is generally located in the mid-frequency band (e.g., 250Hz to 500Hz), and the ending sound of speech is located in the low-frequency band (e.g., 0Hz to 250Hz). In other words, normal speech is a gradual fading process, without abrupt stops. From a frequency domain spectrogram perspective, the speech signal usually exists across frequency bands. In contrast, the prompts generated by applications are usually short sounds, belonging to a single-frequency band, and do not cross frequency bands. In this application embodiment, a single-frequency band sound can be understood as a sound that exists only within a preset frequency band.

[0088] Therefore, the embodiments of this application can identify noise based on the characteristics of the frequency band distribution of the voice signal and the prompt noise, and then perform noise reduction processing on the frequency band where the noise is located when it is determined that noise exists in the audio signal. Compared with the existing technology's overall noise reduction scheme for voice signals and noise, this helps to improve the noise processing effect, thereby improving the audio playback effect and optimizing the user experience.

[0089] As an example, frequency band division in this application embodiment can be performed as follows: obtaining the distribution frequency bands corresponding to multiple target noises, wherein the target noises are application prompts; determining a preset frequency band to be analyzed, and dividing the preset frequency band to be analyzed into multiple preset frequency bands according to the distribution frequency bands corresponding to the multiple target noises, wherein each distribution frequency band is located within a preset frequency band. As an example, the superposition of multiple preset frequency bands can cover the preset frequency band to be analyzed.

[0090] As an example, based on usage requirements, common real-time communication application message tones and system operation prompts can be identified as target noise, their frequency distribution can be analyzed, and the preset frequency bands to be analyzed can be divided into multiple preset frequency bands. These preset frequency bands can be used to represent the frequency coverage range of the effective portion of a user's normal speech.

[0091] As the above analysis shows, if the number of preset frequency bands is too small, i.e., the frequency range of a single preset frequency band is too large, it may affect the accuracy of noise recognition. Referring to the example above, if the preset frequency band is 0Hz to 500Hz, it may identify the user's normal speech and the ending of the speech as a single-frequency sound, and thus misjudge it as a prompt noise, affecting the recognition accuracy.

[0092] Conversely, if the number of preset frequency bands is too large, meaning the frequency span of a single preset frequency band is too small, it may also affect the accuracy of noise recognition. For example, if the frequency band corresponding to the notification tone of application A is 100Hz to 120Hz, and the preset frequency bands are 0Hz to 100Hz and 100Hz to 200Hz, the notification tone may be identified as a non-single-band sound distributed across two preset frequency bands, thus misjudging it as a speech signal and affecting recognition accuracy. In addition, too many preset frequency bands will also increase the resource consumption of the device during noise processing because it is necessary to calculate the energy value of the audio signal distributed within each preset frequency band.

[0093] In summary, this application embodiment, by comprehensively considering the accuracy of noise identification and the resource consumption of noise processing, determines that the number N of preset frequency bands can be: 6≤N≤9.

[0094] In practical applications, considering that human speech does not fall within the 0Hz–80Hz frequency band, audio signals in this band can be directly removed and not played. Therefore, the preset frequency band to be analyzed for noise processing can be defined as 80Hz–4kHz. Taking N=7 preset frequency bands as an example, the resulting preset frequency bands can include: 80Hz–250Hz, 250Hz–500Hz, 500Hz–750Hz, 750Hz–1kHz, 1kHz–2kHz, 2kHz–3kHz, and 3kHz–4kHz.

[0095] This application does not limit the range of the preset frequency band to be analyzed or the specific division method of the preset frequency band. It can be determined according to the usage requirements, such as the distribution frequency band of the prompt tone that needs to be noise processed, the frequency band covered by the user's speech, and other information.

[0096] It should be noted that when dividing the preset frequency band in this application embodiment, the correspondence between the distribution frequency band corresponding to the target noise and the preset frequency band can be one-to-one, that is, one distribution frequency band of a target noise corresponds to one preset frequency band. Alternatively, the correspondence can be many-to-one, that is, when the distribution frequency bands of multiple target noises are relatively similar, there may be a situation where multiple distribution frequency bands of target noise correspond to one preset frequency band. For example, the distribution frequency band corresponding to the prompt tone of application A is 100Hz to 120Hz, and the distribution frequency band corresponding to the prompt tone of application B is 110Hz to 200Hz. Both can be corresponding to the preset frequency band 80Hz to 250Hz, that is, these two types of noise can be identified in the preset frequency band 80Hz to 250Hz.

[0097] As can be seen from the above correspondence, in the embodiments of this application, the distribution frequency band corresponding to a single target noise is located within a preset frequency band and will not cross the preset frequency band.

[0098] In addition, to further improve the accuracy of noise identification, when dividing the preset frequency band according to the distribution frequency band of the target noise, it is preferable not to use the upper limit frequency value of the distribution frequency band of the target noise as the critical point for dividing the preset frequency band, that is, the upper limit frequency value of the distribution frequency band is less than the upper limit frequency value of the corresponding preset frequency band.

[0099] For example, if the upper limit of the frequency band 110Hz to 200Hz for the notification sound of application B is used as the critical point, and the preset frequency bands are divided into 80Hz to 200Hz, 200Hz to 500Hz, ..., the following problems may occur during actual use:

[0100] If the device's speaker malfunctions (e.g., speaker aging), the actual frequency band collected by the microphone may shift after the application B prompt tone is emitted. This means the actual audio signal, reflected in the spectrogram, may exhibit divergent and unclear upper and lower boundaries. For example, the actual collected audio signal might be located in the 100Hz–210Hz frequency band. Based on the energy values ​​corresponding to the audio signal's distribution in the 80Hz–200Hz and 200Hz–500Hz frequency bands, the audio signal might be identified as a single-band sound (the specific identification process is described below). In this case, noise reduction can be performed on the 80Hz–200Hz band, and the sound in the 200Hz–500Hz band can be played. Alternatively, the audio signal might be identified as a non-single-band sound, i.e., a cross-band sound, in which case the sound in both the 80Hz–200Hz and 200Hz–500Hz bands would be played.

[0101] Regardless of which of the above situations occurs, it will affect the accuracy of noise identification in the embodiments of this application. Therefore, the above problem can be solved by setting a frequency difference C.

[0102] S103: Determine a target frequency band from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than a preset energy value.

[0103] In this embodiment, the energy value can be used to determine whether there is sound in different preset frequency bands, and then the number of preset frequency bands with sound can be used to identify whether the audio signal is noise. Specifically, if the energy value corresponding to a preset frequency band is not lower than the preset energy value, it can be determined that there is sound in the preset frequency band, and it can be marked as the target frequency band.

[0104] As an example, the target frequency band can be determined in the following three ways, which will be explained below.

[0105] Method 1: Set different preset energy values ​​E for each preset frequency band. 0i Except for E i The sum of the maximum energy value among all other energy values, and the first preset value A.

[0106] E 0i =E max +A

[0107] Among them, E 0i E represents the preset energy value corresponding to the i-th preset frequency band. max To exclude E i The highest energy value among all other energy values; E iLet A be the energy value corresponding to the i-th preset frequency band; A is the first preset value. If the audio signal is the original audio signal collected by the microphone, 100≤A≤150; if the audio signal is the audio signal after noise processing, for example, the audio signal output after noise reduction processing of the original audio signal using the existing scheme, 200≤A≤300.

[0108] See Figure 4 The spectrogram shown shows that part I is the speech signal. Compared with part II, the energy difference between the two is usually between 100 and 150. Therefore, the range of the first preset value can be set to 100≤A≤150. After noise reduction processing, the energy of part II will be reduced, and part I will become more prominent, resulting in a larger energy difference between the two. Therefore, the range of the first preset value can be set to 200≤A≤300.

[0109] Corresponding to this, when E i ≥E 0i When the target frequency band is determined, the i-th preset frequency band can be identified. Taking the above division into 7 frequency bands as an example, E1>max(E2+E3+E4+E... s When +E6+E7)+A, the preset frequency band 80Hz~250Hz is marked as the target frequency band.

[0110] Method 2: Set the same preset energy value E0 for all preset frequency bands: the sum of the energy values ​​corresponding to all preset frequency bands multiplied by the second preset value B.

[0111]

[0112] Among them, E i Let N be the energy value corresponding to the i-th preset frequency band; N is the number of preset frequency bands; and B is the second preset value, where 40% ≤ B ≤ 60%.

[0113] Corresponding to this, when E i When E1 ≥ E0, the i-th preset frequency band can be determined as the target frequency band. Taking the above division into 7 frequency bands as an example, when E1 > (E1+E2+E3+E4+E5+E6+E7)*B, the preset frequency band 80Hz~250Hz is marked as the target frequency band.

[0114] Method 3, preset energy values ​​include E 0i And E0.

[0115] Corresponding to this, when E i ≥E 0i And E i When E ≥ 0, the i-th preset frequency band can be determined as the target frequency band.

[0116] Specifically, we can first determine E i and E0i The size relationship, if E i ≥E 0i Then we can continue to judge E. i The relationship between E and E0, if E i If E ≥ 0, then the energy value E can be... i The corresponding preset frequency band i is marked as the target frequency band so that the number of frequency bands marked as the target frequency band can be counted later and noise identification can be performed accordingly.

[0117] S104: If the target frequency band is included among the plurality of preset frequency bands, and noise is determined to be included in the audio signal based on the number of target frequency bands, noise reduction processing is performed on the audio signal corresponding to the target frequency band.

[0118] Based on the above analysis, it can be seen that the speech signal of a person speaking usually exists across frequency bands and generally does not cover the entire frequency band. Therefore, the embodiments of this application can perform noise identification in the following manner:

[0119] If a target frequency band is included among multiple preset frequency bands, the audio signal can be determined to be a single-band sound and can be identified as noise. Alternatively, if multiple preset frequency bands are all marked as target frequency bands, the audio signal can be determined to be a full-band sound and can also be identified as noise. In this embodiment, full-band sound can be understood as sound that exists across all preset frequency bands. For example, pink noise and white noise can both be full-band sounds. In actual use, the sound of a finger tapping on a table may be identified as a full-band sound and noise is eliminated.

[0120] In this embodiment of the application, when it is determined that the audio signal contains noise, the audio signal corresponding to the marked target frequency band can be denoised, while the audio signal corresponding to other unmarked preset frequency bands can be directly output.

[0121] As an example, noise reduction of audio signals can be manifested as replacing the audio signal in the target frequency band with a comfortable tone, that is, smoothing the audio signal in the target frequency band, which helps to improve the user's audio listening experience.

[0122] Furthermore, to make the noise reduction effect more natural, the noise processing scheme of this application embodiment can also avoid suppressing background sounds under conditions without speech. Specifically, the sum of the energy values ​​corresponding to the energy values ​​of each of the multiple preset frequency bands can be obtained. The energy values ​​corresponding to the multiple preset frequency bands and the sum of the energy values ​​are used as input to the speech classification model to obtain the output result of the speech classification model. If the output result indicates that the audio signal includes a speech signal, the step of performing noise reduction processing on the audio signal corresponding to the target frequency band is then performed.

[0123] That is, when the speech classification model determines that the audio signal contains a speech signal, noise processing is then performed, which can correspond to the following three cases:

[0124] 1. If the output of the speech classification model indicates that the audio signal does not contain a speech signal, then a noise-free processing result can be returned without noise reduction processing;

[0125] 2. If the output of the speech classification model indicates that the audio signal contains a speech signal, and the audio signal is determined to be free of noise based on the energy value of the preset frequency band, then a noise-free processing result is returned, and no noise reduction processing is performed.

[0126] 3. If the output of the speech classification model indicates that the audio signal contains speech signals, and the audio signal is determined to contain noise based on the energy value of the preset frequency band, then the processing result containing noise is returned, and the audio signal corresponding to the target frequency band is subjected to noise reduction processing.

[0127] Understandably, the execution order of the speech classification model's judgment process and the preset frequency band's energy value judgment process in this application embodiment is not specifically limited. For example, the speech classification model's judgment process can be executed first, and when it is determined that the audio signal contains a speech signal, the preset frequency band's energy value judgment process can be executed then; or, the preset frequency band's energy value judgment process can be executed first, and when it is determined that the audio signal contains noise, the speech classification model's judgment process can be executed then; or, both of the above judgment processes can be executed simultaneously, and then the results of both judgments can be combined for noise reduction processing.

[0128] As an example, the speech classification model in this application embodiment can be a Gaussian Mixture Model (GMM), or other classification models can be selected according to the usage requirements. This application embodiment does not limit this.

[0129] In summary, this application provides a method for identifying whether an audio signal contains noise based on the energy distribution of the audio signal across multiple preset frequency bands by analyzing the characteristics of the frequency band distribution of the speech signal and the prompt noise. When noise is determined to be present in the audio signal, noise reduction processing is performed on the audio signal in the target frequency band where the noise is located. This method avoids directly performing overall noise reduction on the audio signal as in existing technologies, thus preventing the masking of low-energy speech signals and the retention of high-energy noise, which helps improve the noise processing effect. Furthermore, this application can identify the target frequency band where the noise is located and perform targeted noise reduction processing on the audio signal within that target frequency band, ensuring that the original noise is no longer present in the noise-reduced audio signal. This also helps improve the noise processing effect and optimizes the user's auditory experience.

[0130] It should be noted that, in addition to the noise generated by the terminal device, other devices besides the terminal device, such as a mobile phone placed next to the computer when the user is conducting an online meeting through a real-time communication application installed on the computer, may also generate noise as other devices. If the noise is identified as a single-frequency sound, it can also be canceled.

[0131] As an example, users can flexibly set the noise to be denoised according to their needs. That is, noise reduction can be performed on all target noise. Alternatively, noise reduction can be performed on the noise selected by the user from the target noise, based on usage requirements. Specifically, if the noise included in the audio signal is the noise to be processed, the denoising method provided in this application embodiment can be used; otherwise, even if the audio signal is determined to contain noise based on the number of target frequency bands, no denoising is required, and the audio can be played directly. With this method, prompts that the user does not want filtered out can be played normally.

[0132] Based on practical applications, it is known that users' needs for application notification sounds may vary in different real-time communication processes. Taking online meetings as an example, some meetings require maintaining a quiet atmosphere, in which case it is best not to play any application notification sounds aloud, and the noise reduction requirements for notification sounds are high; other meetings do not have high requirements for noise reduction of notification sounds, some noise is acceptable, and it is even desirable to retain some notification sounds that convey information. For example, in the current meeting, the sound of fingers tapping on the table can be used to convey the meaning of warnings and reminders. If this is identified as full-band noise and noise cancellation is performed, it will lead to the loss of effective information. Therefore, this application embodiment can provide a configuration scheme for the noise to be processed. Specifically, it can include:

[0133] A target noise that can be denoised is identified, and the identification information of the target noise is displayed on a preset page. The preset page provides operation options for selecting noise to be denoised from the target noise. After obtaining the noise to be denoised through the operation options, if it is determined that the audio signal contains noise and the noise contained in the audio signal belongs to the noise to be denoised, then the audio signal corresponding to the target frequency band is denoised.

[0134] In practical applications, the noise to be processed can be flexibly set during real-time communication, and this setting applies only to the current real-time communication process; alternatively, different types of real-time communication processes can be pre-set with their associated noise, and noise reduction can be performed based on the preset information. This approach improves the flexibility of noise reduction processing, making it more adaptable to user needs and enhancing the user experience.

[0135] As an example, the noise processing solution provided in this application can be applied to online meetings to reduce the noise of notification sounds generated during the meeting, thereby improving the meeting experience for participants. Specifically, the audio noise reduction method applied to online meetings may include:

[0136] First, when the first user is having an online meeting with at least one second user, the meeting client associated with the first user obtains the audio signal generated in the space where the first user is located.

[0137] The audio signal generated in the space where the first user is located may include at least one of the following sounds: prompts generated by other applications installed on the terminal device where the conference client is deployed; prompts generated by other devices besides the terminal device; the sound of the first user speaking in the conference; and the sound of other participants speaking in the space.

[0138] Secondly, it can be done according to Figure 2 The proposed method is used to process noise and obtain a noise-reduced audio file.

[0139] Specifically: extract the frequency domain features of the audio signal, and obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands based on the frequency domain features; determine a target frequency band from the multiple preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than a preset energy value; if the multiple preset frequency bands include the target frequency band, and it is determined that the audio signal contains noise based on the number of target frequency bands, perform noise reduction processing on the audio signal corresponding to the target frequency band.

[0140] Finally, the conference client associated with the first user plays the noise-reduced audio signal; and sends the noise-reduced audio signal to the conference client associated with at least one second user for audio playback.

[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0142] Corresponding to the foregoing method embodiments, this application also provides an audio signal noise processing apparatus, see below. Figure 5 The device may include:

[0143] The frequency domain feature extraction unit 201 is used to obtain the audio signal generated during real-time communication and extract the frequency domain features of the audio signal;

[0144] The energy value acquisition unit 202 is used to obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands according to the frequency domain characteristics.

[0145] The target frequency band determination unit 203 is used to determine a target frequency band from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than a preset energy value;

[0146] The noise processing unit 204 is used to include the target frequency band among the plurality of preset frequency bands, and to perform noise reduction processing on the audio signal corresponding to the target frequency band when it is determined that the audio signal contains noise based on the number of target frequency bands.

[0147] The device further includes:

[0148] The model output result acquisition unit is used to obtain the sum of the energy values ​​corresponding to the energy values ​​of the multiple preset frequency bands, and to use the energy values ​​corresponding to the multiple preset frequency bands and the sum of the energy values ​​as input to the speech classification model to obtain the output result of the speech classification model.

[0149] The noise processing unit can be specifically used to: when the output result indicates that the audio signal includes a speech signal, perform noise reduction processing on the audio signal corresponding to the target frequency band.

[0150] Specifically, the noise processing unit can be used to: if the output result indicates that the audio signal does not include a speech signal, then no noise processing is performed.

[0151] The device further includes:

[0152] The target noise display unit is used to identify target noise that can be noise-reduced, and to display the identification information of the target noise through a preset page. The preset page provides operation options for selecting noise to be processed from the target noise.

[0153] A noise acquisition unit is configured to acquire the noise to be processed via the operation options.

[0154] The noise processing unit can be specifically used to: if it is determined that the audio signal contains noise, and the noise contained in the audio signal belongs to the noise to be processed, then perform noise reduction processing on the audio signal corresponding to the target frequency band.

[0155] The device further includes:

[0156] The preset energy value acquisition unit is used to obtain the preset energy value E in the following manner. 0i and / or E0:

[0157] The preset energy value E corresponding to the i-th preset frequency band 0i For: except E i The sum of the maximum energy value among all other energy values, and the first preset value A; E i The value is the energy value corresponding to the i-th preset frequency band; if the audio signal is the original audio signal collected by the microphone, 100≤A≤150; if the audio signal is an audio signal that has been noise-processed, 200≤A≤300.

[0158] The preset energy value E0 is the product of the sum of the energy values ​​corresponding to all preset frequency bands and the second preset value B, where 40% ≤ B ≤ 60%.

[0159] The noise processing unit can be specifically used to: determine that the audio signal contains noise if one of the plurality of preset frequency bands includes a target frequency band.

[0160] The noise processing unit can be specifically used to: determine that the audio signal includes noise if all of the multiple preset frequency bands are target frequency bands.

[0161] Corresponding to the foregoing method embodiments, this application also provides a frequency band allocation device, see [link to relevant documentation]. Figure 6 The device may include:

[0162] The frequency band acquisition unit 301 is used to acquire the frequency bands corresponding to each of the multiple target noises, wherein the target noises are application prompt sounds.

[0163] The preset frequency band division unit 302 is used to determine the preset frequency band to be analyzed, and divide the preset frequency band to be analyzed into multiple preset frequency bands according to the distribution frequency bands corresponding to the multiple target noises, and each distribution frequency band is located within a preset frequency band.

[0164] The number N of the preset frequency bands is: 6≤N≤9.

[0165] The range of the preset frequency band to be analyzed is 80Hz to 4KHz.

[0166] Wherein, the single distributed frequency band is located within a preset frequency band, including: the upper limit frequency value of the distributed frequency band is less than the upper limit frequency value of the corresponding preset frequency band.

[0167] Corresponding to the foregoing method embodiments, this application also provides an audio noise reduction processing device for online meetings, which can be applied to a meeting client associated with a first user. See [link to relevant documentation]. Figure 7 The device may include:

[0168] The audio signal acquisition unit 401 is used to acquire the audio signal generated in the space where the first user is located when the first user is having an online meeting with at least one second user.

[0169] Frequency domain feature extraction unit 402 is used to extract the frequency domain features of the audio signal;

[0170] The energy value acquisition unit 403 is used to obtain the energy values ​​of the audio signal distributed in multiple preset frequency bands according to the frequency domain characteristics.

[0171] The target frequency band determination unit 404 is used to determine a target frequency band from the plurality of preset frequency bands, wherein the energy value corresponding to the target frequency band is not lower than a preset energy value;

[0172] The noise processing unit 405 is used to include the target frequency band among the plurality of preset frequency bands, and to perform noise reduction processing on the audio signal corresponding to the target frequency band when it is determined that the audio signal contains noise based on the number of target frequency bands.

[0173] The audio playback unit 406 is used to play the noise-reduced audio signal and send the noise-reduced audio signal to the conference client associated with the at least one second user for audio playback.

[0174] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0175] And an electronic device, comprising:

[0176] One or more processors; and

[0177] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.

[0178] in, Figure 8 The architecture of an electronic device is illustrated by example. For instance, device 1500 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, aircraft, etc.

[0179] Reference Figure 8The device 1500 may include one or more of the following components: processing component 1502, memory 1504, power supply component 1506, multimedia component 1508, audio component 1510, input / output (I / O) interface 1512, sensor component 1514, and communication component 1516.

[0180] Processing component 1502 typically controls the overall operation of device 1500, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 1502 may include one or more processors 1520 to execute instructions to perform all or part of the steps of the methods provided in this disclosure. Furthermore, processing component 1502 may include one or more modules to facilitate interaction between processing component 1502 and other components. For example, processing component 1502 may include a multimedia module to facilitate interaction between multimedia component 1508 and processing component 1502.

[0181] Memory 1504 is configured to store various types of data to support the operation of device 1500. Examples of this data include instructions for any application or method operating on device 1500, contact data, phonebook data, messages, pictures, videos, etc. Memory 1504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0182] Power supply component 1506 provides power to various components of device 1500. Power supply component 1506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 1500.

[0183] Multimedia component 1508 includes a screen that provides an output interface between device 1500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 1508 includes a front-facing camera and / or a rear-facing camera. When device 1500 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0184] Audio component 1510 is configured to output and / or input audio signals. For example, audio component 1510 includes a microphone (MIC) configured to receive external audio signals when device 1500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 1504 or transmitted via communication component 1516. In some embodiments, audio component 1510 also includes a speaker for outputting audio signals.

[0185] I / O interface 1512 provides an interface between processing component 1502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0186] Sensor assembly 1514 includes one or more sensors for providing status assessments of various aspects of device 1500. For example, sensor assembly 1514 may detect the on / off state of device 1500, the relative positioning of components such as the display and keypad of device 1500, changes in the position of device 1500 or a component of device 1500, the presence or absence of user contact with device 1500, the orientation or acceleration / deceleration of device 1500, and temperature changes of device 1500. Sensor assembly 1514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1514 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0187] Communication component 1516 is configured to facilitate wired or wireless communication between device 1500 and other devices. Device 1500 can access wireless networks based on communication standards, such as WiFi, or mobile communication networks such as 2G, 3G, 4G / LTE, and 5G. In one exemplary embodiment, communication component 1516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 1516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0188] In an exemplary embodiment, device 1500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0189] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1504 including instructions, which can be executed by a processor 1520 of device 1500 to perform the method provided by the present disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0190] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0191] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0192] The noise processing solution provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method of noise processing of an audio signal, characterized in that, The method comprises: obtaining an audio signal generated in a real-time communication process, extracting a frequency domain feature of the audio signal; pre-dividing a plurality of preset frequency bands, and obtaining energy values corresponding to the distribution of the audio signal in the plurality of preset frequency bands according to the frequency domain feature; determining a target frequency band from the plurality of preset frequency bands, the energy value corresponding to the target frequency band being not lower than a preset energy value; if the plurality of preset frequency bands include the target frequency band, and it is determined that the audio signal includes noise according to the number of target frequency bands, performing noise reduction processing on the audio signal corresponding to the target frequency band, so that the noise is no longer included in the audio signal after noise reduction processing; the determination that the audio signal includes noise according to the number of target frequency bands comprises: if one target frequency band is included in the plurality of preset frequency bands, it is determined that the audio signal includes prompt tone noise.

2. The method of claim 1, wherein, Further comprising: obtaining the sum of the energy values corresponding to the plurality of preset frequency bands, and taking the energy values corresponding to the plurality of preset frequency bands and the sum of the energy values as inputs of a speech classification model to obtain an output result of the speech classification model; if the output result indicates that the audio signal includes a speech signal, performing the step of performing noise reduction processing on the audio signal corresponding to the target frequency band.

3. The method of claim 1, wherein, The method further comprises: determining a target noise that can be subjected to noise reduction processing, and displaying identification information of the target noise through a preset page, the preset page providing an operation option for selecting a noise to be processed from the target noise; after the noise to be processed is obtained through the operation option, if it is determined that the audio signal includes noise and the noise included in the audio signal belongs to the noise to be processed, performing noise reduction processing on the audio signal corresponding to the target frequency band.

4. The method according to any one of claims 1 to 3, characterized in that, said preset energy value is E 0i and / or E0, said preset energy value is obtained in the following way: The preset energy value E corresponding to the ith preset frequency band 0i The maximum energy value in the other energy values except E i The sum of the maximum energy value in the other energy values except E and the first preset value A; E i The energy value corresponding to the ith preset frequency band; if the audio signal is a raw audio signal collected by a microphone, 100≤A≤150; if the audio signal is a noise-processed audio signal, 200≤A≤300 The preset energy value E0 is the product of the sum of the energy values corresponding to all preset frequency bands and a second preset value B, and 40%≤B≤60%.

5. The method according to any one of claims 1 to 3, characterized in that, The determination that the audio signal includes noise according to the number of target frequency bands comprises: if all the plurality of preset frequency bands are target frequency bands, it is determined that the audio signal includes full-band noise.

6. A frequency band division method characterized by comprising: The method comprises: obtaining a plurality of distribution frequency bands corresponding to a plurality of target noises, the target noises being application prompt tones; determining a preset frequency band to be analyzed, and dividing the preset frequency band to be analyzed into a plurality of preset frequency bands according to the distribution frequency bands corresponding to the plurality of target noises, so as to obtain an audio signal generated in a real-time communication process, calculate energy values corresponding to the distribution of the audio signal in the plurality of preset frequency bands, determine a preset frequency band with an energy value not lower than a preset energy value as a target frequency band, and then perform noise reduction processing on the audio signal corresponding to the target frequency band when it is determined that the audio signal includes noise according to the number of target frequency bands, so that the noise is no longer included in the audio signal after noise reduction processing; wherein each distribution frequency band is located in a preset frequency band; the determination that the audio signal includes noise according to the number of target frequency bands comprises: if one target frequency band is included in the plurality of preset frequency bands, it is determined that the audio signal includes prompt tone noise.

7. The method of claim 6, wherein the number N of the preset frequency bands is 6≤N≤9.

8. The method of claim 6 or 7, wherein the preset frequency band to be analyzed ranges from 80 Hz to 4 kHz. The single distribution frequency band is located in a preset frequency band, including: The upper limit frequency value of the distribution frequency band is less than the upper limit frequency value of the corresponding preset frequency band.

9. The method of claim 6, wherein, Including: When the first user has an online conference with at least one second user, the conference client associated with the first user obtains the audio signal generated in the space where the first user is located; 10. An audio noise reduction processing method of an online conference, characterized by, Extract the frequency domain features of the audio signal; Pre-divide a plurality of preset frequency bands, and obtain the energy values of the audio signal corresponding to the distribution in the plurality of preset frequency bands according to the frequency domain features; Determine a target frequency band from the plurality of preset frequency bands, and the energy value corresponding to the target frequency band is not less than a preset energy value; If the plurality of preset frequency bands include the target frequency band, and it is determined that the audio signal includes noise according to the number of target frequency bands, the audio signal corresponding to the target frequency band is processed to reduce noise, so that the noise is no longer included in the audio signal after noise reduction processing; Play the audio signal after noise reduction processing, and send the audio signal after noise reduction processing to the conference client associated with the at least one second user for audio playback; The determination that the audio signal includes noise according to the number of target frequency bands includes: if the plurality of preset frequency bands include a target frequency band, it is determined that the audio signal includes prompt tone noise. Including: The frequency domain feature extraction unit is configured to obtain an audio signal generated in a real-time communication process, and extract frequency domain features of the audio signal; 11. A noise processing apparatus of an audio signal, characterized by, The energy value obtaining unit is configured to pre-divide a plurality of preset frequency bands, and obtain energy values of the audio signal corresponding to the distribution in the plurality of preset frequency bands according to the frequency domain features; The target frequency band determination unit is configured to determine a target frequency band from the plurality of preset frequency bands, and the energy value corresponding to the target frequency band is not less than a preset energy value; The noise processing unit is configured to, if the plurality of preset frequency bands include the target frequency band, and it is determined that the audio signal includes noise according to the number of target frequency bands, process the audio signal corresponding to the target frequency band to reduce noise, so that the noise is no longer included in the audio signal after noise reduction processing. The noise processing unit is specifically configured to, if the plurality of preset frequency bands include a target frequency band, it is determined that the audio signal includes prompt tone noise. Including: The distribution frequency band obtaining unit is configured to obtain a plurality of distribution frequency bands corresponding to a plurality of target noises respectively, and the target noise is an application prompt tone.

12. A frequency band dividing apparatus characterized by comprising: ​ ​ The preset frequency band division unit is configured to determine a preset to-be-analyzed frequency band, and divide the preset to-be-analyzed frequency band into a plurality of preset frequency bands according to the distribution frequency bands corresponding to the plurality of target noises respectively, so as to obtain an audio signal generated in a real-time communication process, calculate energy values of the audio signal corresponding to the plurality of preset frequency bands, and determine a preset frequency band with an energy value not lower than a preset energy value as a target frequency band, and then determine that the audio signal includes noise according to the number of the target frequency bands, and perform noise reduction processing on the audio signal corresponding to the target frequency band, so that the noise is not included in the audio signal after the noise reduction processing. Each distribution frequency band is located in a preset frequency band. The determination that the audio signal includes noise according to the number of the target frequency bands includes: if one target frequency band is included in the plurality of preset frequency bands, it is determined that the audio signal includes prompt tone noise.

13. An audio noise reduction processing device for online meetings, characterized in that, The apparatus is applied to a conference client associated with a first user, and includes: An audio signal obtaining unit configured to obtain an audio signal generated in a space where the first user is located when the first user and at least one second user perform an online conference. A frequency domain feature extracting unit configured to extract a frequency domain feature of the audio signal. An energy value obtaining unit configured to pre-divide a plurality of preset frequency bands, and obtain energy values of the audio signal corresponding to the plurality of preset frequency bands according to the frequency domain feature. A target frequency band determining unit configured to determine a target frequency band from the plurality of preset frequency bands, the target frequency band corresponding to an energy value not lower than a preset energy value. A noise processing unit configured to, when the plurality of preset frequency bands include the target frequency band, and according to the number of the target frequency bands, determine that the audio signal includes noise, perform noise reduction processing on the audio signal corresponding to the target frequency band, so that the noise is not included in the audio signal after the noise reduction processing. An audio playing unit configured to play the audio signal after the noise reduction processing as audio, and send the audio signal after the noise reduction processing to a conference client associated with the at least one second user to play the audio signal as audio. The noise processing unit is specifically configured to, if one target frequency band is included in the plurality of preset frequency bands, determine that the audio signal includes prompt tone noise.

14. An electronic device, comprising: One or more processors; and A memory associated with the one or more processors, the memory being configured to store program instructions, the program instructions being configured to perform the steps of the method of any one of claims 1 to 10 when read and executed by the one or more processors. ​ ​

Citation Information

Patent Citations

  • Denoising method and mobile terminal

    CN111125423A

  • Adaptive Energy Limiting for Transient Noise Suppression

    US20210151065A1

  • Selective noise cancellation

    US20220238091A1