Audio noise reduction method, computer device, storage medium and product

CN122821979APending Publication Date: 2026-09-25ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610938921.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]在各类语音交互应用场景中,比如日常移动通信通话、网络视频通话、语音助手交互、直播录制等场景中,环境噪声(如风噪、电流声、键盘敲击声、公共场所嘈杂声)会严重影响语音信号的清晰度,导致语音交互质量下降,用户体验变差的问题

Benefits of technology

[0008]第四方面,提供了一种计算机程序产品,所述计算机程序产品包括存储在非暂态计算机可读存储介质上的计算机程序,所述计算机程序包括程序指令,当所述程序指令被计算机执行时,使所述计算机执行以实现上述的音频降噪方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821979A_ABST
    Figure CN122821979A_ABST
Patent Text Reader

Abstract

The application relates to an audio noise reduction method, a computer device, a storage medium and a product, and relates to the technical field of audio processing. The method comprises the following steps: performing keyword recognition on a voice stream to obtain a keyword recognition result; in the case where the keyword recognition result indicates that a preset keyword is not monitored, performing noise suppression on the voice stream by using a first noise suppression strategy; in the case where the keyword recognition result indicates that the preset keyword is monitored, performing noise suppression on the voice stream by using a second noise suppression strategy; the noise reduction strength of the first noise suppression strategy is lower than the noise reduction strength of the second noise suppression strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio noise reduction method, computer device, storage medium and product. Background Technology

[0002] In various voice interaction application scenarios, such as daily mobile communication calls, online video calls, voice assistant interactions, and live recording, environmental noise (such as wind noise, electrical noise, keyboard typing sounds, and public noise) can seriously affect the clarity of voice signals, leading to a decline in voice interaction quality and a worse user experience.

[0003] Existing adaptive noise reduction technology passively adjusts the noise reduction intensity based on the ambient volume, failing to perceive the user's actual voice interaction needs, resulting in poor noise reduction performance. Summary of the Invention

[0004] This application provides an audio noise reduction method, computer device, storage medium, and product that can meet the needs of actual voice interaction and improve the noise reduction effect. The technical solution is as follows.

[0005] Firstly, an audio noise reduction method is provided, the method comprising: Keyword recognition is performed on the speech stream to obtain the keyword recognition results; If the keyword recognition result indicates that no preset keyword has been detected, the first noise suppression strategy is used to suppress noise in the speech stream; When the keyword recognition result indicates that the preset keyword has been detected, a second noise suppression strategy is used to suppress noise in the speech stream; the noise reduction intensity of the first noise suppression strategy is lower than that of the second noise suppression strategy.

[0006] In a second aspect, a computer device is provided, the computer device comprising a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the above-described audio noise reduction method.

[0007] Thirdly, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described audio noise reduction method.

[0008] Fourthly, a computer program product is provided, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to execute to implement the above-described audio noise reduction method.

[0009] The audio noise reduction method provided in this application performs keyword recognition on the speech stream and uses a corresponding noise suppression strategy to suppress noise in the speech stream based on the obtained keyword recognition results. Specifically, when no preset keywords are detected, a first noise suppression strategy with lower noise reduction intensity is used, and when preset keywords are detected, a second noise suppression strategy with higher noise reduction intensity is used. This allows the noise reduction intensity to be dynamically adjusted according to whether the user emits key speech content, thereby meeting the actual voice interaction needs, improving the noise reduction effect, and ultimately enhancing the voice interaction quality and user experience in various voice interaction scenarios.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] Figure 1 A flowchart of an audio noise reduction method provided in an exemplary embodiment of this application is shown; Figure 2 This illustration shows a flowchart of noise suppression for a speech stream provided in an exemplary embodiment of this application; Figure 3 This invention illustrates a schematic diagram of the structure of an audio noise reduction model provided in an exemplary embodiment of this application. Figure 4 A schematic diagram of an audio noise reduction process provided in an exemplary embodiment of this application is shown; Figure 5 A structural block diagram of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0014] This application provides an audio noise reduction method that can sense the user's actual voice interaction needs and improve the audio noise reduction effect in voice interaction. Figure 1A flowchart of an audio noise reduction method provided in an exemplary embodiment of this application is shown. This method can be executed by a computer device, which can be implemented as a terminal device; as... Figure 1 As shown, the audio noise reduction method may include the following steps.

[0015] Step 110: Perform keyword recognition on the speech stream to obtain the keyword recognition results.

[0016] The voice stream is the audio signal generated in a voice interaction scenario, including the audio signal received by the terminal device and the audio signal picked up by the terminal device's microphone. The terminal device can correspond to either the receiver or the sender in the voice interaction. The source form of this voice stream varies depending on the application scenario. For example, when the terminal device is the receiver, in a mobile call scenario, the voice stream is the downlink audio signal sent by the other end of the call; in a web conferencing scenario, it is the audio signal transmitted from other participants to the local device; in a smart home scenario, it is the user's voice command signal continuously picked up by the microphone of a smart home device, etc. When the terminal device is the sender, in a mobile call scenario, the voice stream is... The audio stream is the uplink audio signal collected and uploaded by the sender's terminal microphone. In a web conferencing scenario, the voice stream is the shared audio signal collected by the microphones of the participants' terminals. Correspondingly, different keywords can be preset for different application scenarios, and different keywords can be recognized based on different application scenarios. For example, in a mobile call scenario, natural spoken expressions that indicate call difficulties can be preset as keywords, such as "hello," "can't hear me," "say it again," etc. In a web conferencing scenario, guiding words can be preset as keywords, such as "everyone pay attention," "summarize," etc. In a smart home scenario, wake words and control command words can be preset as keywords, and so on. This application embodiment does not limit the source form of the voice stream or the preset keywords for different application scenarios. In practical applications, the corresponding keyword list can be flexibly configured according to the application scenario.

[0017] After acquiring the speech stream, the computer device can use a pre-trained speech recognition model to perform keyword recognition on the speech stream in order to obtain keyword recognition results for decision-making on subsequent noise suppression strategies. The speech recognition model can be trained based on speech stream samples and corresponding keyword labels. The speech stream samples can come from different application scenarios to improve the applicability of the speech recognition model.

[0018] Step 120: If the keyword recognition result indicates that no preset keyword has been detected, the first noise suppression strategy is used to suppress noise in the speech stream.

[0019] When the keyword recognition result indicates that no preset keyword is detected in the current voice stream, it means that the user has not yet emitted any key voice content that requires high clarity. At this time, the first noise suppression strategy is used to suppress noise in the voice stream. This first noise suppression strategy is a low-intensity steady-state basic noise reduction strategy that takes into account both noise reduction effect and device power consumption. It uses mild, fidelity-first, and low-damage parameters, relies on the local signal processing unit to suppress normal background noise, maintain the basic clarity of human voice, and avoid damage to voice quality due to excessive noise reduction.

[0020] Step 130: When the keyword recognition result indicates that a preset keyword has been detected, a second noise suppression strategy is used to suppress noise in the speech stream; the noise reduction intensity of the first noise suppression strategy is lower than that of the second noise suppression strategy.

[0021] When the keyword recognition result indicates that a preset keyword has been detected in the current speech stream, it means that the user has delivered key semantic content that requires a high degree of clarity. The current environment experiences a sudden deterioration in the signal-to-noise ratio, and the basic noise reduction is insufficient to support clear communication. It can be determined that the semantics meet the criteria for formally triggering enhanced noise reduction. At this time, a second noise suppression strategy is adopted to fully suppress various types of noise in the speech stream, extracting and preserving the clean human voice to the greatest extent possible, so as to ensure the clarity of key speech content and ensure that the recipient of the voice interaction can accurately receive the semantic information expressed by the user.

[0022] The first noise suppression strategy and the second noise suppression strategy can be different degrees of the same noise suppression scheme. That is, they can use the same noise reduction algorithm or model architecture, but take different values ​​for key control parameters, thereby achieving different degrees of noise reduction.

[0023] In one possible implementation, the first noise suppression strategy and the second noise suppression strategy can suppress noise based on the predicted time-frequency mask output by the same audio denoising model. The difference lies in the different decision thresholds used to determine the lowest probability value belonging to human voice. The first noise suppression strategy uses a first decision threshold, and the second noise suppression strategy uses a second decision threshold. The second decision threshold is higher than the first decision threshold corresponding to the first noise suppression strategy.

[0024] In another possible implementation, the first noise suppression strategy and the second noise suppression strategy can suppress noise based on the predicted time-frequency mask output by the same audio denoising model. The difference lies in the attenuation coefficient for the time-frequency units that are identified as noise. The attenuation coefficient is used to control the degree of noise subtraction from the noisy speech. The first noise suppression strategy uses the first attenuation coefficient to subtract a lower proportion of noisy time-frequency units from the time-frequency feature map to achieve mild noise reduction. The second noise suppression strategy uses the second attenuation coefficient to subtract a higher proportion of noisy time-frequency units from the time-frequency feature information to achieve deep noise reduction. The second attenuation coefficient is greater than the first attenuation coefficient.

[0025] By adjusting different control parameters within the same noise reduction scheme, the noise reduction intensity can be flexibly switched, ensuring simplicity while avoiding the increased system complexity caused by maintaining multiple noise reduction schemes.

[0026] In summary, the audio noise reduction method provided in this application performs keyword recognition on the speech stream and uses corresponding noise suppression strategies to suppress noise in the speech stream based on the obtained keyword recognition results. Specifically, when no preset keywords are detected, a first noise suppression strategy with lower noise reduction intensity is adopted, and when preset keywords are detected, a second noise suppression strategy with higher noise reduction intensity is adopted. This allows the noise reduction intensity to be dynamically adjusted according to whether the user emits key speech content, thereby meeting the actual voice interaction needs, improving the noise reduction effect, and ultimately enhancing the voice interaction quality and user experience in various voice interaction scenarios.

[0027] Based on the above embodiments: Taking the use of an audio noise reduction model to suppress noise in a speech stream as an example, Figure 2 This application illustrates a flowchart of noise suppression for a speech stream according to an exemplary embodiment, as shown below. Figure 2 As shown, noise suppression of a speech stream includes the following steps.

[0028] S201, Perform a short-time Fourier transform on the speech stream to obtain the time-frequency feature information of the speech stream.

[0029] This speech stream is a mixed time-domain audio stream containing target human voice, environmental noise, and interfering speech. Through audio framing and frequency domain structure based on short-time Fourier transform technology, the time-domain speech waveform can be transformed into a two-dimensional time-frequency spectral feature map as time-frequency feature information—a lightweight time-frequency feature map. The mixed and superimposed human voice, environmental noise, reverberation, and echo are decomposed into independent time-frequency units (time-frequency points in the time-frequency feature map), ensuring that each frequency category has an independent, distinguishable, and processable feature dimension. Invalid high-frequency bands are truncated, amplitude is normalized, and feature dimensions are compressed. This frequency domain deconstruction upgrades the processing from a unified whole-segment approach to fine-grained frequency-point processing, providing underlying principle support for subsequent human voice separation and differentiated noise reduction.

[0030] S202, input the time-frequency feature information into the audio noise reduction model to obtain the predicted time-frequency mask, which is used to characterize the probability that each time-frequency unit in the time-frequency feature information belongs to human voice.

[0031] The audio denoising model can be pre-trained using a training sample set constructed from speech data in various voice interaction scenarios. During sample construction, noisy speech samples are generated by randomly mixing clean human voice samples, noise samples from multiple scenarios, and interfering human voice samples. Clean human voice samples are used as voice labels for supervised training. All noisy speech samples are fed into a preprocessing layer after being configured with the same sampling parameters, generating lightweight time-frequency feature maps in batches, which serve as input data for the network training phase. Upon receiving the input data, the audio denoising model performs feature extraction, mask generation, and human voice output prediction through forward propagation. Then, a loss function is calculated based on the predicted human voice and voice labels. The model parameters are updated in reverse based on the loss function calculation results. This forward propagation and loss backpropagation process is repeated to train the audio denoising model. When the training completion conditions are met, a trained audio denoising model is obtained. These conditions may include at least one of the following: the human voice separation accuracy of the audio denoising model reaches an accuracy threshold, the number of training iterations reaches an iteration threshold, and the model converges. In the process of accuracy verification, the voice separation performance can be monitored through the validation set. When the mask prediction accuracy and voice separation accuracy no longer improve, the model is considered to have converged.

[0032] After training the audio noise reduction model, the audio noise reduction model is applied to the voice interaction scenario. In this embodiment, since different noise suppression strategies need to be adopted according to different keyword recognition results, it is necessary to first generate a prediction time-frequency mask through the audio noise reduction model, and then apply the corresponding noise suppression strategy in combination with the prediction time-frequency mask to perform audio separation.

[0033] The process of generating a predictive time-frequency mask through the audio noise reduction model (i.e., S201-S202) can be performed in parallel with the keyword recognition process, or it can be performed first or later. This application embodiment does not limit this.

[0034] S203, based on the noise suppression strategy corresponding to the keyword recognition result and the predicted time-frequency mask, processes the time-frequency feature information to obtain the separated human voice audio.

[0035] After determining the noise suppression strategy and the predicted time-frequency mask, the computer device can adjust the predicted time-frequency mask according to the adjustment method indicated by the noise suppression strategy, and then process the time-frequency feature information based on the adjusted time-frequency mask to achieve the separation between human voice audio and noise audio.

[0036] S204 outputs human voice audio.

[0037] After acquiring the human voice audio, the human voice audio is output to the receiver of the voice interaction.

[0038] As one possible implementation, the audio noise reduction model can be a lightweight network structure with small convolutional kernels and few hidden layers. It can be configured on terminal devices to perform real-time noise reduction processing on the voice stream, avoiding network transmission delays, bandwidth consumption, and data privacy leakage risks caused by interaction with the server. Thus, while ensuring the construction effect, it takes into account real-time performance, security, and deployment flexibility.

[0039] As one possible implementation, the audio noise reduction model includes a first feature extraction branch, a second feature extraction branch, a feature fusion layer, and a mask prediction layer. Since human voice has a stable harmonic periodic structure in the frequency domain, consisting of a fundamental frequency and its harmonics in an orderly manner, its spectral distribution is regular and its temporal continuity is strong. In contrast, environmental noise lacks a fixed harmonic distribution, its spectrum is chaotic and disordered, and its energy exhibits random jumps in the temporal domain. Therefore, a feature extraction branch can be constructed using a CNN (Convolutional Neural Network) + RNN (Recurrent Neural Network) architecture, namely, the CNN feature extraction branch (i.e., the first feature extraction branch) and the RNN feature extraction branch (i.e., the second feature extraction branch). The lightweight CNN feature extraction branch can distinguish the frequency domain features between human voice and noise, the lightweight RNN feature extraction branch can capture temporal changes in speech, the feature fusion layer can achieve optimized combination of spatiotemporal features, and the mask prediction layer can locate the effective time-frequency region of human voice.

[0040] Based on the structure of the above audio denoising model, time-frequency feature information is input into the audio denoising model to obtain the predicted time-frequency mask, including: Frequency domain features are extracted from the time-frequency features through the first feature extraction branch to obtain frequency domain features. Temporal feature information is obtained by extracting temporal features from frequency domain feature information through the second feature extraction branch; By fusing frequency domain features and temporal features through a feature fusion layer, spatiotemporal joint fused features are obtained. The spatiotemporal joint fusion features are mapped by the mask prediction layer to obtain the prediction time-frequency mask.

[0041] Figure 3 A schematic diagram of the structure of an audio noise reduction model provided in an exemplary embodiment of this application is shown, as follows: Figure 3 As shown, the audio noise reduction model includes a first feature extraction branch 310, a second feature extraction branch 320, a feature fusion layer 330, and a mask prediction layer 340. The first feature extraction branch 310, i.e., the CNN feature extraction branch, can employ a lightweight shallow convolutional network to learn local frequency domain features from time-frequency feature information, extracting the frequency band energy distribution and the frequency domain differences between human voice and noise, and outputting a frequency domain spatial feature map representing the static frequency domain structure information of the speech, as the frequency domain feature information. The second feature extraction branch 320, i.e., the RNN feature extraction branch, can expand the frequency domain feature information output by the CNN feature extraction branch along the time axis into a temporal feature sequence, and model the inter-frame temporal characteristics of the audio through a lightweight RNN layer. Sequence association, speech prosody features, dynamic noise variation patterns, capture the temporal dimension variation patterns and long-term dependencies in the sequence, and output a temporal context feature sequence that represents the dynamic variation information of speech, i.e., temporal feature information; after the frequency domain feature information output by the first feature extraction branch and the temporal feature information output by the second feature extraction branch are fused by the feature fusion layer 330, a spatiotemporal joint fusion feature containing both frequency domain detail information and temporal global information is obtained. The spatiotemporal joint fusion feature is mapped by the fully connected layer in the mask prediction layer 340 in combination with the Sigmoid activation function to obtain a predicted time-frequency mask, which is used to distinguish the human voice time-frequency feature information and the noise time-frequency feature information in the time-frequency feature information.

[0042] By implementing lightweight configurations for the audio noise reduction model architecture, lightweight deployment on terminal devices can be achieved, adapting to the limited computing resources of these devices.

[0043] As one possible implementation, the noise suppression strategy indicates that the adjustment of the predicted time-frequency mask is to adjust the decision threshold of the predicted time-frequency mask. By adjusting the decision threshold, the classification of noise time-frequency units can be changed. Time-frequency units corresponding to mask values ​​higher than the decision threshold are retained as human voice time-frequency units, while time-frequency units corresponding to mask values ​​lower than the decision threshold are deleted as noise time-frequency units. This approach has low computational overhead and is suitable for terminal devices with limited computing resources.

[0044] As another possible implementation, the noise suppression strategy indicates that the adjustment of the predicted time-frequency mask is to adjust the decision threshold of the predicted time-frequency mask. By adjusting the decision threshold, the classification of noise time-frequency units can be changed. Time-frequency units corresponding to mask values ​​higher than the decision threshold are retained as human voice time-frequency units or enhanced, while time-frequency units corresponding to mask values ​​lower than the decision threshold are attenuated as noise time-frequency units.

[0045] For example, when employing the first noise suppression strategy, based on the noise suppression strategy corresponding to the keyword recognition result and the predicted time-frequency mask, the time-frequency feature information is processed to obtain the separated human voice audio and noise audio, including: When the first noise suppression strategy is adopted, the predicted time-frequency mask is determined and adjusted based on the first determination threshold to obtain the first time-frequency mask; the determination threshold is used to indicate the lowest probability value of the time-frequency unit belonging to human voice. The first human voice time-frequency feature information is obtained by performing mask weighting operation on the time-frequency feature information based on the first time-frequency mask. The audio signal is reconstructed from the time-frequency characteristic information of the first human voice to obtain the separated first human voice audio.

[0046] If, after determining the classification of the corresponding time-frequency unit based on the mask value and the first determination threshold, the noise time-frequency unit is deleted and the human voice time-frequency unit is retained, then after the determination of the first determination threshold, the mask values ​​in the predicted time-frequency mask that are less than the first determination threshold can be set to 0, and the mask values ​​that are greater than or equal to the first determination threshold can be set to 1, thus obtaining the first time-frequency mask.

[0047] If, after classifying the corresponding time-frequency unit based on the mask value and the first determination threshold, further attenuation or enhancement processing is required, after the determination of the first determination threshold, the mask values ​​in the predicted time-frequency mask that are less than the first determination threshold can be multiplied by a first attenuation coefficient to reduce the probability that the corresponding time-frequency unit is obtained as human voice, and the mask values ​​that are greater than or equal to the first determination threshold are retained to obtain the first time-frequency mask; or, the mask values ​​in the predicted time-frequency mask that are less than the first determination threshold can be multiplied by the first attenuation coefficient, and the mask values ​​that are greater than or equal to the first determination threshold can be multiplied by a first enhancement coefficient to increase the probability that the corresponding time-frequency unit is obtained as human voice, thus obtaining the first time-frequency mask; by using a lower first determination threshold to classify the time-frequency unit as human voice and noise and making corresponding adjustments, the naturalness and environmental details of the original speech can be preserved while ensuring the basic noise reduction effect.

[0048] Furthermore, when the first noise suppression strategy is adopted, the first determination threshold can be set to 0. In this case, the mask values ​​of all time-frequency units are greater than or equal to the first determination threshold, and they are all determined to be human voices and preserved. This predicted time-frequency mask is the first time-frequency mask. Based on the first time-frequency mask, the time-frequency feature information is subjected to mask weighting processing to obtain the first human voice time-frequency feature information extracted from the time-frequency feature information by the audio noise reduction model.

[0049] After obtaining the first human voice time-frequency feature information, the audio signal of the first human voice time-frequency feature information is reconstructed to restore it into a time-domain audio signal for playback or transmission. The audio signal reconstruction may refer to performing inverse short-time Fourier transform, frame overlap splicing, and waveform restoration operations on the first human voice time-frequency feature information to convert it from frequency domain representation back to time domain waveform signal, thereby obtaining the separated first human voice audio and avoiding human voice distortion and timbre loss caused by additional processing.

[0050] Human voice noise is classified using a lower first judgment threshold. The time-frequency domain of sudden strong interference areas with low mask values ​​and clear characteristics is defined as the noise domain. Weak signals at the edge of human voice, the transition area between human voice and noise, slight ambient noise, and low-energy intermittent noise are identified as human voice domain and preserved. For time-frequency units in the noise domain, they are deleted or attenuated with a small proportion of attenuation coefficient to reduce the peak noise energy without deep suppression, so as to preserve the original sound field as much as possible.

[0051] The noise suppression strength of the second noise suppression strategy is higher than that of the first noise suppression strategy, which can be manifested by a relatively high decision threshold. As a possible implementation, based on the noise suppression strategy corresponding to the keyword recognition result and the predicted time-frequency mask, the time-frequency feature information is processed to obtain the separated human voice audio and noise audio, including: When the second noise suppression strategy is adopted, the predicted time-frequency mask is determined and adjusted based on the second determination threshold to obtain the second time-frequency mask; the second determination threshold is higher than the first determination threshold corresponding to the first noise suppression strategy. The second human voice time-frequency feature information is obtained by performing mask weighting operation on the time-frequency feature information based on the second time-frequency mask. The audio signal is reconstructed from the time-frequency characteristic information of the second human voice to obtain the separated second human voice audio.

[0052] The values ​​of the first and second judgment thresholds can be set based on actual needs. The process of using the second noise suppression strategy to adjust the predicted time-frequency mask, perform mask weighting operations, and reconstruct audio signals is similar to that of using the first noise suppression strategy. The difference is that by using the second judgment threshold to screen human voice time-frequency units, a purer human voice domain can be retained. All time-frequency units in the noise-noise transition area, weak speech area, intermittent noise area, and low-energy aliasing area retained in the first noise suppression strategy are classified as noise domains. Furthermore, a second screening of the entire time-frequency feature spectrum can be performed to uniformly classify implicit noise such as dynamic noise trails, floating background noise, and edge interference signals into the noise domain, achieving seamless screening and expanding the range of noise filtering or attenuation processing. Furthermore, when attenuation or enhancement processing is required, the attenuation or enhancement coefficient is relatively high. That is, after the judgment of the second judgment threshold, the mask values ​​in the predicted time-frequency mask that are less than the second judgment threshold can be multiplied by the second attenuation coefficient, and the mask values ​​that are greater than or equal to the second judgment threshold can be retained to obtain the second time-frequency mask; or, the mask values ​​in the predicted time-frequency mask that are less than the second judgment threshold can be multiplied by the second attenuation coefficient, and the mask values ​​that are greater than or equal to the second judgment threshold can be multiplied by the second enhancement coefficient to obtain the second time-frequency mask; wherein, the second attenuation coefficient is higher than the first attenuation coefficient, and the second enhancement coefficient is higher than the first enhancement coefficient; through enhanced noise suppression, the mask value corresponding to the noise domain can be adjusted to approach 0. After performing mask weighting operation on the time-frequency feature information based on the second time-frequency mask, the energy of the noise signal is deeply suppressed after point-by-point multiplication operation, and the amplitude of background noise, continuous interference, and impulse noise is greatly reduced, achieving an overall noise zeroing effect.

[0053] Furthermore, when using the second noise suppression strategy for noise suppression, the computer device can also increase the gain of the human voice frequency band to ensure clear human voice; wherein the increase in the human voice frequency band gain can be set based on actual needs, and this application embodiment does not limit this.

[0054] It should be noted that, in the above noise suppression strategies, the noise domain and the human voice domain are divided by determining the threshold and predicting the mask value of each time-frequency unit in the temporal mask. The mask value of each time-frequency unit in the noise domain can be set to 0 or processed by adjusting the attenuation coefficient. The mask value of each time-frequency unit in the human voice domain can be set to 1 or processed by adjusting the enhancement coefficient. The above processing methods for the mask values ​​of the noise domain and the human voice domain can be combined and applied based on actual needs. The combination of processing methods in the embodiments provided in this application is only exemplary.

[0055] To improve the accuracy of keyword triggering and avoid frequent fluctuations in noise reduction intensity due to mis-identification and incorrect strategy switching, which could negatively impact user experience, a possible approach is to employ a second noise suppression strategy to suppress noise in the speech stream when the keyword recognition result indicates the detection of a preset keyword. This strategy includes: If the keyword recognition result indicates that a preset keyword has been detected, and the number of times the preset keyword appears within a first preset time period is greater than a preset occurrence threshold, a second noise suppression strategy is adopted to suppress noise in the speech stream.

[0056] In other words, during voice interaction, the computer device can continuously monitor the occurrence frequency of preset keywords within a first preset duration. For example, a continuous scrolling time detection window of the first preset duration can be set to accumulate the occurrence frequency of preset keywords. If the occurrence frequency of preset keywords is greater than the preset occurrence frequency threshold within the first preset duration, it is determined to be a valid trigger. At this time, the second noise suppression strategy is used to suppress noise in the voice stream. If the occurrence frequency of preset keywords is less than or equal to the occurrence frequency threshold within the first preset duration, it is determined to be an occasional misidentification or non-continuous trigger. In this case, the first noise suppression strategy is still used to suppress noise in the voice stream. The first preset duration and the occurrence frequency threshold can both be set based on actual needs, and this application embodiment does not impose any restrictions on them.

[0057] To avoid a decrease in speech naturalness or an increase in power consumption due to maintaining high-intensity noise reduction for extended periods, as a possible implementation, after employing a second noise suppression strategy to suppress noise in the speech stream, the method further includes: If no preset keyword is detected within the second preset time period, the noise suppression strategy will be switched to the first noise suppression strategy.

[0058] The second preset duration refers to the duration after the noise suppression strategy has been switched to the second noise suppression strategy. If the preset keyword is not detected within the second preset duration, it indicates that the current voice interaction is stable, and the noise suppression strategy will be automatically switched from the second noise suppression strategy to the first noise suppression strategy to reduce the computing power consumption and power consumption of the terminal device. If the preset keyword is detected again within the second preset duration, the timer will be reset and the second noise suppression strategy will be used.

[0059] As one possible approach, in order to improve the accuracy of keyword recognition and avoid keyword omission or misrecognition due to environmental noise interference, after receiving the voice stream, the computer device can default to using the first noise suppression strategy to suppress noise and obtain a denoised voice stream. Then, keyword recognition is performed on the denoised voice stream to further determine the noise suppression strategy to be used subsequently.

[0060] The audio noise reduction method provided in this application involves signal processing, model inference, strategy judgment, and parameter adjustment, all of which are completed locally on the terminal device. Parameters, rules, and thresholds are all stored locally. It does not require network access or rely on cloud servers. It achieves local adaptive dynamic matching of noise reduction intensity according to the call status while protecting the security of voice interaction.

[0061] As one possible application scenario, during voice interaction, the computer device can apply corresponding noise suppression strategies to all global audio information during the voice interaction process based on keyword recognition results, including both uplink and downlink voice streams. As another possible application scenario, the computer device can also identify the recipient of the keyword. If no keyword is identified, a first noise suppression strategy is used to suppress noise in the global voice stream. If a keyword is identified, the computer device can determine whether to switch between noise suppression strategies for the uplink or downlink voice stream based on the recipient of the keyword. For example, if the computer device receives... If a preset keyword is detected in the received uplink voice stream, it indicates that the voice stream uploaded by the other end needs to undergo high-intensity noise suppression. The computer device can use a second noise suppression strategy to suppress the received downlink voice stream sent by the other end, and use a first noise suppression strategy to suppress the uplink voice stream uploaded by the local end. If the computer device detects a preset keyword in the received downlink voice stream, it indicates that the voice stream uploaded by the local end needs to undergo high-intensity noise suppression. The computer device can use a second noise suppression strategy to suppress the locally acquired uplink voice stream, and use a first noise suppression strategy to suppress the downlink voice stream sent by the other end.

[0062] The following uses a phone call scenario as an example to illustrate the application of the audio noise reduction method provided in this application embodiment; Figure 4This illustration shows a schematic diagram of an audio noise reduction process provided in an exemplary embodiment of this application. The method is executed by the local terminal device, such as... Figure 4 As shown, in this scenario, during a call, the local terminal device acquires the raw audio stream in real time, including combining the human voice input from the microphone with ambient sound to form a mixed audio signal as the uplink voice stream, and receiving the downlink voice stream transmitted from the other end. Initially, if no keywords are detected, the local device uses a first noise suppression strategy. If keywords such as "hello," "can't hear clearly," or "say it again" are detected in the downlink voice stream transmitted from the other end, and the number of occurrences exceeds a preset threshold, it is determined that the instantaneous signal-to-noise ratio of the current environment has deteriorated, and the basic noise reduction is insufficient to support a clear call. This indicates that the semantics are satisfied, triggering enhanced noise reduction. The device then switches to a second noise reduction strategy to deeply suppress non-human voice noise in either the local uplink or downlink voice stream, while also increasing the human voice volume gain to output a clear, noise-reduced voice stream. After the switch, timing and keyword recognition are performed. If no keywords are detected within a second preset time (e.g., 10 seconds), the device automatically switches back to the first noise suppression strategy. If keywords are still detected, the timing and keyword recognition are reset, and the second noise suppression strategy is maintained.

[0063] Figure 5 A structural block diagram of a computer device 500 illustrated in an exemplary embodiment of this application is shown. The computer device 500 can be implemented as the aforementioned terminal device, such as a smartphone, tablet computer, laptop computer, desktop computer, etc. The computer device 500 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0064] Typically, computer device 500 includes a processor 501 and a memory 502.

[0065] In some embodiments, the computer device 500 may also optionally include a peripheral device interface 503 and at least one peripheral device. The processor 501, memory 502, and peripheral device interface 503 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 503 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, and a power supply 508.

[0066] In some embodiments, the computer device 500 further includes one or more sensors 509. The one or more sensors 509 include, but are not limited to, an accelerometer 510, a gyroscope 511, a pressure sensor 512, an optical sensor 513, and a proximity sensor 514.

[0067] Those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the computer device 500, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0068] In one exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program, which is loaded and executed by a processor to implement all or part of the steps of the above embodiments.

[0069] In one exemplary embodiment, a computer program product is also provided, comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform all or part of the steps of the above embodiments.

[0070] It should be understood that the training and prediction processes of the AI ​​models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data source legitimacy: All datasets used for AI model training were obtained through legal means, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data originates from compliant data sources that comply with relevant open-source agreements, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology) have been implemented to remove personally identifiable information, fully complying with relevant laws and regulations.

[0071] Data content compliance: The AI ​​model's dataset undergoes multiple screenings and cleaning processes to remove all non-compliant information and ensures that there is no illegal acquisition or use of genetic resources. For data in sensitive fields (such as healthcare and finance), an additional privacy-preserving computation module (including federated learning and secure multi-party computation technologies) is used to ensure that the data is "usable but not visible," avoiding compliance risks during the original data transmission process and ensuring that the data application scenarios and uses comply with public order and good morals and industry regulatory requirements.

[0072] Data governance norms: A complete data traceability system is established during the AI ​​model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions and avoiding reliance on AI-generated data that has not undergone substantial human modification, thus meeting the examination requirements for "human main contributions" in AI patent applications.

[0073] Training objectives and plans are compliant: The AI ​​model training objective focuses on rewriting state selection. The training scheme and final output results do not violate any mandatory provisions of laws and administrative regulations, do not harm the public interest or the legitimate rights and interests of others, and do not pose any potential risks of being used for illegal activities, infringing on privacy, or disrupting public safety. It strictly adheres to the ethical principle of "intelligent for good".

[0074] Training process compliance: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision-making logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on expert system feedback to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with the ethical review requirements of relevant laws.

[0075] Training environment and tool compliance: AI model training is implemented using nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is built using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.

[0076] Training results ethical verification compliance: After the model is trained, it undergoes additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios, a dedicated result verification mechanism is established to ensure that the model always complies with relevant laws and regulations in practical applications.

[0077] In summary, the data and training process used in the AI ​​model in this specification strictly comply with relevant regulations and do not violate any laws, social ethics, public interests, or regulations on the use of genetic resources. Therefore, it fully meets the compliance requirements for patent authorization.

[0078] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0079] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. An audio noise reduction method, characterized in that, The method includes: Keyword recognition is performed on the speech stream to obtain the keyword recognition results; If the keyword recognition result indicates that no preset keyword has been detected, the first noise suppression strategy is used to suppress noise in the speech stream; When the keyword recognition result indicates that the preset keyword has been detected, a second noise suppression strategy is used to suppress noise in the speech stream; the noise reduction intensity of the first noise suppression strategy is lower than that of the second noise suppression strategy.

2. The method according to claim 1, characterized in that, The noise suppression of the speech stream includes: Perform a short-time Fourier transform on the speech stream to obtain the time-frequency feature information of the speech stream; The time-frequency feature information is input into the audio noise reduction model to obtain a predicted time-frequency mask, which is used to characterize the probability that each time-frequency unit in the time-frequency feature information belongs to human voice; Based on the noise suppression strategy corresponding to the keyword recognition result and the predicted time-frequency mask, the time-frequency feature information is processed to obtain the separated human voice audio. Output the human voice audio.

3. The method according to claim 2, characterized in that, The noise suppression strategy corresponding to the keyword recognition result and the predicted time-frequency mask are used to process the time-frequency feature information to obtain the separated human voice audio, including: When the first noise suppression strategy is adopted, the predicted time-frequency mask is determined and adjusted based on the first determination threshold to obtain the first time-frequency mask; the determination threshold is used to indicate the lowest probability value of the time-frequency unit belonging to human voice; Based on the first time-frequency mask, a mask weighting operation is performed on the time-frequency feature information to obtain the first human voice time-frequency feature information; The audio signal is reconstructed from the time-frequency feature information of the first human voice to obtain the separated first human voice audio.

4. The method according to claim 2, characterized in that, The noise suppression strategy corresponding to the keyword recognition result and the predicted time-frequency mask are used to process the time-frequency feature information to obtain the separated human voice audio, including: When the second noise suppression strategy is adopted, the predicted time-frequency mask is determined and adjusted based on the second determination threshold to obtain the second time-frequency mask; the second determination threshold is higher than the first determination threshold corresponding to the first noise suppression strategy. Based on the second time-frequency mask, a mask weighting operation is performed on the time-frequency feature information to obtain the second human voice time-frequency feature information; The audio signal is reconstructed from the time-frequency feature information of the second human voice to obtain the separated second human voice audio.

5. The method according to claim 2, characterized in that, The audio noise reduction model includes a first feature extraction branch, a second feature extraction branch, a feature fusion layer, and a mask prediction layer; The step of inputting the time-frequency feature information into the audio noise reduction model to obtain the predicted time-frequency mask includes: Frequency domain feature information is obtained by extracting frequency domain features from the time-frequency feature information through the first feature extraction branch; Temporal feature information is obtained by extracting temporal features from the frequency domain feature information through the second feature extraction branch; The frequency domain feature information and the temporal feature information are fused by the feature fusion layer to obtain spatiotemporal joint fused features. The spatiotemporal joint fusion features are mapped by the mask prediction layer to obtain the predicted time-frequency mask.

6. The method according to claim 1, characterized in that, When the keyword recognition result indicates that the preset keyword has been detected, a second noise suppression strategy is used to suppress noise in the speech stream, including: If the keyword recognition result indicates that the preset keyword has been detected, and the number of times the preset keyword appears within a first preset time period is greater than a preset occurrence threshold, the second noise suppression strategy is used to suppress noise in the speech stream.

7. The method according to claim 1 or 6, characterized in that, After suppressing noise in the speech stream using the second noise suppression strategy, the method further includes: If the preset keyword is not detected within the second preset time period, the noise suppression strategy will be switched to the first noise suppression strategy.

8. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to implement the audio noise reduction method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the audio noise reduction method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the audio noise reduction method as described in any one of claims 1 to 7.