Audio processing method, device and system, electronic equipment and storage medium
By setting up a dual noise reduction and amplification process and residual noise detection in both the front-end and back-end devices, and adaptively controlling the secondary noise reduction process, the problem of poor noise reduction effect of the front-end device in complex noise environment is solved, and high-quality audio output and resource optimization are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN MICROBT ELECTRONICS TECH CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
AI Technical Summary
Front-end devices are not effective at noise reduction in complex noise environments, leaving a lot of residual noise that affects audio quality. Existing technologies fail to fully utilize the processing capabilities of both front-end and back-end devices.
It employs a dual noise reduction and amplification process, combining primary and secondary noise reduction processing. The secondary noise reduction and amplification is controlled to be turned on and off by residual noise detection. It utilizes the processing capabilities of front-end and back-end devices and combines artificial intelligence and non-artificial intelligence noise reduction methods to adaptively cope with complex noise environments.
It improves audio quality in complex noisy environments, reduces noise residue, enhances speech clarity, optimizes resource utilization, adapts to environments with different signal-to-noise ratios, and ensures high-quality audio output.
Smart Images

Figure CN122090860A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio processing technology, and in particular to an audio processing method, apparatus, system, electronic device, and storage medium. Background Technology
[0002] Real-time noise reduction of audio signals has significant positive implications for information receivers. For example, it makes it easier for listeners to understand spoken content, thereby improving communication efficiency. Removing noise from audio helps protect hearing from damage. Audio noise reduction technology plays an important role in improving sound quality, protecting hearing, and improving communication efficiency, and is one of the indispensable technologies in the field of modern audio processing.
[0003] Due to the need for miniaturization in audio acquisition front-end devices, the effect of real-time noise reduction in audio signal acquisition may be less than ideal, especially in noisy environments such as outdoors or factories. In such cases, relying solely on noise reduction by the front-end device may negatively impact the listener's reception of the audio signal due to the poor noise reduction effect caused by the noisy environment. Summary of the Invention
[0004] In view of this, the present disclosure provides an audio processing method, apparatus, system, electronic device and storage medium to solve problems such as poor noise reduction effect, weak noise reduction strength and residual noise caused by the front-end device being in a noisy environment.
[0005] According to one aspect of the embodiments of this disclosure, an audio processing method is provided, comprising:
[0006] Acquire a first audio signal, wherein the first audio signal is the original audio signal after primary noise reduction and amplification;
[0007] Perform a first residual noise detection on the first audio signal to obtain a first detection result information associated with the first residual noise detection;
[0008] If the first detection result information meets the preset conditions for enabling secondary noise reduction and amplification, secondary noise reduction and amplification processing is performed to obtain the second audio signal;
[0009] Output the second audio signal.
[0010] In one possible implementation, after outputting the second audio signal, the audio processing method further includes:
[0011] Perform a second residual noise detection on the second audio signal to obtain a second detection result information associated with the second residual noise detection;
[0012] If the combined result information composed of the second detection result information and the first detection result information meets the preset secondary noise reduction amplification shutdown condition, the secondary noise reduction amplification process is turned off.
[0013] Output the first audio signal.
[0014] In one possible implementation, the step of performing a first residual noise detection on the first audio signal to obtain first detection result information associated with the first residual noise detection includes:
[0015] Based on the first audio signal, a first signal-to-noise ratio estimation information associated with the first audio signal is obtained;
[0016] Based on the first audio signal, a first speech quality index information associated with the first audio signal is obtained;
[0017] The first detection result information is obtained based on the first signal-to-noise ratio estimation information and the first speech quality index information.
[0018] In one possible implementation, the condition for enabling the secondary noise reduction amplification is:
[0019] The first detection result information reaches the preset secondary noise reduction and amplification activation trigger threshold information.
[0020] In one possible implementation, the step of performing a second residual noise detection on the second audio signal to obtain second detection result information associated with the second residual noise detection includes:
[0021] Based on the second audio signal, a second signal-to-noise ratio estimation information associated with the second audio signal is obtained;
[0022] Based on the second audio signal, a second speech quality index information associated with the second audio signal is obtained;
[0023] The second detection result information is obtained based on the second signal-to-noise ratio estimation information and the second speech quality index information.
[0024] In one possible implementation, the joint result information is the difference information between the second detection result information and the first detection result information;
[0025] The condition for turning off the secondary noise reduction amplification is that the combined result information is within the preset range of the secondary noise reduction amplification turn-off trigger threshold information.
[0026] In one possible implementation, the step of performing secondary noise reduction and amplification processing to obtain a second audio signal includes:
[0027] Two-stage noise reduction processing is performed to obtain a two-stage noise-reduced frequency signal;
[0028] The second audio signal is obtained by performing a second-level audio signal amplification process on the second-level noise-reduced frequency signal.
[0029] In one possible implementation, the secondary noise reduction process includes at least one of artificial intelligence noise reduction processing and non-artificial intelligence noise reduction processing (or conventional methods).
[0030] In one possible implementation, the AI noise reduction process includes:
[0031] Obtain the spectral data of the original audio signal;
[0032] Acquire the spectral data of the primary noise-reduced audio signal obtained after primary noise reduction of the original audio signal during the primary noise reduction amplification process;
[0033] The spectral data of the original audio signal and the spectral data of the primary noise-reduced audio signal are input into the audio noise reduction neural network model, and the audio spectral data associated with the second audio signal is obtained through the audio noise reduction neural network model.
[0034] The audio spectrum data is subjected to an inverse fast Fourier transform to obtain the secondary noise-reduced frequency signal.
[0035] In one possible implementation, the non-AI noise reduction process includes:
[0036] Perform a fast Fourier transform on the first audio signal to obtain the spectrum data of the first audio signal;
[0037] Eliminate the signal components in the frequency portion of the first audio signal that are associated with a specified audio component to obtain the noise-reduced spectrum data;
[0038] The second audio signal is obtained by performing an inverse fast Fourier transform on the noise-reduced spectral data.
[0039] In one possible implementation, the specified audio component includes at least one of wind noise, motor noise, and howling.
[0040] According to another aspect of the embodiments of this disclosure, an audio processing apparatus is provided, comprising:
[0041] The first audio signal acquisition module is configured to acquire a first audio signal, wherein the first audio signal is the original audio signal after primary noise reduction and amplification.
[0042] The detection module is configured to perform a first residual noise detection on the first audio signal and obtain a first detection result information associated with the first residual noise detection;
[0043] The secondary noise reduction and amplification module is configured to perform secondary noise reduction and amplification processing to obtain a second audio signal when the first detection result information meets the preset secondary noise reduction and amplification activation conditions.
[0044] The audio output module is configured to output the second audio signal.
[0045] According to another aspect of the embodiments of this disclosure, an audio processing system is provided, comprising:
[0046] A front-end audio acquisition and processing device is used to acquire raw audio signals and perform primary noise reduction and amplification on the raw audio signals to obtain a first audio signal;
[0047] The back-end audio processing device is used to receive the first audio signal, perform first residual noise detection on the first audio signal, obtain first detection result information associated with the first residual noise detection, and perform second-level noise reduction and amplification processing when the first detection result information meets the preset second-level noise reduction and amplification activation conditions to obtain and output the second audio signal.
[0048] According to another aspect of the embodiments of this disclosure, an electronic device is provided, comprising:
[0049] processor;
[0050] Memory for storing the executable instructions of the processor;
[0051] The processor is configured to execute the executable instructions to implement the audio processing method as described in any of the preceding claims.
[0052] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which, when at least one instruction in the computer-readable storage medium is executed by a processor of an electronic device, enables the electronic device to implement the audio processing method as described in any of the preceding claims.
[0053] As can be seen from the above scheme, the audio processing method, apparatus, system, electronic device, and storage medium disclosed herein overcome the shortcomings of related technologies by introducing a two-stage noise reduction and amplification processing mechanism that is turned on or off based on residual noise detection. This helps solve the technical problem of poor noise reduction effect in complex noise environments, improves noise reduction strength, reduces noise residue, and further eliminates noise. Specifically, primary noise reduction processing can be performed at the front-end device to eliminate most environmental noise, showing good results in high signal-to-noise ratio and near-field environments, and retaining more voice information. Secondary noise reduction processing can be performed at the back-end device, combining artificial intelligence noise reduction and non-artificial intelligence noise reduction methods. Artificial intelligence noise reduction can address nonlinear noise (such as wind noise, motor noise, howling, etc.) to improve sound quality. The amplification process uses automatic gain control. In high signal-to-noise ratio environments, there is less residual noise after noise reduction, and automatic gain control can simply amplify the audio signal. In low signal-to-noise ratio environments, such as far-field and outdoor environments, there is more residual noise after noise reduction, and the noise in the audio signal amplified by the automatic gain control of the front-end device is obvious. Secondary noise reduction processing can further improve sound quality.
[0054] Furthermore, the technical solution disclosed herein also facilitates resource integration. Specifically, the processing power of the front-end device (e.g., embedded device) can be used for primary noise reduction and amplification processing to reduce noise in the audio signal (first audio signal) transmitted to the back-end device. The relatively high processing power of the back-end device (e.g., mobile phone, cloud server) can be used to perform further noise reduction processing after receiving the audio signal, so that even with a low-performance front-end device, a higher quality audio can be output. Alternatively, for some low-end front-end devices (with limited processing power or possibly no noise processing capability), primary noise reduction processing can be performed on the back-end device to improve the quality of the audio signal. Attached Figure Description
[0055] Figure 1 This is a schematic flowchart illustrating an audio processing method according to an illustrative embodiment;
[0056] Figure 2 This is a system flow diagram illustrating an embodiment of an audio processing method;
[0057] Figure 3 This is a schematic flowchart illustrating the process of performing first residual noise detection according to an illustrative embodiment;
[0058] Figure 4 This is a system flow diagram for residual noise detection, illustrated according to an illustrative embodiment.
[0059] Figure 5 This is a schematic diagram illustrating the process of controlling the shutdown of the secondary noise reduction and amplification process according to an illustrative embodiment;
[0060] Figure 6 This is a schematic flowchart illustrating the second residual noise detection process according to an illustrative embodiment;
[0061] Figure 7 This is a schematic diagram illustrating a process for performing artificial intelligence noise reduction according to an illustrative embodiment;
[0062] Figure 8 This is a system flowchart illustrating artificial intelligence noise reduction processing according to an illustrative embodiment;
[0063] Figure 9 This is a schematic diagram illustrating a non-AI noise reduction process according to an illustrative embodiment;
[0064] Figure 10 This is a schematic diagram illustrating an application scenario of an audio processing method according to an illustrative embodiment;
[0065] Figure 11 This is a schematic diagram of an audio processing device structure according to an illustrative embodiment;
[0066] Figure 12 This is a schematic diagram of an audio processing system structure according to an illustrative embodiment;
[0067] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments.
[0069] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0070] Front-end devices for audio acquisition (such as embedded devices), like cameras, often have limited data processing capabilities due to miniaturization requirements. They typically only perform basic noise reduction on the acquired audio. Back-end devices, such as mobile phones, only receive and play the audio, while cloud servers receive and forward the audio. This approach fails to provide satisfactory audio quality in complex noisy environments (such as those with low or fluctuating signal-to-noise ratios), negatively impacting the user experience.
[0071] The existing technologies suffer from several problems, including insufficient single-stage noise reduction processing and inadequate resource utilization. Specifically, current audio acquisition and transmission systems typically perform simple noise reduction at the front-end equipment, which is insufficient to effectively handle complex noisy environments. In situations with low signal-to-noise ratios (SNR) or large SNR variations, residual noise is noticeable, affecting audio quality. Furthermore, in low SNR or far-field environments, the amplified audio signal still exhibits significant residual noise, leading to a poor user experience. Current audio acquisition and transmission systems fail to fully utilize the processing capabilities of both front-end and back-end equipment, particularly lacking further noise reduction processing at the back-end.
[0072] This disclosure provides an audio processing method, apparatus, system, electronic device, and storage medium. By setting up a dual noise reduction and amplification process in the front-end and back-end devices, and controlling the activation and deactivation of the secondary noise reduction and amplification based on the noise reduction effect, the noise reduction and amplification processing is adaptively performed according to the noise condition of the acquired audio signal. This achieves adaptive noise reduction and amplification for audio signals acquired in complex environments, which helps to provide higher quality audio output in complex environments and improves the speech clarity of audio signals acquired in various environments.
[0073] Figure 1 This is a schematic flowchart illustrating an audio processing method according to an exemplary embodiment. Figure 2 This is a system flow diagram illustrating an embodiment of the audio processing method, such as... Figure 1 and combined Figure 2 As shown, the audio processing method mainly includes the following steps 101 to 104.
[0074] Step 101: Obtain the first audio signal, which is the original audio signal after primary noise reduction and amplification.
[0075] Audio signal noise reduction mechanisms typically include noise reduction (NR) and amplification. Noise reduction usually employs a noise reduction module to reduce background noise during speech signal processing, thereby improving speech clarity and intelligibility. Amplification, also known as Automatic Gain Control (AGC), typically uses an AGC module to automatically adjust the audio signal gain to ensure the output signal remains within a stable level range. Audio signal noise reduction can be performed either before or after amplification, depending on the requirements.
[0076] In an illustrative embodiment, the primary noise reduction and amplification process includes a primary noise reduction process and a primary amplification process.
[0077] In an illustrative embodiment, the initial noise reduction process can use traditional noise reduction algorithms to perform preliminary noise reduction processing on the original audio signal, reducing most of the background noise. This method is effective for high signal-to-noise ratio and near-field environments and can retain more speech information.
[0078] In an illustrative embodiment, the primary noise reduction process may include: performing a Fast Fourier Transform (FFT) on the audio frames of the original audio signal to obtain the spectral data of the original audio signal; reducing or eliminating the signal components of the background noise frequency part in the spectral data of the original audio signal, while retaining the signal strength of the speech frequency part, to obtain the spectral data after primary noise reduction; and performing an Inverse Fast Fourier Transform (IFFT) on the spectral data after primary noise reduction to obtain the primary noise-reduced audio signal.
[0079] In an illustrative embodiment, the first audio signal may originate from a front-end device that performs audio acquisition, such as a camera. In this embodiment, the front-end device has a primary noise reduction and amplification function. The front-end device uses a built-in noise reduction and amplification module to perform primary noise reduction and amplification on the acquired raw audio signal to obtain the first audio signal. The first audio signal is then transmitted to a back-end device, such as a mobile phone, via wired and / or wireless communication.
[0080] Step 102: Perform a first residual noise detection on the first audio signal to obtain first detection result information related to the first residual noise detection.
[0081] In an illustrative embodiment, the first residual noise detection can be performed on the backend device.
[0082] Figure 3This is a schematic flowchart illustrating the process of performing first residual noise detection according to an illustrative embodiment, as shown below. Figure 3 As shown, in an illustrative embodiment, step 102 may include steps 301 to 303 as follows.
[0083] Step 301: Based on the first audio signal, obtain the first signal-to-noise ratio estimation information associated with the first audio signal;
[0084] Step 302: Based on the first audio signal, obtain the first speech quality index information associated with the first audio signal;
[0085] Step 303: Obtain the first detection result information based on the first signal-to-noise ratio estimation information and the first speech quality index information.
[0086] Figure 4 This is a system flow diagram for residual noise detection, illustrated according to an illustrative embodiment, as follows: Figure 4 As shown in the illustrative embodiment, the first signal-to-noise ratio (SNR) estimation information can be obtained based on the time-domain and frequency-domain characteristics of the first audio signal. These characteristics can be obtained through frame-by-frame analysis of the audio frames of the first audio signal; that is, the time-domain and frequency-domain characteristics are analyzed on a frame-by-frame basis. Frame-by-frame analysis of the audio frames enables the first detection result information to be real-time, allowing for timely tracking of changes in various audio components within the first audio signal. This real-time nature of the first detection result information facilitates timely response when the secondary noise reduction amplification is activated.
[0087] like Figure 4 As shown in the illustrative embodiment, the first signal-to-noise ratio (SNR) estimation information in step 301 is estimated based on the presence probabilities of speech and noise, i.e., the presence probability of speech and the presence probability of noise. Furthermore, the presence probabilities of speech and noise are obtained by combining the time-domain and frequency-domain features of the first audio signal. This method of combining the time-domain and frequency-domain features of the first audio signal to obtain the presence probabilities of speech and noise in the first audio signal helps improve the reliability of the presence probabilities of speech and noise, thereby improving the reliability of the first SNR estimation information. In the illustrative embodiment, obtaining the presence probabilities of speech and noise in the first audio signal based on the time-domain and frequency-domain features of the first audio signal can be achieved using a recursive averaging algorithm based on the presence probability of speech (e.g., the likelihood ratio method, the minimum value controlled recursive averaging algorithm, etc.), or by combining time-domain features such as short-time energy and short-time zero-crossing rate with frequency-domain features such as spectral entropy and cepstrum. For example, it can be achieved by combining the energy-entropy ratio of short-time energy in the time domain and spectral entropy in the frequency domain.
[0088] like Figure 4As shown in the illustrative embodiment, the first speech quality index information in step 302 is obtained by combining the time-domain and frequency-domain features of the first audio signal. In the illustrative embodiment, the first speech quality index information includes noise energy estimation information and intelligibility information of the first audio signal.
[0089] The noise energy estimation information can be obtained by combining the time-domain and frequency-domain characteristics of the first audio signal. In an illustrative embodiment, noise energy estimation based on time-domain characteristics can include time-domain energy detection methods, peak-to-peak ratio methods, pitch period analysis methods, and average ratio methods.
[0090] Among them, the time-domain energy detection method directly measures the energy of the audio signal in the time domain. It achieves this by calculating the energy (usually the sum of squares) of the audio signal within a certain time window.
[0091] The peak-to-peak ratio method involves the difference between the maximum and minimum values of an audio signal. This difference reflects the range of variation of the audio signal over a period of time. The peak-to-peak value is an important parameter describing the waveform characteristics of an audio signal. In time-domain analysis, the peak-to-peak value is usually used to measure the dynamic range of an audio signal, that is, the difference between the maximum and minimum values of the audio signal.
[0092] The pitch periodic component analysis method mainly focuses on the periodic components of audio signals in the time domain. It identifies and extracts periodic features, such as periodic impacts and fluctuations, by analyzing the time-domain waveform of the audio signal.
[0093] The mean comparison method focuses on the mean of an audio signal at different time points in the time domain. It analyzes the characteristics of an audio signal by comparing the average value of the signal in different time windows.
[0094] In an illustrative embodiment, noise energy estimation based on frequency domain characteristics may include methods such as spectral profile analysis, frequency change exponential analysis, and power spectral density analysis.
[0095] The spectral contour analysis method uses frequency domain information as the feature vector of the image and estimates noise energy based on the contour change characteristics of the spectrum.
[0096] The frequency variation index method analyzes the spectrum of an audio signal and evaluates its characteristics by calculating the distribution changes of the audio signal at different frequencies.
[0097] The frequency variation exponential method estimates noise energy based on the frequency variation exponential of noise. The amplitude spectrum integration method integrates the spectral amplitude of the audio signal to obtain the characteristic information of the audio signal in the frequency domain. The power spectral density analysis method estimates noise energy by analyzing the PSD (Power Spectral Density) of the audio signal. PSD describes the distribution of signal power at different frequencies, i.e., the power characteristics of the signal.
[0098] In an illustrative embodiment, the noise energy estimation information obtained based on frequency domain features and the noise energy estimation information obtained based on time domain features can be combined by weighting to obtain the final noise energy estimation information. For example, if the noise energy estimation information obtained based on frequency domain features is used as the main indicator of the final noise energy estimation information, and the noise energy estimation information obtained based on time domain features is used as the reference of the final noise energy estimation information, then the weight of the noise energy estimation information obtained based on frequency domain features can be set to a larger value, and the weight of the noise energy estimation information obtained based on time domain features can be set to a smaller value. The sum of the weights of the two is 1. The final noise energy estimation information is obtained by adding the results of multiplying the two by their respective weight values.
[0099] Indicators for measuring speech quality typically include intelligibility and clarity. Speech intelligibility refers to a listener's ability to accurately understand the content, which is closely related to the clear presentation of speech units in the speech signal. Highly intelligible speech typically exhibits clear formant trajectories, natural phoneme transitions, and stable modulation frequency characteristics (e.g., around 4Hz, related to syllable rate). Furthermore, short-term energy variations and the appropriateness of phoneme durations also reflect speech intelligibility. If these characteristics are blurred due to noise or distortion, the listener's understanding will be significantly more difficult. Speech clarity describes the purity of the speech signal, i.e., whether the signal is affected by noise or distortion. Clear speech often exhibits smooth spectral attenuation, narrow-band formants, and a clear harmonic structure, while distortion or noise blurs these characteristics, leading to decreased spectral flatness, abnormal zero-crossing rates, and even disrupting the dynamic range of the speech signal. Highly clear speech not only makes the content easy to discern but also provides a natural and realistic listening experience.
[0100] Speech intelligibility, also known as language clarity, refers to the percentage of speech signals that a listener can understand transmitted through a given sound transmission system. It is typically quantified by calculating correctly recognized words or phonemes. Intelligibility can be obtained using band-importance functions (BIF) and signal-to-noise ratio (SNR). The process may include, for example, dividing the audio signal into multiple frequency bands, calculating the SNR of each band, weighting the SNR of each band using the BIF to obtain a weighted intelligibility score for each band, and summing the weighted intelligibility scores of all bands to obtain an overall speech intelligibility score. In illustrative embodiments, intelligibility information can also be obtained using methods such as the Articulation Index (AI) or the Speech Reception Threshold (SRT). The clarity index is calculated based on the intensity of the speech signal relative to background noise in a specific frequency band. It can be calculated using frequency domain features. The speech reception threshold, on the other hand, is the threshold at which a listener's word recognition accuracy is 50% under a specific signal-to-noise ratio. This threshold can be calculated using time domain features. The process of obtaining intelligibility information using the clarity index can include, for example: dividing the audio signal into multiple frequency bands; calculating the difference between the signal strength and noise strength in each band; obtaining weighting coefficients corresponding to the center frequencies of each band from a clarity index weighting coefficient table; obtaining the clarity index value for each band based on the weighting coefficients, signal strength, and noise strength; summing the clarity index values of all bands to obtain the total clarity index; and obtaining the corresponding clarity by consulting a clarity-clarity index relationship graph or table based on the total clarity index. The clarity index is then used as the intelligibility score; a higher clarity index indicates higher intelligibility.The speech reception threshold method is a subjective evaluation method for speech intelligibility. The process of obtaining intelligibility information using the speech reception threshold can include, for example, pre-measuring the speech reception threshold. In a quiet environment, the speech reception threshold is defined as a representation level or intensity level, where the listener's word recognition accuracy is 50%. The speech reception threshold can be obtained by representing speech materials at different intensity levels from low to high and by intensity representation graphs. In a noisy environment, the speech reception threshold is defined as a signal-to-noise (S / N) level, where the listener's word recognition accuracy is also 50%. The speech reception threshold can be obtained by representing different S / N levels from negative S / N (e.g., -10dB) to positive S / N (e.g., 10dB). A negative S / N value indicates poor performance, while a positive S / N value indicates good performance. The S / N performance function is monotonically increasing in an S-shape. Based on this, a mapping relationship is established between the S / N value and intelligibility, allowing intelligibility to be obtained when the S / N value is obtained. In addition, in the illustrative embodiment, intelligibility information can also be obtained using a relevant speech intelligibility model through deep learning. In the illustrative embodiment, the intelligibility information obtained based on frequency domain features and the intelligibility information obtained based on time domain features can be combined using weights to obtain the final intelligibility information.
[0101] In an illustrative embodiment, in step 302, after obtaining the noise energy estimation information and intelligibility information of the first audio signal, the noise energy estimation information and intelligibility information of the first audio signal can be combined in a weighted manner to obtain the first speech quality index information.
[0102] In an illustrative embodiment, in step 303, the first signal-to-noise ratio estimation information and the first speech quality index information can be combined in a weighted manner to obtain the first detection result information.
[0103] Step 103: If the first detection result information meets the preset conditions for enabling secondary noise reduction and amplification, perform secondary noise reduction and amplification processing to obtain the second audio signal.
[0104] In an illustrative embodiment, the condition for enabling the secondary noise reduction amplification is: the first detection result information reaches the preset trigger threshold information for enabling the secondary noise reduction amplification.
[0105] The secondary noise reduction amplification activation trigger threshold information is used to characterize the high noise content in the first audio signal, indicating that the primary noise reduction amplification effect is not ideal. When the first detection result information reaches the preset secondary noise reduction amplification activation trigger threshold information, it indicates that due to the unsatisfactory effect of the primary noise reduction amplification, secondary noise reduction amplification processing needs to be activated to further amplify the first audio signal for noise reduction, thereby improving the noise reduction effect of the audio signal and improving the speech clarity of the audio signal.
[0106] Step 104: Output the second audio signal.
[0107] Generally, when the effect of primary noise reduction and amplification is not ideal, the speech clarity of the second audio signal obtained after secondary noise reduction and amplification is higher than that of the first audio signal. Therefore, after outputting the second audio signal, from the listener's perspective, the speech in the audio information is clearer and there is less noise, which helps to improve the listener's experience.
[0108] In practical applications, temporary changes in ambient sound can cause short-term fluctuations in the quality of the original audio signal. Examples include wind noise from gusts in outdoor environments and brief noise from passing vehicles. In such cases, the degradation of the first audio signal may trigger the preset conditions for activating secondary noise reduction amplification, thus activating the secondary noise reduction amplification process. After the temporary change in ambient sound, the first audio signal returns to its original quality, and the secondary noise reduction amplification remains active. However, sometimes the original signal quality of the first audio signal is good, resulting in minimal difference between the second and first audio signals. In this situation, from the listener's perspective, the second audio signal does not provide a better listening experience. Furthermore, activating the secondary noise reduction amplification consumes system resources, leading to unnecessary waste. Therefore, the secondary noise reduction amplification can be disabled without affecting the listener's experience, allowing the output of the first audio signal to proceed without compromising the secondary noise reduction amplification process.
[0109] In practical applications, there may be situations where the difference between the second audio signal and the first audio signal is small due to changes in the environment of the front-end device. For example, the front-end device may move from a noisy outdoor environment to a quieter indoor environment. In such cases, the second-level noise reduction and amplification process can be turned off without affecting the listener's experience, and the first audio signal can be output.
[0110] For the reasons mentioned above, the audio processing method of this disclosure further includes a step of determining whether to turn off the secondary noise reduction and amplification processing by combining the result information of the second residual noise detection of the second audio signal and the first detection result information.
[0111] Figure 5 This is a schematic diagram illustrating the process of controlling the shutdown of the secondary noise reduction amplification process according to an illustrative embodiment. Additionally, in... Figure 2 The document also illustrates the system flow framework related to controlling the shutdown of the secondary noise reduction and amplification process, such as... Figure 5 and combined Figure 2As shown in the illustrative embodiment, after step 104, the audio processing method of this disclosure embodiment may further include the following steps 501 to 503.
[0112] Step 501: Perform second residual noise detection on the second audio signal to obtain second detection result information related to the second residual noise detection.
[0113] In an illustrative embodiment, a second residual noise detection may also be performed at the back-end device.
[0114] Figure 6 This is a schematic flowchart illustrating the second residual noise detection process according to an illustrative embodiment, as shown below. Figure 6 As shown, in an illustrative embodiment, step 501 may include steps 601 to 603 as follows.
[0115] Step 601: Based on the second audio signal, obtain the second signal-to-noise ratio estimation information associated with the second audio signal;
[0116] Step 602: Obtain second speech quality index information associated with the second audio signal based on the second audio signal;
[0117] Step 603: Obtain the second detection result information based on the second signal-to-noise ratio estimation information and the second speech quality index information.
[0118] The process for detecting the second residual noise described above is similar to the process for detecting the first residual noise described above, and can also be adopted. Figure 4 The system flow framework for residual noise detection is shown. In the illustrative embodiment, the second signal-to-noise ratio (SNR) estimation information can be obtained based on the time-domain and frequency-domain characteristics of the second audio signal. These characteristics can be obtained through frame-by-frame analysis of the audio frames of the second audio signal. Frame-by-frame analysis of the audio frames enables the second detection result information to be real-time, allowing for timely tracking of changes in various audio components within the second audio signal. This real-time nature of the second detection result information facilitates a timely response to the shutdown of the secondary noise reduction amplification.
[0119] Similar to the first signal-to-noise ratio (SNR) estimation information, the second SNR estimation information in step 601 is also estimated based on the probability of speech and noise presence, i.e., the probability of speech presence and the probability of noise presence. Furthermore, the probability of speech presence and the probability of noise presence are obtained by combining the time-domain and frequency-domain features of the second audio signal, which helps to improve the reliability of the second SNR estimation information.
[0120] Similar to the first speech quality metric, the second speech quality metric in step 602 is obtained by combining the time-domain and frequency-domain features of the second audio signal. In an illustrative embodiment, the second speech quality metric includes noise energy estimation information and intelligibility information of the second audio signal.
[0121] Similar to the first speech quality index information, in step 602, after obtaining the noise energy estimation information and intelligibility information of the second audio signal, the noise energy estimation information and intelligibility information of the second audio signal can be combined in a weighted manner to obtain the second speech quality index information.
[0122] Similar to the first detection result information, in step 603, the second signal-to-noise ratio estimation information and the second speech quality index information can be combined in a weighted manner to obtain the second detection result information.
[0123] Step 502: If the combined result information composed of the second detection result information and the first detection result information meets the preset conditions for turning off the second-level noise reduction amplification, then turn off the second-level noise reduction amplification process.
[0124] In an illustrative embodiment, the joint result information is the difference information between the second detection result information and the first detection result information. Based on this, the condition for turning off the secondary noise reduction amplification is: the joint result information is within the preset range of the secondary noise reduction amplification turn-off trigger threshold information.
[0125] The joint result information is used to characterize the perceived difference between the first and second audio signals. If the perceived difference is small, it indicates that the improvement in the effect of the secondary noise reduction amplification process is minimal, and the secondary noise reduction amplification process can be turned off. This perceived difference can be quantified using a preset trigger threshold range for turning off the secondary noise reduction amplification. If the joint result information falls within the preset trigger threshold range, it indicates that the difference between the second and first audio signals is not significant, and the result of the secondary noise reduction amplification is not substantial. The trigger threshold range for turning off the secondary noise reduction amplification can be adaptively set and adjusted according to the actual application scenario.
[0126] Step 503: Output the first audio signal.
[0127] When the combined result information is within the preset threshold information for turning off the secondary noise reduction amplification, the difference between the second audio signal and the first audio signal is not significant, and the result of the secondary noise reduction amplification is not significant. At this time, turning off the secondary noise reduction amplification process and changing the output of the second audio signal to the output of the first audio signal will not have a significant impact on the listener's listening experience.
[0128] To avoid potential discomfort to the listener when switching between the output of the first audio signal and the second audio signal, in the illustrative embodiment, speech smoothing processing may also be performed on the output audio signal when switching from the output of the first audio signal to the output of the second audio signal, or from the output of the second audio signal to the output of the first audio signal.
[0129] In an illustrative embodiment, the secondary noise reduction and amplification process can be performed on a backend device with relatively strong data processing capabilities. In an illustrative embodiment, step 103 may specifically include: performing secondary noise reduction processing to obtain a secondary noise-reduced frequency signal; and performing secondary audio signal amplification processing on the secondary noise-reduced frequency signal to obtain a second audio signal.
[0130] In the illustrative embodiment, different secondary noise reduction processing methods can be adopted according to the data processing capabilities of the backend device. Based on this, in the illustrative embodiment, the secondary noise reduction processing may include at least one of artificial intelligence noise reduction processing and non-artificial intelligence noise reduction processing. Non-artificial intelligence noise reduction processing includes other noise reduction processing methods in related technologies besides artificial intelligence noise reduction processing, such as noise reduction processing performed using existing traditional noise reduction models for specific noise.
[0131] Figure 7 This is a schematic diagram illustrating the process of performing artificial intelligence noise reduction according to an illustrative embodiment. Figure 8 This is a system flowchart illustrating artificial intelligence noise reduction processing according to an illustrative embodiment, such as... Figure 7 , Figure 8 and combined Figure 2 As shown in the illustrative embodiment, the artificial intelligence noise reduction process mainly includes the following steps 701 to 704.
[0132] Step 701: Obtain the spectrum data of the original audio signal.
[0133] like Figure 8 and combined Figure 2 As shown, in the illustrative embodiment, the spectral data of the original audio signal is obtained from the front-end device. This spectral data can be obtained from the primary noise reduction process within the primary noise reduction and amplification process. For example, in the primary noise reduction and amplification process, the original audio signal undergoes a Fast Fourier Transform (FFT) to obtain its spectral data. In the illustrative embodiment, in step 701, the back-end device can directly obtain the spectral data of the original audio signal from the front-end device. In the illustrative embodiment, noise reduction is performed based on audio frames; therefore, a frame segmentation operation of the original audio signal is included before the FFT.
[0134] In an illustrative embodiment, the spectral data of the original audio signal includes characteristic information about the noise intensity in the original audio signal.
[0135] Step 702: Obtain the spectrum data of the primary noise-reduced audio signal obtained after primary noise reduction of the original audio signal during the primary noise reduction amplification process.
[0136] In an illustrative embodiment, the spectral data of the primary noise-reduced audio signal contains characteristic information of the noise intensity in the audio signal after primary noise reduction, that is, it contains characteristic information of the noise intensity in the first audio signal.
[0137] Step 703: Input the spectrum data of the original audio signal and the spectrum data of the primary noise-reduced audio signal into the audio noise reduction neural network model, and obtain the audio spectrum data associated with the second audio signal through the audio noise reduction neural network model.
[0138] In an illustrative embodiment, the audio noise reduction neural network model further eliminates noise components in the audio signal based on the noise intensity characteristics of the original audio signal and the first audio signal. For example, if the noise intensity in the first audio signal is significantly reduced compared to the noise intensity in the original audio signal, it indicates that the primary noise reduction effect is good. In this case, the audio noise reduction neural network model continues to eliminate a small amount of noise components based on the spectral data of the primary noise-reduced audio signal. If the noise intensity in the first audio signal is only slightly reduced compared to the noise intensity in the original audio signal, it indicates that the primary noise reduction effect is poor. In this case, the audio noise reduction neural network model continues to eliminate a large amount of noise components based on the spectral data of the primary noise-reduced audio signal.
[0139] By utilizing an audio noise reduction neural network model and selecting a suitable training sample set, noise reduction processing can be achieved for various specific noise components. However, considering the limitations of backend device computing power and the real-time requirements of noise reduction, the artificial intelligence noise reduction process preferably adopts a lightweight audio noise reduction neural network model.
[0140] Step 704: Perform an inverse fast Fourier transform on the audio spectrum data to obtain a second-level noise-reduced frequency signal.
[0141] In an illustrative embodiment, non-AI noise reduction processing may employ noise reduction processing performed by a conventional noise reduction model for a specific type of noise. Figure 9 This is a schematic diagram illustrating a non-AI noise reduction process according to an illustrative embodiment, such as... Figure 9 As shown, non-AI noise reduction processing may include the following steps 901 to 903.
[0142] Step 901: Perform a fast Fourier transform on the first audio signal to obtain the spectrum data of the first audio signal;
[0143] Step 902: Reduce the signal strength of the frequency portion of the first audio signal that is associated with the specified audio component in the spectral data to obtain the noise-reduced spectral data;
[0144] Step 903: Perform an inverse fast Fourier transform on the denoised spectral data to obtain the second audio signal.
[0145] Step 902 can be implemented using a denoising model for specific noise.
[0146] In the illustrative embodiment, the specified audio components include at least one of wind noise, motor noise, and howling. Specifically, wind noise can be denoised using a wind noise denoising model, motor noise can be denoised using a motor noise denoising model, and howling can be denoised using a howling denoising model.
[0147] The audio processing method of this disclosure addresses the shortcomings of related technologies, such as insufficient single noise reduction processing, lack of multi-level processing mechanisms, and inadequate resource utilization. It overcomes these shortcomings by introducing a two-stage noise reduction amplification processing mechanism that is activated and deactivated based on residual noise detection. This helps solve the technical problem of poor noise reduction performance in complex noisy environments, improves noise reduction intensity, reduces noise residue, and further eliminates noise. Specifically, a primary noise reduction process can be performed at the front-end device using a noise reduction algorithm to eliminate most environmental noise, showing good performance in high signal-to-noise ratio and near-field environments while preserving more voice information. A secondary noise reduction process can be performed at the back-end device, combining artificial intelligence (AI) noise reduction with non-AI noise reduction methods. AI noise reduction can process non-linear noise (such as wind noise, motor noise, and howling) to improve sound quality. The amplification process employs automatic gain control. In high signal-to-noise ratio (SNR) environments, the residual noise after noise reduction is minimal, and automatic gain control can simply amplify the audio signal. However, in low SNR environments, such as far-field or outdoor environments, the residual noise after noise reduction is significant, and the noise in the amplified audio signal from the front-end device's automatic gain control is noticeable. Secondary noise reduction processing can further improve sound quality. This disclosure also facilitates resource integration. The processing power of the front-end device can be utilized for primary noise reduction and amplification, reducing noise in the audio signal (first audio signal) transmitted to the back-end device. The relatively higher processing power of the back-end device can be utilized for further noise reduction processing after receiving the audio signal, enabling high-quality audio output even with low-performance front-end devices. Alternatively, for some low-end front-end devices (with limited processing power or potentially no noise processing capabilities), primary noise reduction can be performed at the back-end device to improve the audio signal quality.
[0148] The audio processing method of this disclosure helps improve the audio quality acquired in complex noisy environments. Through primary noise reduction and amplification processing and secondary noise reduction and amplification processing performed on the front-end device, background noise can be effectively reduced, helping to enhance the clarity of the speech signal. In particular, by combining artificial intelligence noise reduction technology, it can handle nonlinear noise that is difficult to eliminate with traditional noise reduction algorithms, thus improving the applicability of the system and the user experience.
[0149] To address potential latency issues during the joint primary and secondary noise reduction and amplification processes of front-end and back-end devices, more efficient algorithms and optimized computing architectures (such as chip acceleration and GPU acceleration on the back-end device side) can be used to reduce latency. In addition, data compression techniques can be employed during processing to further reduce the burden of transmission and processing.
[0150] When employing artificial intelligence (AI) noise reduction, especially using small neural networks for secondary noise reduction, significant computational resources may be required. This places high demands on the computing power of backend devices. To address this, the size and complexity of the AI model can be optimized, and techniques such as model pruning and quantization can be used to reduce computational resource consumption. Alternatively, cloud computing resources can be utilized, leveraging cloud devices to perform the computationally intensive secondary noise reduction tasks. The backend devices communicate with the cloud devices to send and receive audio signals before and after the secondary noise reduction process.
[0151] To address the issue of existing denoising algorithms and AI models' limited adaptability to certain types of noise (such as wind noise under extreme weather conditions), robustness can be improved by enhancing the algorithm's adaptability and training the model to accommodate a wider range of noise types. Adaptive denoising techniques can also be incorporated, dynamically adjusting relevant denoising parameters based on changes in environmental noise. In certain situations, specialized denoising processing can also be applied to specific types of noise.
[0152] Achieving primary and secondary noise reduction amplification may require compatibility support between different devices, particularly in data transmission and processing between front-end and back-end devices. Therefore, in this illustrative embodiment, standardized data interfaces and protocols are used between the front-end and back-end devices to ensure data compatibility and transmission efficiency. By designing a more flexible system architecture to adapt to the capabilities and requirements of different devices, the front-end and back-end devices can be applied to a wider range of devices, including action cameras, drones, and robots.
[0153] To address the issue that users may need to manually adjust noise reduction and gain settings based on environmental conditions, increasing the complexity of use, the system can automatically optimize settings by recognizing the environment and learning user habits, thus achieving intelligent automatic adjustment. Furthermore, by providing an intuitive graphical user interface and related human-computer interaction functions, users can more easily make manual adjustments and view the processing effects.
[0154] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0155] Figure 10 This is a schematic diagram illustrating an application scenario of an audio processing method according to an illustrative embodiment, such as... Figure 10 As shown, the application scenario process mainly includes the following steps 1001 to 1009.
[0156] Step 1001: Obtain the first audio signal obtained after the original audio signal has undergone primary noise reduction and amplification, and then proceed to step 1002.
[0157] Step 1002: Perform first residual noise detection on the first audio signal to obtain first detection result information associated with the first residual noise detection, and then execute step 1003.
[0158] Step 1003: Determine whether the first detection result information meets the preset conditions for enabling secondary noise reduction and amplification. If yes, proceed to step 1004; otherwise, proceed to step 1005.
[0159] Step 1004: Perform two-stage noise reduction and amplification processing on the first audio signal to obtain the second audio signal, and then execute step 1006.
[0160] Step 1005: Output the acquired first audio signal and execute step 1001.
[0161] Step 1006: Perform second residual noise detection on the second audio signal to obtain second detection result information associated with the second residual noise detection, and then execute step 1007.
[0162] Step 1007: Determine whether the joint result information composed of the second detection result information and the first detection result information meets the preset second-level noise reduction and amplification shutdown condition. If yes, proceed to step 1008; otherwise, proceed to step 1009.
[0163] Step 1008: Stop the secondary noise reduction and amplification process on the first audio signal and proceed to step 1005.
[0164] Step 1009: Output the second audio signal and execute step 1001.
[0165] The specific execution process of each of the above steps can be found in the descriptions of the above embodiments, and will not be repeated here.
[0166] Figure 11 This is a schematic diagram of an audio processing device structure according to an illustrative embodiment, such as... Figure 11 As shown, the audio processing device mainly includes a first audio signal acquisition module 1101, a detection module 1102, a secondary noise reduction and amplification module 1103, and an audio output module 1104. The first audio signal acquisition module 1101 is configured to acquire a first audio signal, which is the original audio signal after primary noise reduction and amplification. The detection module 1102 is configured to perform a first residual noise detection on the first audio signal to obtain first detection result information associated with the first residual noise detection. The secondary noise reduction and amplification module 1103 is configured to perform secondary noise reduction and amplification processing when the first detection result information meets preset secondary noise reduction and amplification activation conditions to obtain a second audio signal. The audio output module 1104 is configured to output the second audio signal.
[0167] In an illustrative embodiment, the detection module 1102 is further configured to perform: second residual noise detection on the second audio signal to obtain second detection result information associated with the second residual noise detection; and the secondary noise reduction amplification module 1103 is further configured to perform: when the joint result information composed of the second detection result information and the first detection result information satisfies the preset secondary noise reduction amplification shutdown condition, the secondary noise reduction amplification process is turned off; and the audio output module 1104 is further configured to perform: output the first audio signal.
[0168] In an illustrative embodiment, the detection module 1102 is further configured to perform: obtaining first signal-to-noise ratio estimation information associated with the first audio signal based on the first audio signal; obtaining first speech quality index information associated with the first audio signal based on the first audio signal; and obtaining first detection result information based on the first signal-to-noise ratio estimation information and the first speech quality index information.
[0169] In an illustrative embodiment, the condition for enabling the secondary noise reduction amplification is: the first detection result information reaches the preset trigger threshold information for enabling the secondary noise reduction amplification.
[0170] In an illustrative embodiment, the detection module 1102 is further configured to perform: obtaining second signal-to-noise ratio estimation information associated with the second audio signal based on the second audio signal; obtaining second speech quality index information associated with the second audio signal based on the second audio signal; and obtaining second detection result information based on the second signal-to-noise ratio estimation information and the second speech quality index information.
[0171] In the illustrative embodiment, the joint result information is the difference information between the second detection result information and the first detection result information; the secondary noise reduction amplification shutdown condition is: the joint result information is within the preset secondary noise reduction amplification shutdown trigger threshold information range.
[0172] In an illustrative embodiment, the secondary noise reduction and amplification module 1103 is further configured to perform: secondary noise reduction processing to obtain a secondary noise-reduced frequency signal; and secondary audio signal amplification processing on the secondary noise-reduced frequency signal to obtain a second audio signal.
[0173] In an illustrative embodiment, the secondary noise reduction process includes at least one of artificial intelligence noise reduction processing and non-artificial intelligence noise reduction processing.
[0174] In an illustrative embodiment, the secondary noise reduction and amplification module 1103 is further configured to perform the following: acquiring the spectrum data of the original audio signal; acquiring the spectrum data of the primary noise-reduced frequency signal obtained after primary noise reduction of the original audio signal during the primary noise reduction and amplification process; inputting the spectrum data of the original audio signal and the spectrum data of the primary noise-reduced frequency signal into an audio noise reduction neural network model, and obtaining audio spectrum data associated with the second audio signal through the audio noise reduction neural network model; and performing an inverse fast Fourier transform on the audio spectrum data to obtain the secondary noise-reduced frequency signal.
[0175] In an illustrative embodiment, the secondary noise reduction and amplification module 1103 is further configured to perform: performing a fast Fourier transform on the first audio signal to obtain the spectrum data of the first audio signal; reducing the signal strength of the frequency portion of the spectrum data of the first audio signal associated with a specified audio component to obtain the noise-reduced spectrum data; and performing an inverse fast Fourier transform on the noise-reduced spectrum data to obtain the second audio signal.
[0176] In an illustrative embodiment, the specified audio components include at least one of wind noise, motor noise, and howling.
[0177] Regarding the audio processing apparatus in the above embodiments, the specific manner in which each unit performs its operation has been described in detail in the embodiments relating to the audio processing method, and will not be elaborated upon here.
[0178] It should be noted that the above embodiments are only examples of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0179] Figure 12 This is a schematic diagram of an audio processing system structure according to an illustrative embodiment, such as... Figure 12As shown, the audio processing system includes a front-end audio acquisition and processing device 1201 and a back-end audio processing device 1202. The front-end audio acquisition and processing device 1201 acquires the original audio signal and performs primary noise reduction and amplification on the original audio signal to obtain a first audio signal. The back-end audio processing device 1202 receives the first audio signal, performs primary residual noise detection on the first audio signal, obtains primary detection result information related to the primary residual noise detection, and performs secondary noise reduction and amplification processing when the primary detection result information meets preset conditions for enabling secondary noise reduction and amplification, thereby obtaining and outputting a second audio signal.
[0180] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. In some embodiments, the electronic device is a server. The electronic device 1300 can vary considerably due to different configurations or performance, and may include one or more Central Processing Units (CPUs) 1301 and one or more memories 1302, wherein the memory 1302 stores at least one line of program code, which is loaded and executed by the processor 1301 to implement the audio processing methods provided in the various embodiments described above. Of course, the electronic device 1300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The electronic device 1300 may also include other components for implementing device functions, which will not be elaborated here.
[0181] In an exemplary embodiment, a computer-readable storage medium including at least one instruction is also provided, such as a memory including at least one instruction, which can be executed by a processor in a computer device to perform the audio processing method in the above embodiments.
[0182] Optionally, the aforementioned computer-readable storage medium may be a non-transitory computer-readable storage medium, such as ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices.
[0183] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An audio processing method, comprising: Acquire a first audio signal, wherein the first audio signal is the original audio signal after primary noise reduction and amplification; Perform a first residual noise detection on the first audio signal to obtain a first detection result information associated with the first residual noise detection; If the first detection result information meets the preset conditions for enabling secondary noise reduction and amplification, secondary noise reduction and amplification processing is performed to obtain the second audio signal; Output the second audio signal.
2. The audio processing method according to claim 1, characterized in that, After outputting the second audio signal, the audio processing method further includes: Perform a second residual noise detection on the second audio signal to obtain a second detection result information associated with the second residual noise detection; If the combined result information composed of the second detection result information and the first detection result information meets the preset secondary noise reduction amplification shutdown condition, the secondary noise reduction amplification process is turned off. Output the first audio signal.
3. The audio processing method according to claim 1, characterized in that, The step of performing a first residual noise detection on the first audio signal to obtain first detection result information associated with the first residual noise detection includes: Based on the first audio signal, a first signal-to-noise ratio estimation information associated with the first audio signal is obtained; Based on the first audio signal, a first speech quality index information associated with the first audio signal is obtained; The first detection result information is obtained based on the first signal-to-noise ratio estimation information and the first speech quality index information.
4. The audio processing method according to claim 1, characterized in that, The conditions for enabling the secondary noise reduction amplification are as follows: The first detection result information reaches the preset secondary noise reduction and amplification activation trigger threshold information.
5. The audio processing method according to claim 2, characterized in that, The step of performing a second residual noise detection on the second audio signal to obtain second detection result information associated with the second residual noise detection includes: Based on the second audio signal, a second signal-to-noise ratio estimation information associated with the second audio signal is obtained; Based on the second audio signal, a second speech quality index information associated with the second audio signal is obtained; The second detection result information is obtained based on the second signal-to-noise ratio estimation information and the second speech quality index information.
6. The audio processing method according to claim 2, characterized in that: The combined result information is the difference information between the second detection result information and the first detection result information; The condition for turning off the secondary noise reduction amplification is that the combined result information is within the preset range of the secondary noise reduction amplification turn-off trigger threshold information.
7. The audio processing method according to claim 1, characterized in that, The second audio signal is obtained by performing secondary noise reduction and amplification processing, including: Two-stage noise reduction processing is performed to obtain a two-stage noise-reduced frequency signal; The second audio signal is obtained by performing a second-level audio signal amplification process on the second-level noise-reduced frequency signal.
8. The audio processing method according to claim 7, characterized in that: The secondary noise reduction process includes at least one of artificial intelligence noise reduction processing and non-artificial intelligence noise reduction processing.
9. The audio processing method according to claim 8, characterized in that, The artificial intelligence noise reduction process includes: Obtain the spectral data of the original audio signal; Acquire the spectral data of the primary noise-reduced audio signal obtained after primary noise reduction of the original audio signal during the primary noise reduction amplification process; The spectral data of the original audio signal and the spectral data of the primary noise-reduced audio signal are input into the audio noise reduction neural network model, and the audio spectral data associated with the second audio signal is obtained through the audio noise reduction neural network model. The audio spectrum data is subjected to an inverse fast Fourier transform to obtain the secondary noise-reduced frequency signal.
10. The audio processing method according to claim 8, characterized in that, The non-AI noise reduction processing includes: Perform a fast Fourier transform on the first audio signal to obtain the spectrum data of the first audio signal; Eliminate the signal components in the frequency portion of the first audio signal that are associated with a specified audio component to obtain the noise-reduced spectrum data; The second audio signal is obtained by performing an inverse fast Fourier transform on the noise-reduced spectral data.
11. The audio processing method according to claim 10, characterized in that: The specified audio components include at least one of wind noise, motor noise, and howling.
12. An audio processing device, characterized in that, include: The first audio signal acquisition module is configured to acquire a first audio signal, wherein the first audio signal is the original audio signal after primary noise reduction and amplification. The detection module is configured to perform a first residual noise detection on the first audio signal and obtain a first detection result information associated with the first residual noise detection; The secondary noise reduction and amplification module is configured to perform secondary noise reduction and amplification processing to obtain a second audio signal when the first detection result information meets the preset secondary noise reduction and amplification activation conditions. The audio output module is configured to output the second audio signal.
13. An audio processing system, characterized in that, include: A front-end audio acquisition and processing device is used to acquire raw audio signals and perform primary noise reduction and amplification on the raw audio signals to obtain a first audio signal; The back-end audio processing device is used to receive the first audio signal, perform first residual noise detection on the first audio signal, obtain first detection result information associated with the first residual noise detection, and perform second-level noise reduction and amplification processing when the first detection result information meets the preset second-level noise reduction and amplification activation conditions to obtain and output the second audio signal.
14. An electronic device, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the executable instructions to implement the audio processing method as described in any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that, When at least one instruction in the computer-readable storage medium is executed by the processor of the electronic device, the electronic device is able to implement the audio processing method as described in any one of claims 1 to 11.