Communication noise reduction method and system for explosion-proof industrial telephone
The audio signal is collected through the main and auxiliary dual microphones, wavelet transformation and instantaneous mutual correlation coefficient analysis, and combined with adaptive algorithm processing, the problem of unstable call quality in explosion-proof industrial telephones in a noisy environment is solved, and efficient communication noise reduction effect is achieved.
Patent Information
- Application Number
- CN202510664471.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
AI Technical Summary
In an industrial production environment, explosion-proof industrial telephones have unstable call quality due to mechanical noise, electromagnetic interference, etc., and the existing adaptive filtering algorithms are difficult to quickly adjust when the position of the noise source changes rapidly, resulting in noise leakage and poor call continuity.
Audio signals are collected through the main and auxiliary dual microphones, wavelet transformation is performed to obtain multi-scale time-frequency characteristics, calculate the instantaneous mutual correlation coefficient to construct the time-frequency correlation map, identify mutation frames, and use fast convergence or conventional step size adaptive algorithms to perform noise reduction according to the band energy distribution to ensure the continuity and stability of the speech signal.
It realizes the rapid suppression of noise while maintaining the mutation characteristics of voice signals, reducing voice distortion, improving the continuity and stability of calls, and enhancing the communication noise reduction ability of explosion-proof industrial telephones.
Smart Images

Figure CN120544591A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication noise reduction, and in particular to a communication noise reduction method and system for explosion-proof industrial telephones. Background Art
[0002] In industrial production environments, explosion-proof industrial telephones are important communication equipment. However, the mechanical noise, electromagnetic interference, and environmental noise generated by the operation of various equipment that are prevalent in industrial sites seriously affect the quality of calls, making it difficult to understand the content of calls. Especially in high-noise environments, it is easy to cause information transmission errors, affecting production safety and work efficiency.
[0003] In related technologies, an industrial telephone noise reduction method based on an adaptive filtering algorithm can be adopted. This method uses a built-in dual-microphone array in the telephone, uses the main microphone to collect noise signals with target speech, and the auxiliary microphone to collect environmental reference noise. The collected signals are processed in real time in combination with an adaptive minimum mean square error (LMS) algorithm, thereby achieving dynamic elimination of industrial environmental noise.
[0004] However, when the location of noise sources in industrial sites changes rapidly, the adaptive filtering algorithm converges slowly, making it difficult to adjust the noise reduction parameters in a timely manner. This results in brief noise leakage at the moment when the noise source location changes suddenly, reducing the continuity and stability of call quality. Summary of the Invention
[0005] The present application provides a communication noise reduction method and system for an explosion-proof industrial telephone, which are used to maintain the continuity and stability of calls and improve the communication noise reduction capability of the explosion-proof industrial telephone.
[0006] In a first aspect, the present application provides a communication noise reduction method for an explosion-proof industrial telephone, wherein a first audio signal containing a target voice is collected by a main microphone, and a second audio signal containing ambient noise is collected by an auxiliary microphone; Calculating the instantaneous cross-correlation coefficient of the first audio signal and the second audio signal at each scale at a preset reference time interval, and constructing a time-frequency correlation graph based on the instantaneous cross-correlation coefficient; Calculate the change rate of the correlation coefficient between adjacent time points in the time-frequency correlation map. When the change rate is greater than a preset threshold, mark the signal frame corresponding to the time point where the change rate is greater than the preset threshold as a mutation frame. Calculate the energy value of the mutation frame in each frequency band, and use the average value of the energy value as the energy reference value, wherein the frequency band with an energy value greater than the energy reference value is determined as a high-energy frequency band, and the frequency band with an energy value less than or equal to the energy reference value is determined as a low-energy frequency band; A first preset step-size fast-converging adaptive algorithm is used for the high-energy frequency band, and a second preset step-size fast-converging adaptive algorithm is used for the low-energy frequency band to perform noise reduction processing, thereby obtaining a noise reduction signal corresponding to the sudden change frame; The noise reduction signal is subjected to inverse wavelet transform to obtain a time domain signal, which is then converted into an audio signal for output.
[0007] By adopting the above technical solution, audio signals are collected through dual primary and secondary microphones, wavelet transforms are performed on the signals to obtain multi-scale time-frequency features, and time-frequency correlation maps are constructed by calculating the instantaneous cross-correlation coefficient. This accurately captures the changes in the time-frequency features of the audio signal. Mutation frames are identified based on the rate of change of the cross-correlation coefficients between adjacent time points. Combined with the energy distribution characteristics of the mutation frames in each frequency band, a fast-converging adaptive algorithm is used for high-energy bands, and a regular-step-size adaptive algorithm is used for low-energy bands to perform noise reduction. This allows the system to suppress noise while maintaining the mutation characteristics of the speech signal. By using a fast-converging adaptive algorithm for high-energy bands, the changing characteristics of the speech signal can be quickly tracked, reducing speech distortion. By using a regular-step-size adaptive algorithm for low-energy bands, background noise can be smoothly suppressed, avoiding audio distortion, minimizing damage to the speech signal, maintaining call continuity and stability, and improving the noise reduction capabilities of explosion-proof industrial telephones.
[0008] In conjunction with some embodiments of the first aspect, in some embodiments, calculating the instantaneous cross-correlation coefficient of the first audio signal and the second audio signal at each scale at a preset reference time interval specifically includes: Performing sliding analysis on the first audio signal and the second audio signal according to a preset reference time interval to obtain a signal segment within each preset reference time interval; Perform multi-scale wavelet decomposition on the signal segment to obtain wavelet coefficient sequences at different scales; Calculate the signal envelope at each scale based on the wavelet coefficient sequence and perform normalization on the signal envelope; Based on the normalized signal envelope, Hilbert transform is used to extract the instantaneous phase features; The phase difference is calculated based on the instantaneous phase characteristics, and the instantaneous cross-correlation coefficient is obtained by combining the amplitude of the signal envelope.
[0009] By adopting the above technical solution, the audio signal is analyzed through a sliding time window, combined with multi-scale wavelet decomposition to obtain coefficient sequences at different scales, the signal envelope is calculated and normalized, and the instantaneous phase characteristics are extracted using the Hilbert transform, ultimately obtaining an accurate instantaneous cross-correlation coefficient. This cross-correlation calculation method based on signal envelope and phase characteristics can capture subtle variations in the audio signal in the time-frequency domain. By normalizing the signal envelope, the impact of signal amplitude changes on the cross-correlation calculation is reduced, improving the stability and reliability of the cross-correlation coefficient. Combined with the phase characteristics extracted by the Hilbert transform, the time-varying characteristics of the signal can be more accurately reflected, making subsequent mutation frame detection more precise. While ensuring calculation accuracy, this method reduces the amount of calculation through sliding analysis, improving the real-time performance of the system.
[0010] In conjunction with some embodiments of the first aspect, in some embodiments, calculating the energy value of the mutation frame in each frequency band specifically includes: Dividing the frequency spectrum of the mutation frame into multiple sub-bands according to a preset bandwidth; Perform weighted accumulation on the spectral components in each sub-band to obtain the initial sub-band energy value; Calculate the energy gradient between adjacent sub-bands and adjust the band boundaries according to the energy gradient; recalculating the spectral components of each sub-band based on the adjusted band boundaries; The recalculated spectral components are weighted and accumulated to obtain the energy value of each frequency band.
[0011] By adopting this technical solution, an adaptive frequency band division method is used to calculate the frequency band energy of the sudden change frame. The frequency band boundaries are dynamically adjusted by calculating the energy gradient between adjacent sub-bands. This makes the frequency band division more consistent with the energy distribution characteristics of the actual signal, improving the accuracy of the frequency band energy calculation. By weighted accumulation of spectral components, the contribution of different frequency components to the overall energy is considered, making the energy calculation result more reasonable.
[0012] In conjunction with some embodiments of the first aspect, in some embodiments, after converting the time domain signal into an audio signal for output, the method further includes: Mark the signal frames that are not marked as mutation frames as steady-state frames; Calculate the correlation coefficient between the stable frame and the adjacent mutation frame; Classify the steady-state frames whose correlation coefficient is greater than a preset coefficient threshold as transition frames, and classify the steady-state frames whose correlation coefficient is less than or equal to the preset coefficient threshold as pure steady-state frames; The same noise reduction method as the adjacent mutation frames is used for the transition frames, and the adaptive algorithm with a preset fixed step size is used for the pure steady-state frames to obtain the noise reduction signal corresponding to the steady-state frames; The noise reduction signal corresponding to the steady-state frame is subjected to inverse wavelet transform to obtain a time domain signal, which is then converted into an audio signal for output.
[0013] By employing the above technical solution, the correlation between steady-state frames and adjacent abrupt frames is calculated, further subdividing steady-state frames into transition frames and pure steady-state frames. Corresponding noise reduction strategies are then applied to each frame type. Transition frames are treated with the same noise reduction method as adjacent abrupt frames, ensuring the continuity and smoothness of the audio signal near the abrupt region. A fixed-step adaptive algorithm is used for noise reduction of pure steady-state frames, which stably suppresses background noise. This classification processing method based on frame correlation avoids sudden changes in the noise reduction effect between abrupt and steady-state frames, reducing the perceived discontinuity of the processed audio signal. By rationally classifying audio frames and applying corresponding noise reduction strategies, this method improves the coherence and naturalness of the processed audio signal while maintaining effective noise reduction.
[0014] In conjunction with some embodiments of the first aspect, in some embodiments, noise reduction processing is performed on the purely steady-state frame using an adaptive algorithm with a preset fixed step size, specifically including: Calculate the signal-to-noise ratio of the pure steady-state frame and determine the order of the adaptive filter according to the signal-to-noise ratio; Construct a Wiener filter based on the order; The coefficients of the Wiener filter are iteratively optimized using the minimum mean square error criterion with a preset fixed step size; The optimized filter coefficients are applied to the noise reduction process of pure steady-state frames.
[0015] By adopting the above technical solution, the order of the Wiener filter is determined by calculating the signal-to-noise ratio of the purely steady-state frame. The filter complexity can be dynamically adjusted according to the characteristics of the actual signal, achieving a balance between filter performance and computational overhead. When the signal-to-noise ratio is high, a lower-order filter can achieve good noise reduction, reducing unnecessary computational resource consumption. When the signal-to-noise ratio is low, the noise reduction capability is enhanced by increasing the filter order. Using the minimum mean square error criterion with a preset fixed step size for iterative coefficient optimization can improve convergence speed while ensuring convergence stability, avoiding the step size fluctuation problem that may occur with adaptive step size in purely steady-state signal processing. This Wiener filtering scheme based on adaptive signal-to-noise ratio adjustment not only ensures effective noise reduction for purely steady-state frames, but also avoids speech distortion caused by excessive noise reduction, while optimizing the system's computational efficiency.
[0016] In conjunction with some embodiments of the first aspect, in some embodiments, after calculating the correlation coefficient between the stable frame and the adjacent sudden change frame, the method further includes: Extract the short-term energy and zero-crossing rate of the steady-state frame; Calculate the energy ratio of the stable frame and the adjacent mutation frame; When the energy ratio is greater than a first energy threshold and the zero-crossing rate is less than a zero-crossing rate threshold, the steady-state frame is re-determined as a sudden change frame; When the correlation coefficient is greater than a preset coefficient threshold and the energy ratio is greater than a second energy threshold, the steady-state frame is determined to be a transition frame.
[0017] By employing this technical solution, when the inter-frame energy ratio is large and the zero-crossing rate is low, it indicates that the frame may contain important speech information. In this case, reclassifying it as a transition frame ensures more refined processing of the speech signal. Determining transition frames using the dual constraints of energy ratio and correlation coefficient accurately captures the gradual changes in the speech signal. This refined frame classification method improves the targeted noise reduction processing, enabling the system to adopt the most appropriate noise reduction strategy for speech segments with different characteristics, thereby achieving better noise reduction while maintaining speech continuity.
[0018] In conjunction with some embodiments of the first aspect, in some embodiments, extracting the short-term energy and zero-crossing rate of the steady-state frame specifically includes: Divide the steady-state frame into several subframes according to the preset number of sampling points; Calculate the short-time energy of each subframe, where the short-time energy is the square sum of the amplitudes of the sampling points in the subframe; Calculate the zero-crossing rate of each subframe, where the zero-crossing rate is the number of times the signs of adjacent sampling points in the subframe change; The short-time energy of the steady-state frame is obtained by adding up the short-time energies of all subframes and dividing by the number of subframes. The zero-crossing rate of the steady-state frame is obtained by adding up the zero-crossing rates of all subframes and dividing by the number of subframes.
[0019] By employing this technical solution, short-time energy and zero-crossing rate are calculated using a subframe approach. By finely dividing the sample points according to a pre-set number, local variations in the speech signal can be more accurately captured. In short-time energy calculation, the sum of the squared amplitudes of the sample points is used to highlight variations in signal strength. In zero-crossing rate calculation, the number of sign changes between adjacent sample points is counted to reflect variations in signal frequency. By averaging the features across all subframes, the overall signal characteristics are preserved while suppressing the influence of local random fluctuations, improving the accuracy and reliability of the calculated time-domain feature parameters.
[0020] In a second aspect, an embodiment of the present application provides a communication noise reduction system for an explosion-proof industrial telephone, the communication noise reduction system for the explosion-proof industrial telephone comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code comprising computer instructions, the one or more processors calling the computer instructions to enable the system to execute the method described in the first aspect and any possible implementation of the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on a system, enables the system to execute the method described in the first aspect and any possible implementation of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer program product, which, when executed on a system, enables the system to execute the method described in any possible implementation manner in the first aspect.
[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. The present application provides a communication noise reduction method for explosion-proof industrial telephones. The audio signal is collected by a main and auxiliary dual microphone, and the signal is subjected to wavelet transform to obtain multi-scale time-frequency features. The time-frequency correlation spectrum is constructed by calculating the instantaneous cross-correlation coefficient, which can accurately capture the time-frequency feature changes in the audio signal. The mutation frame is identified based on the change rate of the cross-correlation coefficient of adjacent time points. Combined with the energy distribution characteristics of the mutation frame in each frequency band, a fast-converging adaptive algorithm is used for the high-energy frequency band, and a regular step-size adaptive algorithm is used for the low-energy frequency band to perform noise reduction processing, so that the system can suppress noise while maintaining the mutation characteristics of the voice signal. By using a fast-converging adaptive algorithm for the high-energy frequency band, the changing characteristics of the voice signal can be quickly tracked and voice distortion can be reduced; by using a regular step-size adaptive algorithm for the low-energy frequency band, background noise can be smoothly suppressed to avoid audio distortion, reduce damage to the voice signal, maintain the continuity and stability of the call, and improve the communication noise reduction capability of the explosion-proof industrial telephone.
[0024] 2. The present application provides a communication noise reduction method for explosion-proof industrial telephones. By calculating the correlation between steady-state frames and adjacent mutation frames, the steady-state frames are further subdivided into transition frames and pure steady-state frames, and corresponding noise reduction strategies are adopted for different types of frames. The same noise reduction processing method as the adjacent mutation frames is adopted for transition frames to ensure the continuity and smoothness of the audio signal near the mutation area; the fixed-step adaptive algorithm is used for pure steady-state frames to perform noise reduction, which can stably suppress background noise. This classification processing method based on frame correlation avoids the sudden change in the noise reduction effect between mutation frames and steady-state frames, and reduces the sense of discontinuity after audio signal processing. This method improves the coherence and naturalness of the processed audio signal while ensuring the noise reduction effect by reasonably dividing the types of audio frames and adopting corresponding noise reduction strategies.
[0025] 3. This application provides a communication noise reduction method for explosion-proof industrial telephones. When the inter-frame energy ratio is large and the zero-crossing rate is low, it indicates that the frame may contain important voice information. At this time, re-dividing it into mutation frames can ensure that the voice signal is processed more finely. By determining the transition frame through the dual constraints of energy ratio and correlation coefficient, the gradual characteristics of the voice signal can be accurately captured. This refined frame type division method improves the targeted noise reduction processing, enabling the system to adopt the most suitable noise reduction strategy for voice segments with different characteristics, thereby achieving better noise reduction effects while maintaining voice continuity. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a flow chart of a method for reducing communication noise of an explosion-proof industrial telephone in an embodiment of the present application.
[0027] Figure 2 This is another flow chart of a method for reducing communication noise of an explosion-proof industrial telephone in an embodiment of the present application.
[0028] Figure 3 This is a schematic diagram of the physical device structure of a communication noise reduction system for an explosion-proof industrial telephone provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "said," "above," "the," and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to any or all possible combinations comprising one or more of the listed items.
[0030] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0031] The following uses an embodiment and combines Figure 1 , a communication noise reduction method for an explosion-proof industrial telephone in an embodiment of the present application is described: See also Figure 1 , which is a flow chart of a communication noise reduction method for an explosion-proof industrial telephone in an embodiment of the present application.
[0032] S101, collecting a first audio signal containing a target voice through a primary microphone, and simultaneously collecting a second audio signal containing ambient noise through an auxiliary microphone; This step involves collecting audio signals. The system uses the primary and secondary microphones installed on the explosion-proof industrial telephone to collect a first audio signal containing the target speech and a second audio signal containing ambient noise. The primary microphone primarily collects the user's speech signal, while the secondary microphone primarily collects the ambient noise signal. Collecting these two different types of audio signals provides raw data for subsequent noise suppression and speech enhancement processing.
[0033] In a specific implementation, the system can control the synchronization of the primary and secondary microphones to ensure that the collected first and second audio signals are synchronized in time. In addition, the system can also pre-process the collected audio signals, such as filtering, noise reduction, and gain control, to improve the quality of the audio signals.
[0034] In this step, if the distance between the primary and secondary microphones is not set properly, the collected first and second audio signals may be highly correlated, affecting the effectiveness of subsequent processing. To address this issue, the system can reasonably set the distance between the primary and secondary microphones, for example, by setting the distance between the two microphones to greater than a certain threshold, to ensure that the correlation between the two collected audio signals is minimized, thereby improving the effectiveness of subsequent processing.
[0035] S102: Calculate the instantaneous cross-correlation coefficient of the first audio signal and the second audio signal at each scale at a preset reference time interval, and construct a time-frequency correlation graph based on the instantaneous cross-correlation coefficient; The system calculates the instantaneous cross-correlation coefficients of the first and second audio signals at each scale at a preset reference time interval. Specifically, the system performs sliding analysis on the first and second audio signals according to the preset reference time interval to obtain signal segments within each preset reference time interval; performs multi-scale wavelet decomposition on the signal segments to obtain wavelet coefficient sequences at different scales; calculates the signal envelope at each scale based on the wavelet coefficient sequence and normalizes the signal envelope; extracts instantaneous phase features based on the normalized signal envelope using a Hilbert transform; calculates the phase difference based on the instantaneous phase features and combines the amplitude of the signal envelope to obtain the instantaneous cross-correlation coefficient. Subsequently, a time-frequency correlation map is constructed based on the instantaneous cross-correlation coefficients.
[0036] The system performs sliding analysis on the first and second audio signals, respectively, at preset reference time intervals, to obtain signal segments within each reference time interval. Sliding analysis is a commonly used time window processing method that captures local features of audio signals within different time periods by continuously sliding the time window. The size of the preset reference time interval determines the temporal resolution of the sliding analysis. Smaller intervals result in higher temporal resolution, but also increase the computational effort. Larger intervals result in lower temporal resolution, but also improve computational efficiency. Therefore, the selection of the preset reference time interval requires a comprehensive consideration of the balance between temporal resolution and computational efficiency.
[0037] By performing multiscale wavelet decomposition on the signal segments within each reference time interval, we can obtain sequences of wavelet coefficients at different scales. Wavelet decomposition is a time-frequency analysis method that, by selecting appropriate wavelet basis functions, decomposes the signal into wavelet coefficients at different scales, reflecting the local characteristics of the signal at different frequency bands. Larger scales correspond to lower frequencies, and the resulting wavelet coefficients reflect the signal's low-frequency components; smaller scales correspond to higher frequencies, and the resulting wavelet coefficients reflect the signal's high-frequency components. Multiscale wavelet decomposition can comprehensively characterize the signal's time-frequency distribution characteristics at different frequency bands.
[0038] After obtaining the wavelet coefficient sequences at different scales, the system calculates the signal envelope at each scale based on the wavelet coefficient sequences. The signal envelope reflects the energy distribution of the signal at each scale. By analyzing the signal envelope, we can obtain information about the energy variations of the signal at different frequency bands. To facilitate subsequent processing and comparison, the system normalizes the signal envelope to a value within the [0, 1] range. Normalization eliminates differences in signal amplitude at different scales, making the signal envelopes at different scales comparable.
[0039] While the normalized signal envelope reflects the signal's energy distribution, it does not contain phase information. To obtain the signal's phase characteristics, the system uses the Hilbert transform to process the normalized signal envelope and extract the instantaneous phase characteristics. The Hilbert transform is a commonly used signal processing method that converts a real signal into a complex signal, obtaining the real and imaginary parts, and then calculating the signal's instantaneous phase. The instantaneous phase reflects the signal's phase value at each point in time. By analyzing the changes in the instantaneous phase, the signal's phase characteristics can be obtained.
[0040] After extracting the instantaneous phase characteristics of the first and second audio signals, the system calculates the phase difference between the two signals at each time point, known as the instantaneous phase difference. This instantaneous phase difference reflects the degree of phase difference between the two signals at each time point. A larger phase difference indicates a lower correlation between the two signals at that time point; a smaller phase difference indicates a higher correlation between the two signals at that time point. Combining the instantaneous phase difference with the normalized signal envelope amplitude yields the instantaneous cross-correlation coefficient between the first and second audio signals at each scale. This coefficient comprehensively considers the signal's energy and phase information, comprehensively reflecting the strength of the correlation between the two signals at different time and frequency points.
[0041] Finally, the system constructs a time-frequency correlation map based on the instantaneous cross-correlation coefficient. The time-frequency correlation map is an intuitive time-frequency analysis tool that maps the instantaneous cross-correlation coefficient onto a two-dimensional plane, using time and frequency as the two dimensions, to form an image. In the time-frequency correlation map, the horizontal axis represents time, the vertical axis represents frequency, and the brightness or color of the image represents the magnitude of the instantaneous cross-correlation coefficient. By observing the time-frequency correlation map, one can intuitively understand the distribution of the correlation between the first and second audio signals at different time and frequency points, identify areas with strong or weak correlation, and provide an important reference for subsequent noise suppression and speech enhancement processing.
[0042] S103, calculating the change rate of the mutual correlation coefficient of adjacent time points in the time-frequency correlation graph, and when the change rate is greater than a preset threshold, marking the signal frame corresponding to the time point where the change rate is greater than the preset threshold as a mutation frame; This step is called mutation frame detection. The system determines the temporal mutation of the audio signal by calculating the rate of change of the cross-correlation coefficients between adjacent time points in the time-frequency correlation graph. When the rate of change of the cross-correlation coefficients between adjacent time points exceeds a preset threshold, the system marks the signal frame corresponding to the time point where the rate of change exceeds the threshold as a mutation frame.
[0043] In practical implementation, the system can use a differential method to calculate the rate of change of the cross-correlation coefficient between adjacent time points. This involves subtracting the cross-correlation coefficient of the previous time point from the current time point, and then dividing the result by the interval between the two time points. Furthermore, the system can flexibly set a preset threshold for the rate of change based on the actual application scenario and the characteristics of the audio signal, thereby balancing the accuracy and recall of sudden frame detection.
[0044] S104, calculating the energy value of the mutation frame in each frequency band, and taking the average value of the energy value as the energy reference value; The system calculates the energy value of each frequency band of the mutation frame. Specifically, the system divides the spectrum of the mutation frame into multiple sub-bands according to a preset bandwidth; performs weighted accumulation of the spectral components within each sub-band to obtain an initial sub-band energy value; calculates the energy gradient between adjacent sub-bands and adjusts the band boundaries based on the energy gradient; recalculates the spectral components of each sub-band based on the adjusted band boundaries; and performs weighted accumulation of the recalculated spectral components to obtain the energy value of each band. The average of the energy values is used as the energy reference value. Bands with energy values greater than the energy reference value are determined as high-energy bands, and bands with energy values less than or equal to the energy reference value are determined as low-energy bands.
[0045] The system performs spectral analysis on detected mutation frames and calculates the energy values of each frequency band. Band energy reflects the energy distribution of the mutation frame within different frequency ranges. The calculation and analysis of band energy provides important reference information for subsequent adaptive noise reduction processing.
[0046] In its implementation, the system first divides the frequency spectrum of the mutation frame into multiple sub-bands according to a preset bandwidth. The preset bandwidth determines the frequency range of each sub-band, and the bandwidth size needs to be reasonably set based on parameters such as the audio signal's sampling rate and frequency resolution. Generally, the smaller the bandwidth, the more sub-bands there are and the higher the frequency resolution, but the computational complexity also increases accordingly; the larger the bandwidth, the fewer sub-bands there are and the lower the frequency resolution, but the computational complexity also decreases accordingly. Therefore, the selection of the preset bandwidth requires a trade-off between frequency resolution and computational complexity to meet the needs of practical applications.
[0047] After dividing the frequency subbands, the system performs a weighted accumulation of the spectral components within each subband to obtain the initial subband energy value. Weighted accumulation aims to emphasize the importance of different frequency components. Common weighting methods include energy weighting and perceptual weighting. Energy weighting directly uses the square of the amplitude of the spectral components as weights, while perceptual weighting takes into account the human ear's sensitivity to different frequencies and applies different weights to different frequency components. The initial subband energy value obtained after weighted accumulation reflects the total energy within each subband.
[0048] To further improve the accuracy of band energy calculation, the system can also calculate the energy gradient between adjacent sub-bands and adjust the band boundaries based on the energy gradient. The energy gradient reflects the changing trend of the energy values between adjacent sub-bands. By analyzing the energy gradient, it is possible to detect energy mutations at band boundaries. When the energy gradient between adjacent sub-bands is large, it indicates that important spectral components may exist at the band boundary. In this case, it is possible to consider adjusting the band boundary in the direction of the larger energy gradient to better preserve these important components. The adjusted band boundaries are more consistent with the actual spectral characteristics of the audio signal, which helps to improve the accuracy of band energy calculation.
[0049] Based on the adjusted band boundaries, the system recalculates the spectral components of each sub-band and performs a weighted accumulation of the recalculated spectral components to obtain the final energy values for each band. These energy values comprehensively consider factors such as band division, energy weighting, and band boundary adjustment, and can more accurately reflect the energy distribution of the mutation frame in different bands.
[0050] After calculating the energy values for each frequency band, the system uses the average of the energy values as the energy baseline. The energy baseline reflects the overall energy level of the mutation frame and divides the frequency band energy values into two intervals: high energy and low energy. Frequency bands with energy values greater than the energy baseline are identified as high-energy bands, indicating that the spectral components within these bands have stronger energy and may contain more critical information, requiring more protection and enhancement during noise reduction. Frequency bands with energy values less than or equal to the energy baseline are identified as low-energy bands, indicating that the spectral components within these bands have weaker energy and may be mainly noise components, requiring more suppression and reduction during noise reduction.
[0051] By dividing the energy baseline into different frequency bands, the system can adopt different noise reduction strategies based on the energy level of the frequency bands, providing targeted processing for high-energy and low-energy frequency bands, thereby improving noise reduction effectiveness and audio quality. For example, for high-energy frequency bands, the system can use a more conservative noise reduction method to preserve more voice information; for low-energy frequency bands, the system can use a more aggressive noise reduction method to effectively suppress noise components.
[0052] S105, performing noise reduction processing on the high-energy frequency band using an adaptive algorithm with a first preset step size and a second preset step size for the low-energy frequency band, to obtain a noise reduction signal corresponding to the sudden change frame; This step is the adaptive noise reduction process. The system applies different adaptive algorithms to the sudden change frames based on the energy level of the frequency bands. For high-energy frequency bands, the system uses a fast-converging adaptive algorithm with a first preset step size to quickly suppress the noise components in the high-energy bands. For low-energy frequency bands, the system uses an adaptive algorithm with a second preset step size to avoid over-suppressing the speech components in the low-energy bands. By processing the high- and low-energy frequency bands separately, the system can obtain the noise reduction signal corresponding to the sudden change frames.
[0053] In specific implementations, the system can employ various adaptive algorithms for noise reduction, such as the Least Mean Squared Error (LMS) algorithm, the Normalized Least Mean Squared Error (NLMS) algorithm, and the Affine Projection (AP) algorithm. The first preset step size is typically larger than the second preset step size to achieve rapid convergence in high-energy frequency bands. Furthermore, the system can adaptively adjust the preset step size based on the statistical characteristics of the signal. For example, a variable-step-size LMS algorithm can be employed to dynamically adjust the step size based on the instantaneous gradient of the signal, thereby improving the convergence speed and stability of the noise reduction process.
[0054] S106 , performing inverse wavelet transform on the noise reduction signal to obtain a time domain signal, and converting the time domain signal into an audio signal for output.
[0055] This step is the audio signal reconstruction step. The system performs an inverse wavelet transform on the noise-reduced signal, reconstructing the frequency-domain signal into a time-domain signal. The system then converts the reconstructed time-domain signal into a standard audio signal format, such as PCM, and ultimately outputs the noise-reduced audio signal.
[0056] During implementation, the system needs to perform an inverse wavelet transform on the noise-reduced signal based on the wavelet transform parameter settings, such as the type of wavelet basis function and the number of decomposition levels. At the same time, the system also needs to convert the format of the reconstructed time-domain signal and adjust parameters, such as the sampling rate and quantization bit number, according to the requirements of the output device to meet the playback requirements of the output device.
[0057] In this step, if the inverse wavelet transform parameters are not set appropriately, the reconstructed time domain signal may be distorted or contain residual noise. To address this issue, the system can improve the quality of signal reconstruction by optimizing the selection of wavelet basis functions and the number of decomposition layers, such as using the same parameter settings as the wavelet transform stage, or selecting the optimal parameter combination through experimentation and evaluation. In addition, the system can also perform post-processing on the time domain signal before converting it to an audio signal, such as smoothing filtering or dynamic range compression, to further improve the sound quality and listening experience of the audio signal.
[0058] In the above embodiment, audio signals are collected using dual primary and secondary microphones, wavelet transforms are performed on the signals to obtain multi-scale time-frequency features, and a time-frequency correlation map is constructed by calculating the instantaneous cross-correlation coefficient. This accurately captures the time-frequency feature changes in the audio signal. Mutation frames are identified based on the rate of change of the cross-correlation coefficients between adjacent time points. Combined with the energy distribution characteristics of the mutation frames in each frequency band, a fast-converging adaptive algorithm is used for high-energy bands, while a regular-step-size adaptive algorithm is used for low-energy bands to perform noise reduction. This allows the system to suppress noise while maintaining the mutation characteristics of the speech signal. By using a fast-converging adaptive algorithm for high-energy bands, the changing characteristics of the speech signal can be quickly tracked, reducing speech distortion; while using a regular-step-size adaptive algorithm for low-energy bands, background noise can be smoothly suppressed, avoiding audio distortion, minimizing damage to the speech signal, maintaining call continuity and stability, and improving the noise reduction capabilities of explosion-proof industrial telephones.
[0059] In the above embodiment, the processing of mutation frames is mainly described. After a series of processing, the audio signals collected by the main and auxiliary dual microphones are effectively noise-reduced for the mutation frames. However, in actual applications, in addition to the mutation frames, the audio signal also contains a large number of non-mutation signal segments. In order to achieve comprehensive noise reduction processing for the entire audio signal, the system also needs to effectively reduce the noise of the non-mutation signal segments. Figure 2 , another method for reducing communication noise of an explosion-proof industrial telephone in an embodiment of the present application is described: See also Figure 2 , is another flow chart of a communication noise reduction method for an explosion-proof industrial telephone in an embodiment of the present application.
[0060] S201, marking signal frames that are not marked as mutation frames as steady-state frames; This step marks steady-state frames. During the previous mutation frame detection process, the system marked signal frames corresponding to time points where the rate of change exceeded a preset threshold as mutation frames. For all other signal frames not marked as mutation frames, the system marks them as steady-state frames. A steady-state frame indicates that the energy and spectral characteristics of the audio signal within that frame change relatively smoothly, without noticeable mutations.
[0061] In specific implementations, the system can use different methods to mark steady-state frames, such as setting a specific flag bit in the attribute field of the signal frame or adding the index of the steady-state frame to the steady-state frame list. The marked steady-state frame will be treated as a special frame type in subsequent processing, distinguished from the sudden change frame.
[0062] In this step, if the previous mutation frame detection results are inaccurate, some mutation frames may be mistakenly marked as steady-state frames, affecting subsequent processing. To address this issue, the system can improve the accuracy of mutation frame detection by optimizing the mutation frame detection algorithm. For example, by introducing techniques such as adaptive thresholds or multi-feature fusion to reduce missed mutation frame detections. Furthermore, after marking steady-state frames, the system can perform a secondary verification of the marking results. For example, by performing energy or spectrum analysis on signal frames marked as steady-state frames, the system can identify potentially mislabeled mutation frames and re-mark them.
[0063] S202, calculating the correlation coefficient between the stable frame and the adjacent mutation frame; The system measures the degree of similarity in energy or spectral characteristics between a steady-state frame and its adjacent sudden-change frame by calculating the correlation coefficient. A larger correlation coefficient indicates closer characteristics between the steady-state and sudden-change frames, suggesting a potential transition relationship between them. A smaller correlation coefficient indicates a significant difference in characteristics between the steady-state and sudden-change frames, suggesting a weaker transition relationship.
[0064] In specific implementations, the system can use various correlation metrics to calculate the correlation coefficient between steady-state frames and sudden change frames, such as the Pearson correlation coefficient, Spearman rank correlation coefficient, and mutual information. The Pearson correlation coefficient is the most commonly used linear correlation metric, measuring the degree of linear correlation between two variables by calculating the ratio of their covariance and standard deviation. Furthermore, the system can select an appropriate correlation metric based on the characteristics of the audio signal. For example, for speech signals, a correlation metric based on Mel-Frequency Cepstral Coefficients (MFCCs) can be used to better reflect the perceptual characteristics of speech signals.
[0065] S203, classifying steady-state frames whose correlation coefficients are greater than a preset coefficient threshold as transition frames, and classifying steady-state frames whose correlation coefficients are less than or equal to the preset coefficient threshold as pure steady-state frames; The system classifies steady-state frames into two categories: transition frames and pure steady-state frames based on the correlation coefficient between them and adjacent abrupt frames. Stable-state frames with correlation coefficients greater than a preset threshold are classified as transition frames, indicating that these steady-state frames and adjacent abrupt frames are similar in energy or spectral characteristics and have a certain transition relationship. Stable-state frames with correlation coefficients less than or equal to the preset threshold are classified as pure steady-state frames, indicating that these steady-state frames and adjacent abrupt frames have significantly different energy or spectral characteristics and a weak transition relationship.
[0066] During implementation, the system can flexibly set the preset coefficient threshold based on the characteristics of the audio signal and actual application requirements. Generally, a larger preset coefficient threshold results in fewer steady-state frames being classified as transition frames and more pure steady-state frames; a smaller preset coefficient threshold results in more steady-state frames being classified as transition frames and fewer pure steady-state frames. The system can select the optimal threshold setting to achieve a balanced ratio of transition frames to pure steady-state frames through experimentation and statistical analysis of different preset coefficient thresholds.
[0067] S204, applying the same noise reduction method as that used for adjacent sudden change frames to the transition frames, and applying a preset fixed-step adaptive algorithm to the pure steady-state frames to perform noise reduction processing, thereby obtaining a noise reduction signal corresponding to the steady-state frames; The system uses the same noise reduction processing method as the adjacent mutation frames for transition frames, and uses an adaptive algorithm with a preset fixed step size to perform noise reduction processing on pure steady-state frames to obtain the noise reduction signal corresponding to the steady-state frame, wherein the adaptive algorithm with a preset fixed step size is used to perform noise reduction processing on the pure steady-state frame, specifically including: calculating the signal-to-noise ratio of the pure steady-state frame, and determining the order of the adaptive filter according to the signal-to-noise ratio; constructing a Wiener filter based on the order; iteratively optimizing the coefficients of the Wiener filter using the minimum mean square error criterion with a preset fixed step size; and applying the optimized filter coefficients to the noise reduction processing of the pure steady-state frame.
[0068] The system applies different noise reduction methods to steady-state frames depending on their type. For transition frames, since their characteristics are similar to those of adjacent abrupt frames, the system uses the same noise reduction method as adjacent abrupt frames. This method employs an adaptive noise reduction algorithm based on frequency band energy, adaptively adjusting the noise reduction intensity based on the energy distribution of the transition frame across different frequency bands. This method fully utilizes the noise reduction results of adjacent abrupt frames, reducing the complexity of the transition frame noise reduction process while ensuring a smooth transition between transition frames and abrupt frames, avoiding sudden changes and discontinuities in the noise reduction process.
[0069] For purely steady-state frames, since their energy and spectral characteristics change slowly, the system uses an adaptive algorithm with a preset fixed step size for noise reduction. That is, a fixed update step size is used during the noise reduction process to ensure the stability of the noise reduction process. Specifically, the system first calculates the signal-to-noise ratio of the purely steady-state frame. The signal-to-noise ratio reflects the relative strength of the speech component and the noise component in the purely steady-state frame. Based on the size of the signal-to-noise ratio, the system determines the order of the adaptive filter, and the order determines the length and complexity of the adaptive filter. The higher the signal-to-noise ratio, the greater the proportion of the speech component and the smaller the proportion of the noise component. In this case, a lower-order adaptive filter can be used; the lower the signal-to-noise ratio, the greater the proportion of the noise component and the smaller the proportion of the speech component. In this case, a higher-order adaptive filter is required.
[0070] After determining the order, the system constructs a Wiener filter based on that order. The Wiener filter is a classic adaptive filter that optimizes the filter coefficients by minimizing the mean square error between the output signal and the desired signal. In noise reduction, the Wiener filter input is a pure steady-state frame, and the desired signal is the pure speech component of the pure steady-state frame. The filter coefficients are iteratively optimized using an adaptive algorithm to make the filter output signal as close as possible to the desired signal, thereby achieving the purpose of noise reduction.
[0071] The system iteratively optimizes the Wiener filter coefficients using the Least Mean Squared Error (LMS) criterion with a preset fixed step size. The LMS criterion is a simple and effective adaptive filter optimization criterion. It estimates the error between the filter output signal and the desired signal and uses this error to update the filter coefficients using gradient descent, continuously bringing the filter output signal closer to the desired signal. The preset fixed step size controls the magnitude of the filter coefficient adjustment during each iterative optimization. The step size requires a trade-off between convergence speed and stability. Through multiple iterative optimizations, the Wiener filter coefficients are continuously updated, continuously enhancing its noise suppression capabilities.
[0072] Finally, the system applies the optimized filter coefficients to the purely steady-state frames for noise reduction. Specifically, the optimized Wiener filter is used to filter the purely steady-state frames, generating the corresponding noise-reduced signals. In these purely steady-state frames, most noise components are effectively suppressed, while the speech components are preserved and enhanced, improving signal quality and clarity.
[0073] S205 , performing inverse wavelet transform on the noise reduction signal corresponding to the steady-state frame to obtain a time domain signal, and converting the time domain signal into an audio signal for output.
[0074] The system reconstructs the noise-reduced signal corresponding to the steady-state frame after noise reduction into a time-domain signal through an inverse wavelet transform. The inverse wavelet transform is the inverse process of the wavelet transform, restoring the frequency-domain signal to the time-domain signal, restoring the signal's original form. During the inverse wavelet transform, the system uses the same wavelet basis functions and decomposition levels as the wavelet transform to ensure accurate signal reconstruction. The reconstructed time-domain signal undergoes necessary format conversion and parameter adjustments before being converted to a standard audio signal format, such as PCM, for output to subsequent playback devices or storage media.
[0075] In practice, the system uses the same processing flow and parameter settings as for signal reconstruction of sudden-frame signals, performing inverse wavelet transform and audio signal conversion on the noise-reduced signals corresponding to the steady-state frames. Furthermore, the system can perform post-processing on the reconstructed time-domain signals based on actual application requirements, such as removing residual background noise and increasing the signal's dynamic range, to further improve the audio signal quality.
[0076] In the above embodiment, the correlation between steady-state frames and adjacent abrupt frames is calculated to further subdivide steady-state frames into transition frames and pure steady-state frames, and corresponding noise reduction strategies are adopted for different types of frames. Transition frames are treated with the same noise reduction method as adjacent abrupt frames, ensuring the continuity and smoothness of the audio signal near the abrupt region. A fixed-step adaptive algorithm is used for noise reduction in pure steady-state frames, which can stably suppress background noise. This classification processing method based on frame correlation avoids sudden changes in the noise reduction effect between abrupt and steady-state frames, reducing the perceived discontinuity of the processed audio signal. By rationally classifying audio frames and adopting corresponding noise reduction strategies, this method improves the coherence and naturalness of the processed audio signal while ensuring effective noise reduction.
[0077] Furthermore, in another embodiment, after calculating the correlation coefficient between the steady-state frame and the adjacent sudden change frame, the system further includes: extracting the short-time energy and zero-crossing rate of the steady-state frame, specifically including: dividing the steady-state frame into a plurality of subframes according to a preset number of sampling points; calculating the short-time energy of each subframe, where the short-time energy is the sum of the squares of the amplitudes of the sampling points in the subframe; calculating the zero-crossing rate of each subframe, where the zero-crossing rate is the number of times the signs of adjacent sampling points in the subframe change; summing the short-time energies of all subframes and dividing the sum by the number of subframes to obtain the short-time energy of the steady-state frame, and summing the zero-crossing rates of all subframes and dividing the sum by the number of subframes to obtain the zero-crossing rate of the steady-state frame; Calculate the energy ratio of the stable frame and the adjacent mutation frame; When the energy ratio is greater than a first energy threshold and the zero-crossing rate is less than a zero-crossing rate threshold, the steady-state frame is re-determined as a sudden change frame; When the correlation coefficient is greater than a preset coefficient threshold and the energy ratio is greater than a second energy threshold, the steady-state frame is determined to be a transition frame.
[0078] The system extracts the short-time energy and zero-crossing rate features of the steady-state frame. Short-time energy reflects the energy change of the signal within a short time interval, and the zero-crossing rate reflects the frequency change of the signal within a short time interval. To extract these two features, the system divides the steady-state frame into several subframes according to a preset number of sampling points, and each subframe contains a fixed number of sampling points. The system then calculates the short-time energy and zero-crossing rate of each subframe. The short-time energy is calculated by squaring the amplitude of each sampling point in the subframe and adding all the squared values to obtain the short-time energy of the subframe. The zero-crossing rate is calculated by counting the number of sign changes between adjacent sampling points in the subframe, that is, the number of times the positive and negative signs of adjacent sampling points change, to obtain the zero-crossing rate of the subframe.
[0079] After calculating the short-term energy and zero-crossing rate of all subframes, the system adds up the short-term energies of all subframes and divides the sum by the number of subframes to obtain the short-term energy of the entire steady-state frame. Similarly, the zero-crossing rate of all subframes is added up and divided by the number of subframes to obtain the zero-crossing rate of the entire steady-state frame. This averaging process can reduce the impact of individual subframe feature fluctuations on the overall feature, improving the stability and reliability of the feature.
[0080] Next, the system calculates the energy ratio between the steady-state frame and the adjacent abrupt frame. This energy ratio reflects the relative energy of the steady-state frame and the abrupt frame, and can be used to determine whether the steady-state frame exhibits abrupt frame characteristics. If the steady-state frame's energy is significantly greater than the abrupt frame's, and the energy ratio exceeds a preset first energy threshold, the steady-state frame may have been misidentified as a abrupt frame and needs to be re-identified.
[0081] In addition to the energy ratio determination, the system also incorporates the zero-crossing rate as a criterion. Generally speaking, speech signals have a low zero-crossing rate, while noise signals have a high zero-crossing rate. If the zero-crossing rate of a steady-state frame is below the preset zero-crossing rate threshold, it indicates that the frame is primarily speech-based. Combined with the energy ratio determination, it can be reclassified as a sudden change frame to ensure accurate sudden change frame detection.
[0082] Finally, the system combines the correlation coefficient and energy ratio to determine the final classification of steady-state frames. If the correlation coefficient between a steady-state frame and an adjacent sudden change frame exceeds a preset coefficient threshold, and the energy ratio is also greater than a second energy threshold, it indicates that the steady-state frame and the sudden change frame are highly correlated in time-frequency characteristics and have similar energy characteristics, and can be identified as a transition frame. Transition frames are a special type of frame between steady-state frames and sudden change frames. They have certain sudden change characteristics but also have a transitional relationship with sudden change frames. They require similar processing as sudden change frames to ensure signal continuity and smoothness.
[0083] In the above embodiment, when the inter-frame energy ratio is large and the zero-crossing rate is low, it indicates that the frame may contain important speech information. In this case, reclassifying it as a transition frame ensures more refined processing of the speech signal. Determining transition frames using the dual constraints of energy ratio and correlation coefficient accurately captures the gradual changes in the speech signal. This refined frame classification method improves the targeted noise reduction processing, enabling the system to adopt the most appropriate noise reduction strategy for speech segments with different characteristics, thereby achieving better noise reduction while maintaining speech continuity.
[0084] The following describes the system in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , which is a schematic diagram of the physical device structure of a communication noise reduction system for an explosion-proof industrial telephone provided in an embodiment of the present application.
[0085] It should be noted that Figure 3 The structure of the system shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0086] like Figure 3 As shown, the system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes, such as the methods described in the above embodiments, based on programs stored in a read-only memory (ROM) 302 or programs loaded from a storage unit 308 into a random access memory (RAM) 303. RAM 303 also stores various programs and data required for system operation. CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.
[0087] The following components are connected to the I / O interface 305: an input section 306 including a camera, infrared sensor, and the like; an output section 307 including a liquid crystal display (LCD) and speakers; a storage section 308 including a hard disk and the like; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is installed in the drive 310 as needed, so that computer programs read from the media can be installed in the storage section 308 as needed.
[0088] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 309 and / or installed from removable media 311. When executed by the central processing unit (CPU) 301, the computer program performs the various functions defined in the present invention.
[0089] It should be noted that the computer-readable medium described in the embodiments of the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium may include a data signal transmitted in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal may take any of a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0091] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the system described in the above embodiments, or may exist independently and not incorporated into the system. The storage medium carries one or more computer programs, and when executed by a processor of a system, the system implements the methods provided in the above embodiments.
[0092] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0093] As used in the above embodiments, the term “when…” may be interpreted as “if…” or “after…” or “in response to determining…” or “in response to detecting…”, depending on the context. Similarly, the phrases “upon determining…” or “if (stated condition or event) is detected” may be interpreted as “if determining…” or “in response to determining…” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.
[0094] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0095] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for reducing communication noise of an explosion-proof industrial telephone, characterized in that: include: The first audio signal containing the target speech is collected by the main microphone, and the second audio signal containing the ambient noise is collected by the auxiliary microphone; calculating the instantaneous cross-correlation coefficient of the first audio signal and the second audio signal at each scale at a preset reference time interval, and constructing a time-frequency correlation graph based on the instantaneous cross-correlation coefficient; Calculating a change rate of a correlation coefficient between adjacent time points in the time-frequency correlation graph, and when the change rate is greater than a preset threshold, marking a signal frame corresponding to a time point where the change rate is greater than the preset threshold as a mutation frame; Calculating the energy value of the mutation frame in each frequency band, and taking the average of the energy values as an energy reference value, determining the frequency band whose energy value is greater than the energy reference value as a high-energy frequency band, and determining the frequency band whose energy value is less than or equal to the energy reference value as a low-energy frequency band; Performing noise reduction processing on the high-energy frequency band using an adaptive algorithm with a first preset step size and fast convergence, and on the low-energy frequency band using an adaptive algorithm with a second preset step size, to obtain a noise reduction signal corresponding to the sudden change frame; The noise reduction signal is subjected to inverse wavelet transform to obtain a time domain signal, and the time domain signal is converted into an audio signal for output.
2. The method according to claim 1, characterized in that The calculating the instantaneous cross-correlation coefficient of the first audio signal and the second audio signal at each scale at a preset reference time interval specifically includes: Performing sliding analysis on the first audio signal and the second audio signal according to a preset reference time interval to obtain a signal segment within each of the preset reference time intervals; Performing multi-scale wavelet decomposition on the signal segment to obtain wavelet coefficient sequences at different scales; Calculating the signal envelope at each scale according to the wavelet coefficient sequence, and normalizing the signal envelope; Extracting instantaneous phase features based on the normalized signal envelope using Hilbert transform; The phase difference is calculated according to the instantaneous phase feature, and the instantaneous cross-correlation coefficient is obtained in combination with the amplitude of the signal envelope.
3. The method according to claim 1, characterized in that The calculating the energy value of the mutation frame in each frequency band specifically includes: Dividing the frequency spectrum of the mutation frame into a plurality of sub-bands according to a preset bandwidth; Performing weighted accumulation on the spectral components within each of the sub-bands to obtain an initial sub-band energy value; Calculating energy gradients between adjacent sub-frequency bands, and adjusting frequency band boundaries according to the energy gradients; recalculating the spectral components of each of the sub-bands based on the adjusted frequency band boundaries; The recalculated spectral components are weighted and accumulated to obtain the energy value of each frequency band.
4. The method according to claim 1, wherein After converting the time domain signal into an audio signal for output, the method further includes: marking the signal frames that are not marked as the mutation frames as stable frames; Calculating the correlation coefficient between the stable frame and the adjacent mutation frame; Classifying the steady-state frames whose correlation coefficient is greater than a preset coefficient threshold as transition frames, and classifying the steady-state frames whose correlation coefficient is less than or equal to the preset coefficient threshold as pure steady-state frames; The transition frame is subjected to the same noise reduction processing as that of the adjacent sudden change frame, and the pure steady-state frame is subjected to noise reduction processing using an adaptive algorithm with a preset fixed step size to obtain a noise reduction signal corresponding to the steady-state frame; Performing inverse wavelet transform on the noise reduction signal corresponding to the steady-state frame to obtain a time domain signal, and converting the time domain signal into an audio signal for output.
5. The method according to claim 4, characterized in that The noise reduction process of the pure steady-state frame using an adaptive algorithm with a preset fixed step size specifically includes: Calculating a signal-to-noise ratio of the pure steady-state frame, and determining an order of an adaptive filter according to the signal-to-noise ratio; constructing a Wiener filter based on the order; Iteratively optimizing the coefficients of the Wiener filter using a minimum mean square error criterion with a preset fixed step size; The optimized filter coefficients are applied to the noise reduction process of the pure steady-state frame.
6. The method according to claim 4, characterized in that After calculating the correlation coefficient between the steady-state frame and the adjacent mutation frame, the method further includes: extracting the short-time energy and zero-crossing rate of the steady-state frame; Calculating an energy ratio between the steady-state frame and the adjacent mutation frame; When the energy ratio is greater than a first energy threshold and the zero-crossing rate is less than a zero-crossing rate threshold, re-determining the steady-state frame as the sudden change frame; When the correlation coefficient is greater than the preset coefficient threshold and the energy ratio is greater than a second energy threshold, the steady-state frame is determined to be a transition frame.
7. The method according to claim 6, characterized in that Extracting the short-time energy and zero-crossing rate of the steady-state frame specifically includes: Dividing the steady-state frame into a plurality of subframes according to a preset number of sampling points; Calculating the short-time energy of each subframe, where the short-time energy is the sum of the squares of the amplitudes of the sampling points in the subframe; Calculating a zero-crossing rate of each subframe, where the zero-crossing rate is the number of times that signs of adjacent sampling points in the subframe change; The short-time energy of the steady-state frame is obtained by adding up the short-time energies of all subframes and dividing the sum by the number of subframes, and the zero-crossing rate of the steady-state frame is obtained by adding up the zero-crossing rate of all subframes and dividing the sum by the number of subframes.
8. A communication noise reduction system for explosion-proof industrial telephones, characterized in that: The system comprises: One or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a system, the system is caused to perform the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that When the computer program product is run on a system, the system is caused to perform the method according to any one of claims 1 to 7.