An airborne high-reliability intelligent voice call noise reduction method

Through the hierarchical design of mixing, noise reduction and automatic gain, the noise suppression and voice intelligibility problems of the airborne voice communication system in complex noise environments are solved, and high-reliability voice calls with high clarity and low latency are achieved, meeting the top standards of avionics.

CN120564745BActive Publication Date: 2025-09-26西安赛普特信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511062745.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-26
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Airborne voice communication systems face problems such as insufficient noise suppression, high computing resource requirements, reduced voice intelligibility, and difficulty meeting low latency requirements in complex noisy environments.

Method used

It adopts a hierarchical design of mixing, noise reduction, squelch and automatic gain, and achieves high-reliability voice calls through ADC conversion, FFT transformation, spectral entropy weighting, auditory masking compensation, pre-trained noise reduction model, dynamic sub-band GMM modeling and automatic gain processing.

Benefits of technology

Achieve high-reliability voice calls with high clarity and low latency in complex noisy environments, meeting top avionics standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564745B_ABST
    Figure CN120564745B_ABST
Patent Text Reader

Abstract

The present application relates to the field of audio processing technology, and in particular to an airborne high-reliability intelligent voice call noise reduction method, comprising: performing ADC conversion on an acquired analog audio signal to obtain a digital audio signal; performing frame division and FFT transformation on the digital audio signal, performing phase optimization, calculating dynamic spectral entropy weights, combining with an auditory masking compensation factor, performing spectrum synthesis, and obtaining a mixed audio signal after IFFT transformation and overlap elimination; performing noise reduction processing on the mixed audio signal through multimodal feature extraction to obtain a noise-reduced audio signal; performing energy feature extraction on the noise-reduced audio signal to perform Bayesian decision optimization noise reduction judgment to obtain a noise-reduced audio signal; performing automatic gain processing on the noise-reduced audio signal to obtain an automatically gained audio signal; performing DAC conversion on the automatically gained audio signal to output a noise-reduced analog audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of airborne audio processing technology, and in particular to an airborne high-reliability intelligent voice call noise reduction method. Background Art

[0002] The Airborne Voice Communication System (AVC) is a crucial component of aircraft communications systems, playing a vital role in ensuring flight safety, improving operational efficiency, and enhancing the cabin service experience. The AVC system primarily facilitates voice communications within the aircraft crew, between the aircraft and ground control, between aircraft and other aircraft, and between the cockpit and the passenger cabin. With the advancement of communication technology, AVC systems are moving towards digitalization, networking, and intelligentization, with their technological evolution revolving around high reliability, low latency, strong compatibility, and wide coverage.

[0003] Airborne voice communication systems face significant challenges, including complex and high-intensity noise environments (such as engine noise, airflow, and alarms), distortion from multi-channel audio mixing, reduced speech intelligibility in low signal-to-noise ratios, and the low-latency requirements of emergency scenarios. These factors severely hinder their further development. Traditional solutions include noise filtering, audio mixing processing, and ZYNQ noise reduction, but each has its limitations. Traditional noise filtering provides insufficient noise suppression, making it difficult to eliminate non-stationary noise (such as sudden mechanical vibration) in real time. ZYNQ noise reduction requires high computing resources and suffers from poor stability. Traditional audio mixing processing can easily mask high-frequency speech (such as consonants) and weak signals (such as bass), resulting in a loss of speech detail and unreliable communication. Summary of the Invention

[0004] In order to solve the above technical problems, the embodiments of the present application propose an airborne high-reliability intelligent voice call noise reduction method. Through the hierarchical design of mixing, noise reduction, noise reduction, automatic gain, and software and hardware joint filtering, the airborne voice call system is realized to operate with high clarity, low latency, and high reliability in complex noise environments. The technical indicators meet the top standards of avionics and well meet the voice call needs of aircraft.

[0005] In order to achieve the above-mentioned purpose, an embodiment of the present application proposes an airborne high-reliability intelligent voice call noise reduction method, which includes: performing ADC conversion on the acquired analog audio signal to obtain a digital audio signal; framing and FFT transforming the digital audio signal, determining the main spectrum in each complex spectrum, optimizing the phase of each complex spectrum based on the main spectrum, and calculating the spectral entropy of each complex spectrum at the same time, assigning dynamic spectral entropy weights to each complex spectrum based on the spectral entropy, and then combining the auditory masking compensation factor to perform spectral synthesis on each complex spectrum after phase optimization, and obtaining the mixed audio signal after IFFT transformation and overlap elimination; pre-emphasis, framing and FFT transforming the mixed audio signal, and based on each frame signal The amplitude spectrum and phase spectrum of the noise reduction signal are extracted to obtain the multimodal features of each frame signal, and the multimodal features of each frame signal are input into the pre-trained noise reduction model for noise reduction processing to obtain the audio signal after noise reduction processing; feature extraction is performed on the audio signal after noise reduction processing to obtain energy features including subband energy and differential energy, and a Bayesian decision optimization is performed on the squelch decision based on the energy features through dynamic subband GMM modeling to obtain the audio signal after squelch processing; long-term static gain, medium-time dynamic compression and short-time transient protection are performed on the audio signal after squelch processing to achieve automatic gain processing to obtain the audio signal after automatic gain; DAC conversion is performed on the audio signal after automatic gain to output the squelch analog audio signal.

[0006] In order to achieve the above-mentioned purpose, the embodiment of the present application also proposes an airborne high-reliability intelligent voice call noise reduction system, which includes: a receiving and ADC conversion module for acquiring an analog audio signal, performing ADC conversion on the acquired analog audio signal to obtain a digital audio signal; a mixing module for framing and FFT transforming the digital audio signal, determining the main spectrum in each complex spectrum, performing phase optimization on each complex spectrum based on the main spectrum, and calculating the spectral entropy of each complex spectrum at the same time, assigning a dynamic spectral entropy weight to each complex spectrum based on the spectral entropy, and then combining the auditory masking compensation factor to perform spectral synthesis on each complex spectrum after phase optimization, and obtaining the mixed audio signal after IFFT transformation and overlap elimination; a noise reduction module for pre-emphasis, framing and FFT transforming the mixed audio signal, and based on each The amplitude spectrum and phase spectrum of the frame signal are used for feature extraction to obtain the multimodal features of each frame signal, and the multimodal features of each frame signal are input into the pre-trained noise reduction model for noise reduction processing to obtain the audio signal after noise reduction processing; the noise reduction module is used to extract features of the audio signal after noise reduction processing to obtain energy features including subband energy and differential energy, and through dynamic subband GMM modeling, the noise reduction decision is optimized based on the energy features by Bayesian decision to obtain the audio signal after noise reduction processing; the automatic gain module is used to perform long-term static gain, medium-term dynamic compression and short-term transient protection on the audio signal after noise reduction processing to achieve automatic gain processing, and obtain the audio signal after automatic gain; the DAC conversion and output module is used to perform DAC conversion on the audio signal after automatic gain, obtain the noise reduction analog audio signal and output it.

[0007] In order to achieve the above-mentioned purpose, an embodiment of the present application also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an onboard high-reliability intelligent voice call noise reduction method as described above.

[0008] In order to achieve the above-mentioned purpose, an embodiment of the present application further proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement an airborne high-reliability intelligent voice call noise reduction method as described above.

[0009] This application proposes a method for reducing noise in airborne high-reliability intelligent voice calls. Through the hierarchical design of mixing, noise reduction, squelch, and automatic gain, it effectively solves the problem of multi-channel signal distortion, achieves a balance between noise suppression and voice fidelity, and improves the robustness of noise judgment. The automatic gain design with three-level control and clipping protection ensures output stability, ultimately achieving the high-definition, low-latency, and high-reliability requirements of airborne calls under complex noise conditions, and its technical indicators meet the top standards of avionics.

[0010] Optionally, the acquired analog audio signal is road, is an integer greater than 1, and the digital audio signal obtained after ADC conversion is also Road, will The digital audio signal is recorded as , ;

[0011] The digital audio signal is framed and FFT transformed. The main spectrum is determined in each complex spectrum. The phase of each complex spectrum is optimized based on the main spectrum. The spectral entropy of each complex spectrum is calculated. Dynamic spectral entropy weights are assigned to each complex spectrum based on the spectral entropy. Then, the phase-optimized complex spectrum is synthesized in combination with the auditory masking compensation factor. After IFFT transformation and overlap elimination, the mixed audio signal is obtained, including:

[0012] for The digital audio signals are divided into frames according to the preset frame length and overlap rate, and then the FFT transformation is performed on each frame signal to obtain The complex number spectrum; where The corresponding complex spectrum is recorded as ;

[0013] Will The complex spectrum with the largest energy among the complex spectrums is taken as the main spectrum , respectively calculated The complex spectrum and The phase offset between them is calculated based on their respective phase offsets. The complex spectrum of the two paths is mixed non-uniformly in phase to align the phases and obtain the complex spectrum after phase optimization;

[0014] The first The complex spectrum after phase optimization is recorded as , It is expressed by the formula:

[0015] ;

[0016] ;

[0017] ;

[0018] in, express The phase, express The phase, express and The phase shift between Indicates the center frequency of the sensitive frequency band, Indicates the maximum frequency;

[0019] definition The spectral entropy is , , Frequency The critical frequency band centered at express Any frequency point in It is the normalized energy probability. The higher the spectral entropy is, the more complex the frequency band information is.

[0020] Based on the following formula, ,calculate Corresponding dynamic spectral entropy weight :

[0021] ;

[0022] in, is the preset spectral entropy threshold, is the preset slope factor;

[0023] Using the MPEG psychoacoustic model, calculate The masking threshold of the complex spectrum is The maximum value of the masking threshold of the complex number spectrum is taken as the global masking threshold; The masking threshold is denoted as , the global masking threshold is recorded as , ;

[0024] Based on the following formula, calculate Compensation factor corresponding to the complex spectrum:

[0025] ;

[0026] in, express The corresponding compensation factor;

[0027] The following formula is used to perform spectrum synthesis based on the complex spectrum after dynamic spectrum entropy weight, compensation factor and phase optimization to obtain the final spectrum synthesis result;

[0028] ;

[0029] in, Represents the final spectrum synthesis result;

[0030] The final spectrum synthesis result is subjected to IFFT transformation and window overlap elimination operations in sequence to obtain the mixed audio signal. .

[0031] Optionally, the mixed audio signal is pre-emphasized, framed, and FFT-transformed, and feature extraction is performed based on the amplitude spectrum and phase spectrum of each frame signal to obtain multimodal features of each frame signal, including:

[0032] The mixed audio signal is pre-processed including pre-emphasis, framing and FFT transformation to extract the amplitude spectrum and phase spectrum ;

[0033] based on and Perform multimodal feature extraction to obtain 34-dimensional multimodal features, including 16-dimensional MFCC (Mel-frequency cepstral coefficients), 8-dimensional first-order difference ΔMFCC, 8-dimensional second-order difference ΔΔMFCC, 1-dimensional logarithmic energy, and 1-dimensional phase value; where Δ represents the first-order difference and ΔΔ represents the second-order difference;

[0034] The multimodal features of each frame signal are input into the pre-trained noise reduction model for noise reduction processing to obtain the noise-reduced audio signal, including:

[0035] The multimodal features of each frame signal are input into a pre-trained denoising model consisting of a CNN spatial feature branch, an LSTM temporal feature branch, a fusion layer, and a fully connected layer;

[0036] The CNN spatial feature branch is used to capture the spectrogram spatial features in the multimodal features, and the LSTM temporal feature branch is used to model the temporal dependency features in the multimodal features.

[0037] The fusion layer is used to concatenate and fuse the spatial features and temporal dependency features of the spectrogram to obtain fused features, and then the fully connected layer is used to output the denoised audio signal based on the fused features.

[0038] Optionally, feature extraction is performed on the audio signal after noise reduction processing to obtain energy features including subband energy and differential energy. Dynamic subband GMM modeling is used to perform Bayesian decision optimization based on the energy features to obtain the audio signal after noise reduction processing, including:

[0039] The sound activation detection algorithm is used to determine whether the audio signal after noise reduction processing contains human voice. If it does not contain human voice, it needs to be muted. If it does contain human voice, it directly performs automatic gain processing;

[0040] When squelch processing is required, the audio signal after noise reduction is subjected to frame processing and feature extraction to obtain a 48-dimensional energy feature. The 48-dimensional energy feature includes a 24-dimensional Mel subband energy feature and a 24-dimensional differential energy feature.

[0041] The 48-dimensional energy features are divided into 6 groups according to the frequency band. Each group of energy features is independently modeled by GMM to obtain speech GMM and noise GMM. Each GMM contains four Gaussian components. The speech GMM corresponding to the group energy feature is recorded as , will The noise GMM corresponding to the group energy feature is recorded as , ;

[0042] Use the silent period data to train the noise GMM, and train the speech GMM based on the clean speech library to obtain the model parameters of the noise GMM and speech GMM;

[0043] The subband likelihood ratio of each set of energy features is calculated using the following formula:

[0044] ;

[0045] in, represents the subband characteristics, Indicates the Subband likelihood ratio of group energy features, represents a Gaussian distribution, 、 、 express For the The model parameters of Gaussian components, 、 、 express For the The model parameters of the Gaussian components;

[0046] The following formula is used to update the speech prior probability and noise prior probability of the current frame according to the squelch decision result of the previous frame:

[0047] ;

[0048] ;

[0049] in, Indicates the current frame, Indicates the previous frame, represents the speech prior probability of the current frame, represents the noise prior probability of the current frame, represents the prior probability of speech in the previous frame, is the preset smoothing factor, Indicates the squelch decision result of the previous frame, Indicates that the previous frame is a speech frame, Indicates that the previous frame is a noise frame;

[0050] The following formula is used to perform a squelch decision on the current frame based on the speech prior probability and noise prior probability of the current frame and the subband likelihood ratio of each set of energy features to obtain the squelch decision result of the current frame:

[0051] ;

[0052] ;

[0053] in, Indicates the squelch decision threshold that is dynamically adjusted with the signal-to-noise ratio. Indicates the initial squelch decision threshold, represents the threshold adjustment factor, represents the signal-to-noise ratio, Indicates the squelch decision result of the current frame. Indicates that the current frame is a speech frame. Indicates that the current frame is a noise frame;

[0054] The noise frame is attenuated in the frequency domain, and a 10ms fade-in and fade-out window is applied at the switching boundary between the speech frame and the noise frame to achieve a smooth transition. Finally, the audio signal after the squelch process is output.

[0055] Optionally, performing long-term static gain, medium-term dynamic compression, and short-term transient protection on the audio signal after the squelch processing to implement automatic gain processing, thereby obtaining an audio signal after automatic gain, including:

[0056] Calculate the long-term static gain based on the target level and the average energy of the audio signal after squelching.

[0057] Using a compressor to perform mid-time dynamic compression, determine mid-time dynamic gain, and determine the upper limit of dynamic gain;

[0058] When the instantaneous energy of the audio signal after squelch processing is greater than 4 times the average energy and the zero-crossing rate is greater than 3 times the sampling rate, short-term transient protection is performed to determine the short-term transient linear attenuation;

[0059] The total gain is calculated based on the long-term static gain, medium-term dynamic gain, dynamic gain upper limit, and short-term transient linear attenuation using the following formula:

[0060] ;

[0061] ;

[0062] ;

[0063] in, represents the long-term static gain, Indicates the target level, Represents the average energy of the audio signal after squelch processing, represents the mid-time dynamic gain, Indicates the upper limit of dynamic gain, Represents short-time transient linear attenuation, Indicates freezing adjustment time, represents the total gain;

[0064] The total gain is smoothed using the following formula:

[0065] ;

[0066] in, represents the total gain after smoothing, Indicates the sampling point;

[0067] use And anti-clipping processing mechanism, the audio signal after squelch processing Perform automatic gain processing to obtain the audio signal after automatic gain , It is expressed by the formula:

[0068] ;

[0069] ;

[0070] in, Represents a symbolic function.

[0071] Optionally, the audio signal after automatic gain is subjected to DAC conversion, and a noise reduction analog audio signal is output, including: performing DAC conversion on the audio signal after automatic gain to obtain a noise reduction analog audio signal, filtering the noise reduction analog audio signal before output, and finally outputting the filtered noise reduction analog audio signal; wherein, the pre-output filtering adopts a multi-order high-pass filtering and low-pass filtering cascade architecture, including two groups of high-pass filters and two groups of low-pass filters, the two groups of high-pass filters are respectively a 300Hz high-pass filter and a 290Hz high-pass filter, and the two groups of low-pass filters are respectively a 4KHz low-pass filter and a 6.48KHz low-pass filter.

[0072] Optionally, the analog audio signal obtained is a four-channel microphone input, which are respectively the first microphone, the second microphone, the third microphone and the fourth microphone, and the output noise reduction analog audio signal is a three-channel headphone output, which are respectively the first headphone, the second headphone and the third headphone; before obtaining the analog audio signal, first determine the current working mode, which is one of the normal mode, the emergency mode or the follower mode; if the current working mode is the normal mode, input selection is performed in the first microphone, the second microphone and the third microphone, and the analog audio signal of the connected microphone is received. After being converted into a digital audio signal by ADC, the airborne signal processing module performs mixing, noise reduction, squelch and automatic gain processing, and transmits the automatically gained audio signal to the host computer through the SPI protocol for confirmation, and the received audio signal is returned by the host computer through the SPI protocol Perform DAC conversion to obtain a noise reduction analog audio signal and output it through the first earphone; if the current working mode is emergency mode, the analog audio signals of the first microphone, the second microphone and the third microphone are received at the same time, converted into digital audio signals by ADC, and input into the onboard signal processing module for mixing, noise reduction, squelching and automatic gain processing, and the audio signal after automatic gain is DAC converted to obtain a noise reduction analog audio signal and output it through the second earphone; if the current working mode is follower mode, the analog audio signal of the fourth microphone is received, converted into a digital audio signal by ADC, and transmitted to the host computer through the SPI protocol, which performs mixing, noise reduction, squelching and automatic gain processing, and returns the audio signal after automatic gain processing through the SPI protocol, and then after DAC conversion, obtains a noise reduction analog audio signal and outputs it through the third earphone. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the following is a brief introduction to the drawings required for use in the embodiments of the present application or the description of the related technologies. Obviously, the following drawings are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings described here are only used to explain the present application and are not used to limit the present application.

[0074] Figure 1 This is a flow chart of an onboard high-reliability intelligent voice call noise reduction method provided in one embodiment of the present application;

[0075] Figure 2 is a schematic diagram of the principle of mixing processing provided in one embodiment of the present application;

[0076] Figure 3 is a structural diagram of a noise reduction model provided in one embodiment of the present application;

[0077] Figure 4 Schematic diagram of the structure of the CNN spatial feature branch provided in one embodiment of the present application;

[0078] Figure 5 This is a schematic diagram of the structure of the LSTM time series feature branch provided in one embodiment of the present application;

[0079] Figure 6 is a schematic diagram of the principle of noise reduction processing provided in one embodiment of the present application;

[0080] Figure 7 is a schematic diagram of the principle of automatic gain processing provided in one embodiment of the present application;

[0081] Figure 8 1 is a schematic diagram of the principle of pre-output filtering provided in one embodiment of the present application;

[0082] Figure 9 This is a comparison diagram of the original human voice, the human voice mixed with strong noise, and the noise reduction simulated audio provided in one embodiment of the present application;

[0083] Figure 10 This is a comparison chart of mixed audio of human voices in an onboard high-intensity noise background and noise-reduced simulated audio provided in one embodiment of the present application;

[0084] Figure 11 This is a comparison chart of recorded and muted analog audio in an airborne high-intensity noise environment provided in one embodiment of the present application;

[0085] Figure 12 is a structural diagram of a redundant design provided in one embodiment of the present application;

[0086] Figure 13 This is a structural diagram of an airborne high-reliability intelligent voice call noise reduction system provided in another embodiment of the present application;

[0087] Figure 14 It is a structural diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION

[0088] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will appreciate that in the embodiments of the present application, many technical details are proposed to enable the reader to better understand. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of the present application. The following embodiments can be combined with each other and referenced to each other under the premise of no contradiction.

[0089] An embodiment of the present application proposes an airborne high-reliability intelligent voice call noise reduction method. The implementation details of the airborne high-reliability intelligent voice call noise reduction method proposed in this embodiment are described in detail below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this solution.

[0090] The specific process of the airborne high-reliability intelligent voice call noise reduction method proposed in this embodiment can be as follows: Figure 1 As shown, including:

[0091] Step 11: Perform ADC conversion on the acquired analog audio signal to obtain a digital audio signal.

[0092] In the specific implementation, the object of noise reduction and squelch processing is the acquired analog audio signal (which can be collected through a microphone). The analog audio signal is not suitable for direct processing and needs to be converted into a digital audio signal through ADC before subsequent mixing, noise reduction, squelch, and automatic gain processing can be performed.

[0093] In one example, the analog audio signal obtained is road, is an integer greater than 1, and the digital audio signal obtained after ADC conversion is also Road, No. The digital audio signal is recorded as , . The sound amplitudes of the analog audio signals of the different channels are different, so they will be mixed first after ADC conversion.

[0094] Step 12: Frame the digital audio signal and perform FFT transformation, determine the main spectrum in each complex spectrum, optimize the phase of each complex spectrum based on the main spectrum, calculate the spectral entropy of each complex spectrum, assign dynamic spectral entropy weights to each complex spectrum based on the spectral entropy, and then perform spectrum synthesis on the phase-optimized complex spectrum in combination with the auditory masking compensation factor. After IFFT transformation and overlap elimination, the mixed audio signal is obtained.

[0095] In the specific implementation, The digital audio signals first need to be mixed. The digital audio signals are framed and FFT transformed. The main spectrum is determined in each complex spectrum. The phase of each complex spectrum is optimized based on the main spectrum. At the same time, the spectral entropy of each complex spectrum is calculated. Based on the spectral entropy, dynamic spectral entropy weights are calculated and assigned to each complex spectrum. Auditory masking compensation is then performed in combination with the auditory masking compensation factor. The spectrum of each complex spectrum after phase optimization is synthesized. After IFFT transformation and overlap elimination, the mixed audio signal is obtained.

[0096] In one example, the mixing process can be as follows Figure 2 As shown. The digital audio signal needs to be divided into frames according to the preset frame length and overlap rate (1024 frames, Hanning window, 50% overlap rate), and then the FFT transformation of each frame signal is performed to obtain The complex spectrum, The corresponding complex spectrum is recorded as , Indicates frequency.

[0097] Next, The complex spectrum with the largest energy among the complex spectrums is taken as the main spectrum , respectively calculated The complex spectrum and The phase offset between them is calculated based on their respective phase offsets. The complex spectrum of the two paths is mixed non-uniformly in phase to align the phases and obtain the complex spectrum after phase optimization.

[0098] The first The complex spectrum after phase optimization is recorded as , It is expressed by the formula:

[0099] ;

[0100] ;

[0101] ;

[0102] in, express The phase, express The phase, express and The phase shift between Indicates the center frequency of the sensitive frequency band (1KHz to 4KHz). Indicates the maximum frequency.

[0103] While performing phase optimization, it is also necessary to calculate the dynamic spectral entropy weight and auditory masking compensation factor. The spectral entropy is , , Frequency The critical frequency band centered at express Any frequency point in It is the normalized energy probability. The higher the spectral entropy, the more complex the frequency band information is and the more details need to be retained. The lower the entropy value, the more concentrated the energy is (such as pure tone).

[0104] Then, based on the following formula, ,calculate Corresponding dynamic spectral entropy weight :

[0105] ;

[0106] in, is the preset spectral entropy threshold (generally set to 0.8), For a preset slope factor (generally set to 5.0), the dynamic spectral entropy weight allocated to the high spectral entropy region approaches 1, and the dynamic spectral entropy weight allocated to the low spectral entropy region approaches 0.

[0107] While calculating the dynamic spectral entropy weights, the MPEG psychoacoustic model is used to calculate The masking threshold of the complex spectrum is The maximum value of the masking threshold of the complex spectrum is used as the global masking threshold. The masking threshold is denoted as , the global masking threshold is recorded as , .

[0108] Then, based on the following formula, calculate Compensation factor corresponding to the complex spectrum:

[0109] ;

[0110] in, express The corresponding compensation factor is improved after being compensated by the auditory masking compensation factor.

[0111] At this point, the complex spectrum after dynamic spectral entropy weights, compensation factors, and phase optimization is ready. We will use the following formula to perform spectrum synthesis based on the complex spectrum after dynamic spectral entropy weights, compensation factors, and phase optimization to obtain the final spectrum synthesis result.

[0112] ;

[0113] in, Represents the final spectrum synthesis result, Indicates the total number of complex spectrum paths, express The corresponding dynamic spectral entropy weight, express The corresponding compensation factor, Indicates the Complex spectrum after phase optimization.

[0114] Finally, the final spectrum synthesis result is subjected to IFFT transformation and window overlap elimination operations in sequence to obtain the mixed audio signal. .

[0115] The mixing processing proposed in this embodiment assigns a larger dynamic spectral entropy weight to the high spectral entropy area, thereby achieving detail preservation. For masked weak signals, it is compensated by the auditory masking compensation factor, effectively avoiding the occurrence of distortion problems. The phase optimization design avoids the problem of monophony and has good compatibility.

[0116] Step 13: Pre-emphasize, frame, and perform FFT transformation on the mixed audio signal, perform feature extraction based on the amplitude spectrum and phase spectrum of each frame signal to obtain the multimodal features of each frame signal, input the multimodal features of each frame signal into a pre-trained noise reduction model for noise reduction processing, and obtain the audio signal after noise reduction processing.

[0117] In the specific implementation, after completing the mixing processing, the noise reduction processing stage can be entered. During the noise reduction processing, the audio signal after the mixing processing is pre-emphasized, framed and FFT transformed, and feature extraction is performed based on the amplitude spectrum and phase spectrum of each frame signal to obtain the multimodal features of each frame signal. The multimodal features of each frame signal are input into the pre-trained noise reduction model for noise reduction processing to obtain the audio signal after noise reduction processing.

[0118] In one example, when performing noise reduction processing, the mixed audio signal is first pre-processed including pre-emphasis (implemented by a first-order filter), framing (frame length 25ms, frame shift 10ms, Hamming window) and FFT transformation to extract the amplitude spectrum. and phase spectrum , then based on and Multimodal feature extraction is performed to obtain 34-dimensional multimodal features, including 16-dimensional MFCC, 8-dimensional first-order difference ΔMFCC, 8-dimensional second-order difference ΔΔMFCC, 1-dimensional logarithmic energy and 1-dimensional phase value; where Δ represents the first-order difference and ΔΔ represents the second-order difference.

[0119] Next, the multimodal features of each frame signal are input into the denoising model composed of CNN spatial feature branch, LSTM temporal feature branch, fusion layer and full connection layer for denoising. The specific structure of the denoising model can be as follows: Figure 3 shown.

[0120] The denoising model captures the spectrogram spatial features in the multimodal features through the CNN spatial feature branch, models the temporal dependency features in the multimodal features through the LSTM temporal feature branch, and then performs feature splicing and fusion of the spectrogram spatial features and temporal dependency features through the fusion layer to obtain the fused features. Finally, the fully connected layer outputs the denoised audio signal based on the fused features. The denoised audio signal is recorded as .

[0121] In an example, the specific structure of the CNN spatial feature branch is as follows Figure 4 As shown in Figure 2, the CNN spatial feature branch consists of convolution layer 1, ReLU activation function, maximum pooling, convolution layer 2, ReLU activation function, and global average pooling. The structure of the LSTM temporal feature branch is shown in Figure 2. Figure 5 As shown in Figure 1, the LSTM time series feature branch consists of a bidirectional LSTM unit, a unidirectional LSTM unit, and time pooling. The parameter configuration of the denoising model is shown in Table 1, all in px.

[0122] Table 1: Parameter configuration and output dimensions in the denoising model

[0123]

[0124] In one example, when training a noise reduction model, a clean speech library (TIMIT library) and environmental noise were used as data sets. The signal-to-noise ratio range was set between -5dB and 20dB to cover non-stationary noise. The Adam optimizer was used as the optimizer, the initial learning rate was set to 0.001, the attenuation factor was set to 0.9, and the mean square error loss was used as the loss function to measure the difference in the amplitude spectrum between the clean speech and the noise-reduced speech.

[0125] In one example, the CNN spatial feature branch uses DSP SIMD instructions to parallelize convolution calculations, and the LSTM temporal feature branch expands the time step loop and hardware accelerates matrix multiplication and addition.

[0126] In step 14, feature extraction is performed on the audio signal after noise reduction processing to obtain energy features including subband energy and differential energy. Dynamic subband GMM modeling is used to perform Bayesian decision optimization based on the energy features to obtain the audio signal after noise reduction processing.

[0127] In specific implementations, squelch processing is required after noise reduction. Before squelch processing, a sound activation detection algorithm is used to determine whether the noise-reduced audio signal contains human voices. If no human voices are present, squelch processing is performed. This involves feature extraction of the noise-reduced audio signal to obtain energy features including subband energy and differential energy. Dynamic subband GMM modeling is then used to perform Bayesian decision optimization based on the energy features to determine the squelch decision. If human voices are present, automatic gain processing is performed directly.

[0128] In one example, the principle of squelch processing is as follows Figure 6 As shown in the figure. If squelch processing is determined to be necessary, the denoised audio signal is framed (with a frame length of 20ms and a frame shift of 10ms) and feature extracted to obtain a 48-dimensional energy feature. The 48-dimensional energy feature consists of a 24-dimensional Mel subband energy feature and a 24-dimensional differential energy feature. The differential energy feature is the difference between the subband energy feature of the current frame and the subband energy feature of the previous frame.

[0129] After obtaining the 48-dimensional energy features, the 48-dimensional energy features are divided into 6 groups according to the frequency band. Each group of energy features is independently modeled using a Gaussian mixture model (GMM) to obtain speech GMM and noise GMM. Each GMM contains four Gaussian components (this choice can balance computational complexity and accuracy).

[0130] The first The speech GMM corresponding to the group energy feature is recorded as , will The noise GMM corresponding to the group energy feature is recorded as , .

[0131] The noise GMM is trained using the data from the silence period (the first 0.5s), and the speech GMM is trained based on the pure speech library (TIMIT) to obtain the model parameters of the noise GMM and speech GMM.

[0132] Next, the subband likelihood ratio of each set of energy features is calculated using the following formula:

[0133] ;

[0134] in, represents the subband characteristics, Indicates the Subband likelihood ratio of group energy features, represents a Gaussian distribution, 、 、 express For the The model parameters of the Gaussian components (i.e. weights, means, covariances), 、 、 express For the The model parameters of the Gaussian components.

[0135] Then, the following formula is used to update the speech prior probability and noise prior probability of the current frame according to the squelch decision result of the previous frame:

[0136] ;

[0137] ;

[0138] in, Indicates the current frame, Indicates the previous frame, represents the speech prior probability of the current frame, represents the noise prior probability of the current frame, represents the prior probability of speech in the previous frame, is the preset smoothing factor, Indicates the squelch decision result of the previous frame, Indicates that the previous frame is a speech frame, Indicates that the previous frame is a noise frame.

[0139] Then, based on the speech prior probability and noise prior probability of the current frame and the subband likelihood ratio of each set of energy features, the following formula is used to make a squelch decision for the current frame and obtain the squelch decision result for the current frame:

[0140] ;

[0141] ;

[0142] in, Indicates the squelch decision threshold that is dynamically adjusted with the signal-to-noise ratio. Indicates the initial squelch decision threshold, represents the threshold adjustment factor, Indicates the signal-to-noise ratio (given by the host computer through UART serial port communication), Indicates the squelch decision result of the current frame. Indicates that the current frame is a speech frame. Indicates that the current frame is a noise frame.

[0143] Finally, the noise frame is attenuated in the frequency domain (-20dB gain), and a 10ms fade-in and fade-out window is applied at the switching boundary between the speech frame and the noise frame to achieve a smooth transition. Finally, the audio signal after the squelch is output, which effectively avoids the occurrence of "clicking" sounds. The audio signal after the squelch is recorded as , after sampling, it can be recorded as .

[0144] In one example, the GMM model needs to be adaptively updated.

[0145] Noise GMM updates are divided into fast and slow updates. Fast updates immediately update the mean and variance of the corresponding subband GMM for frames judged as noise. Slow updates re-estimate the GMM weights and covariance every 10 seconds.

[0146] The update of the speech GMM is a protected update and is only performed when there are 100 consecutive high-confidence speech frames to prevent it from being contaminated by noise.

[0147] This embodiment proposes equal noise processing, dividing the spectrum into non-uniform subbands and independently modeling each subband with a GMM model. This effectively addresses the sensitivity of traditional full-band GMM modeling to local noise. This embodiment also introduces prior probabilities to dynamically adjust the speech / noise decision boundary, reducing the false positive rate at low signal-to-noise ratios. This embodiment also designs a dual-rate update mechanism based on noise non-stationarity, achieving an optimal balance between tracking speed and stability.

[0148] Step 15: Perform long-term static gain, medium-term dynamic compression, and short-term transient protection on the audio signal after the squelch process to achieve automatic gain processing, thereby obtaining an audio signal after automatic gain.

[0149] In the specific implementation, in order to prevent the sound from fluctuating, it is necessary to perform long-term static gain, medium-term dynamic compression and short-term transient protection on the audio signal after squelch processing (or the audio signal after noise reduction processing) to achieve automatic gain processing and obtain the audio signal after automatic gain.

[0150] In one example, the automatic gain processing process can be as follows Figure 7 shown.

[0151] First, a long-term static gain is calculated based on the target level (conversation comfort level, -26 dBFS) and the average energy of the squelched audio signal. The long-term static gain is updated only during noisy segments to avoid semantic distortion.

[0152] Subsequently, a compressor is used to perform mid-time dynamic compression, determine the mid-time dynamic gain, and determine the upper limit of the dynamic gain of the speech segment and the upper limit of the dynamic gain of the noise segment. The upper limit of the dynamic gain of the speech segment is 15dB, and the upper limit of the dynamic gain of the noise segment is 6dB.

[0153] The compressor has an attack time of 20ms to prevent sudden gain changes, a release time of 500ms to achieve natural decay, a compression ratio of 2 to 1, and a compression threshold of -30dB. No compression is performed below -30dB.

[0154] Next, when the instantaneous energy of the audio signal after the squelch process is greater than 4 times the average energy and the zero-crossing rate is greater than 3 times the sampling rate, short-time transient protection is performed to determine short-time transient linear attenuation.

[0155] The total gain is then calculated based on the long-term static gain, medium-term dynamic gain, dynamic gain ceiling, and short-term transient linear attenuation using the following formula:

[0156] ;

[0157] ;

[0158] ;

[0159] in, represents the long-term static gain, Indicates the target level, Represents the average energy of the audio signal after squelch processing, represents the mid-time dynamic gain, Indicates the upper limit of dynamic gain, Represents short-time transient linear attenuation, Indicates freezing adjustment time, Indicates the total gain.

[0160] In getting Then, through the following formula, To perform smoothing:

[0161] ;

[0162] in, represents the total gain after smoothing, Indicates the sampling point.

[0163] Finally, use And anti-clipping processing mechanism, the audio signal after squelch processing Perform automatic gain processing to obtain the audio signal after automatic gain , It is expressed by the formula:

[0164] ;

[0165] ;

[0166] in, Represents a symbolic function.

[0167] The automatic gain processing proposed in this embodiment utilizes a three-layer gain architecture: long-term static gain, medium-term dynamic compression, and short-term transient protection. This architecture prevents signal distortion. During gain control, the upper limit of gain is dynamically adjusted based on human hearing characteristics to minimize noise audibility. The short-term transient protection design identifies impact signals and freezes gain adjustments, eliminating "pops."

[0168] Step 16: Perform DAC conversion on the automatically amplified audio signal and output a noise-reduced analog audio signal.

[0169] In a specific implementation, after the automatic gain processing is completed, the audio signal after the automatic gain can be subjected to DAC conversion to obtain a noise-reduced analog audio signal and output it.

[0170] In one example, after performing DAC conversion on the automatically-gained audio signal to obtain a noise-reduced analog audio signal, it is necessary to perform pre-output filtering on the noise-reduced analog audio signal, and finally output the filtered noise-reduced analog audio signal. Figure 8 The multi-order cascaded high-pass and low-pass filtering architecture shown here consists of two sets of high-pass filters (300Hz and 290Hz), and two sets of low-pass filters (4kHz and 6.48kHz). By precisely setting the cutoff frequencies of the filter circuits of different orders, a highly targeted audio signal "filtering" channel can be constructed. This effectively isolates the target audio frequency band (100Hz to 4.5kHz) while simultaneously filtering out non-target noise, such as low-frequency environmental noise (such as equipment hum and low-frequency interference from ambient vibration) and high-frequency clutter (such as high-frequency noise from electromagnetic radiation), ensuring audio signal purity. The AIP4558S8 operational amplifier is used in the filter configuration. Its low noise and high gain-bandwidth product make it ideal for audio signal amplification and filtering. It balances signal transmission efficiency with amplitude stability, avoiding signal attenuation and distortion, and ensuring that the audio signal maintains excellent sound quality after filtering.

[0171] In one example, the cascade processing effects of mixing, noise reduction, squelch, and automatic gain can be as follows: Figure 9 、 Figure 10 、 Figure 11 As shown, the cascade processing architecture designed in this embodiment greatly improves the call quality and suppresses strong noise in complex environments.

[0172] In one example, Figure 12 As shown, the acquired analog audio signal is input by four microphones, namely the first microphone, the second microphone, the third microphone and the fourth microphone, and the output noise reduction analog audio signal is output by three headphones, namely the first headphone, the second headphone and the third headphone.

[0173] Before acquiring the analog audio signal, it is first necessary to determine the current working mode, which is one of the normal mode, the emergency mode, or the follower mode.

[0174] If the current working mode is normal mode, input selection is performed in the first microphone, the second microphone and the third microphone, and the analog audio signal of the connected microphone is received. After being converted into a digital audio signal by ADC, the onboard signal processing module performs mixing, noise reduction, squelch and automatic gain processing, and the audio signal after automatic gain is transmitted to the host computer for confirmation through the SPI protocol. The audio signal received from the host computer through the SPI protocol is converted by DAC to obtain a noise-reduced analog audio signal and output through the first earphone.

[0175] If the current working mode is emergency mode, the analog audio signals of the first microphone, the second microphone and the third microphone are received at the same time, converted into digital audio signals by ADC, and input into the onboard signal processing module for mixing, noise reduction, squelch and automatic gain processing. The audio signal after automatic gain is converted by DAC to obtain a noise-reduced and squelched analog audio signal and output through the second earphone.

[0176] If the current working mode is follower mode, the analog audio signal of the fourth microphone is received, converted into a digital audio signal by ADC, and transmitted to the host computer through the SPI protocol. The host computer performs mixing, noise reduction, squelch and automatic gain processing, and returns the audio signal after automatic gain processing through the SPI protocol. After conversion by DAC, the squelch analog audio signal is obtained and output through the third earphone.

[0177] Such redundant design greatly improves the reliability of the onboard voice communication system.

[0178] This embodiment proposes a high-reliability, intelligent airborne voice call noise reduction method. Through a hierarchical design of mixing, noise reduction, squelching, and automatic gain, it effectively solves the problem of multi-channel signal distortion, achieves a balance between noise suppression and voice fidelity, and improves the robustness of noise judgment. The automatic gain design with three-level control and clipping protection ensures output stability, ultimately achieving the high clarity, low latency, and high reliability requirements of airborne calls in complex noise environments, with technical indicators meeting top avionics standards.

[0179] The steps of the various methods above are divided for clarity of description only. During implementation, they can be combined into a single step, or some steps can be broken down into multiple steps. As long as they contain the same logical relationships, they are all within the scope of protection of this application. Adding minor modifications or introducing minor design changes to the algorithm or process, but not changing the core design of the algorithm or process, are also within the scope of protection of this application.

[0180] Another embodiment of the present application proposes an airborne high-reliability intelligent voice call noise reduction system. The details of the airborne high-reliability intelligent voice call noise reduction system proposed in this embodiment are described in detail below. The following content is only the implementation details provided for easy understanding and is not necessary for the implementation of this example. Figure 13 This is a structural diagram of an airborne high-reliability intelligent voice call noise reduction system proposed in this embodiment, including: a receiving and ADC conversion module 21, a mixing module 22, a noise reduction module 23, a noise reduction module 24, an automatic gain module 25 and a DAC conversion and output module 26.

[0181] The receiving and ADC conversion module 21 is used to obtain an analog audio signal, and perform ADC conversion on the obtained analog audio signal to obtain a digital audio signal.

[0182] The mixing module 22 is used to frame and perform FFT transformation on the digital audio signal, determine the main spectrum in each complex spectrum, optimize the phase of each complex spectrum based on the main spectrum, and simultaneously calculate the spectral entropy of each complex spectrum. Based on the spectral entropy, dynamic spectral entropy weights are assigned to each complex spectrum. Then, combined with the auditory masking compensation factor, the complex spectrum after phase optimization is synthesized. After IFFT transformation and overlap elimination, the audio signal after mixing is obtained.

[0183] The noise reduction module 23 is used to pre-emphasize, frame and perform FFT transformation on the audio signal after mixing processing, extract features based on the amplitude spectrum and phase spectrum of each frame signal to obtain the multimodal features of each frame signal, input the multimodal features of each frame signal into the pre-trained noise reduction model for noise reduction processing, and obtain the audio signal after noise reduction processing.

[0184] The squelch module 24 is used to extract features from the audio signal after noise reduction processing to obtain energy features including subband energy and differential energy. By using dynamic subband GMM modeling, a squelch decision is made based on the energy features using Bayesian decision optimization to obtain the squelched audio signal.

[0185] The automatic gain module 25 is used to perform long-term static gain, medium-term dynamic compression and short-term transient protection on the audio signal after the squelch process to achieve automatic gain processing and obtain an audio signal after automatic gain.

[0186] The DAC conversion and output module 26 is used to perform DAC conversion on the audio signal after automatic gain, obtain a noise-reduced analog audio signal, and output it.

[0187] It is not difficult to find that this embodiment is a system embodiment corresponding to the above-mentioned method embodiment, and this embodiment can be implemented in conjunction with the above-mentioned method embodiment. The relevant technical details and technical effects mentioned in the above-mentioned method embodiment are still valid in this embodiment, and to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above-mentioned method embodiment.

[0188] It is worth mentioning that all modules and modules involved in this embodiment are logical modules. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovation of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that other units do not exist in this embodiment.

[0189] Another embodiment of the present application provides an electronic device, such as Figure 14 As shown, it includes: at least one processor 31; and a memory 32 communicatively connected to the at least one processor 31; wherein the memory 32 stores instructions that can be executed by the at least one processor 31, and the instructions are executed by the at least one processor 31 to enable the at least one processor 31 to execute an onboard high-reliability intelligent voice call noise reduction method as described in the above method embodiment.

[0190] The memory and processor are connected using a bus, which includes any number of interconnected buses and bridges. The bus connects various circuits of one or more processors and memories. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and therefore will not be described further in this article. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. Data processed by the processor is transmitted on a wireless medium via an antenna. Furthermore, the antenna also receives data and transmits it to the processor.

[0191] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.

[0192] Another embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement an onboard high-reliability intelligent voice call noise reduction method as described in the above method embodiment.

[0193] That is, those skilled in the art will understand that all or part of the steps in the above-described method embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (such as a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps in the method embodiments of the present application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk.

[0194] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application. In actual applications, various modifications may be made to the embodiments in form and detail without departing from the spirit and scope of the present application. Those skilled in the art will appreciate that improvements and modifications may be made without departing from the principles of the present application, and such improvements and modifications are also considered to be within the scope of protection of the present application.

Claims

1. An airborne high-reliability intelligent voice call noise reduction method, characterized in that: include: Perform ADC conversion on the acquired analog audio signal to obtain a digital audio signal; The digital audio signal is framed and FFT-transformed, the main spectrum is determined in each complex spectrum, and the phase of each complex spectrum is optimized based on the main spectrum. The spectral entropy of each complex spectrum is calculated, and dynamic spectral entropy weights are assigned to each complex spectrum based on the spectral entropy. Then, combined with the auditory masking compensation factor, the complex spectra after phase optimization are synthesized. After IFFT transformation and overlap elimination, the mixed audio signal is obtained; The mixed audio signal is pre-emphasized, framed, and subjected to FFT transformation. Feature extraction is performed based on the amplitude spectrum and phase spectrum of each frame signal to obtain the multimodal features of each frame signal. The multimodal features of each frame signal are input into a pre-trained noise reduction model for noise reduction, thereby obtaining the noise-reduced audio signal. Feature extraction is performed on the denoised audio signal to obtain energy features including subband energy and differential energy. Dynamic subband GMM modeling is used to perform Bayesian decision optimization based on the energy features to obtain the denoised audio signal. Performing long-term static gain, medium-term dynamic compression, and short-term transient protection on the audio signal after squelch processing to achieve automatic gain processing and obtain an audio signal after automatic gain; Perform DAC conversion on the audio signal after automatic gain and output the noise-reduced analog audio signal; The audio signal after squelch processing is subjected to long-term static gain, medium-term dynamic compression, and short-term transient protection to achieve automatic gain processing, and the obtained audio signal after automatic gain includes: Calculate the long-term static gain based on the target level and the average energy of the audio signal after squelching. Using a compressor to perform mid-time dynamic compression, determine mid-time dynamic gain, and determine the upper limit of dynamic gain; When the instantaneous energy of the audio signal after squelch processing is greater than 4 times the average energy and the zero-crossing rate is greater than 3 times the sampling rate, short-term transient protection is performed to determine the short-term transient linear attenuation; The total gain is calculated based on the long-term static gain, medium-term dynamic gain, dynamic gain upper limit, and short-term transient linear attenuation using the following formula: ; ; ; in, represents the long-term static gain, Indicates the target level, Represents the average energy of the audio signal after squelch processing, represents the mid-time dynamic gain, Indicates the upper limit of dynamic gain, Represents short-time transient linear attenuation, Indicates freezing adjustment time, represents the total gain; The total gain is smoothed using the following formula: ; in, represents the total gain after smoothing, Indicates the sampling point; use And anti-clipping processing mechanism, the audio signal after squelch processing Perform automatic gain processing to obtain the audio signal after automatic gain , It is expressed by the formula: ; ; in, Represents a symbolic function.

2. The method for reducing noise during airborne high-reliability intelligent voice calls according to claim 1, characterized in that: The analog audio signal obtained is road, is an integer greater than 1, and the digital audio signal obtained after ADC conversion is also Road, will The digital audio signal is recorded as , ; The digital audio signal is framed and FFT transformed. The main spectrum is determined in each complex spectrum. The phase of each complex spectrum is optimized based on the main spectrum. The spectral entropy of each complex spectrum is calculated. Dynamic spectral entropy weights are assigned to each complex spectrum based on the spectral entropy. Then, the phase-optimized complex spectrum is synthesized in combination with the auditory masking compensation factor. After IFFT transformation and overlap elimination, the mixed audio signal is obtained, including: for The digital audio signals are divided into frames according to the preset frame length and overlap rate, and then the FFT transformation is performed on each frame signal to obtain The complex number spectrum; where The corresponding complex spectrum is recorded as ; Will The complex spectrum with the largest energy among the complex spectrums is taken as the main spectrum , respectively calculated The complex spectrum and The phase offset between them is calculated based on their respective phase offsets. The complex spectrum of the two paths is mixed non-uniformly in phase to align the phases and obtain the complex spectrum after phase optimization; The first The complex spectrum after phase optimization is recorded as , It is expressed by the formula: ; ; ; in, express The phase, express The phase, express and The phase shift between Indicates the center frequency of the sensitive frequency band, Indicates the maximum frequency; definition The spectral entropy is , , Frequency The critical frequency band centered at express Any frequency point in It is the normalized energy probability. The higher the spectral entropy is, the more complex the frequency band information is. Based on the following formula, ,calculate Corresponding dynamic spectral entropy weight : ; in, is the preset spectral entropy threshold, is the preset slope factor; Using the MPEG psychoacoustic model, calculate The masking threshold of the complex spectrum is The maximum value of the masking threshold of the complex number spectrum is taken as the global masking threshold; The masking threshold is denoted as , the global masking threshold is recorded as , ; Based on the following formula, calculate Compensation factor corresponding to the complex spectrum: ; in, express The corresponding compensation factor; The following formula is used to perform spectrum synthesis based on the complex spectrum after dynamic spectrum entropy weight, compensation factor and phase optimization to obtain the final spectrum synthesis result; ; in, Represents the final spectrum synthesis result; The final spectrum synthesis result is subjected to IFFT transformation and window overlap elimination operations in sequence to obtain the mixed audio signal. .

3. The method for reducing noise during airborne high-reliability intelligent voice calls according to claim 2, wherein: The mixed audio signal is pre-emphasized, framed, and FFT-transformed. Feature extraction is performed based on the amplitude spectrum and phase spectrum of each frame signal to obtain the multimodal features of each frame signal, including: The mixed audio signal is pre-processed including pre-emphasis, framing and FFT transformation to extract the amplitude spectrum and phase spectrum ; based on and Perform multimodal feature extraction to obtain 34-dimensional multimodal features, including 16-dimensional MFCC, 8-dimensional first-order difference ΔMFCC, 8-dimensional second-order difference ΔΔMFCC, 1-dimensional logarithmic energy and 1-dimensional phase value; where Δ represents the first-order difference and ΔΔ represents the second-order difference; The multimodal features of each frame signal are input into the pre-trained noise reduction model for noise reduction processing to obtain the noise-reduced audio signal, including: The multimodal features of each frame signal are input into a pre-trained denoising model consisting of a CNN spatial feature branch, an LSTM temporal feature branch, a fusion layer, and a fully connected layer; The CNN spatial feature branch is used to capture the spectrogram spatial features in the multimodal features, and the LSTM temporal feature branch is used to model the temporal dependency features in the multimodal features. The fusion layer is used to concatenate and fuse the spatial features and temporal dependency features of the spectrogram to obtain fused features, and then the fully connected layer is used to output the denoised audio signal based on the fused features.

4. The method for reducing noise during airborne high-reliability intelligent voice calls according to claim 3, wherein: Feature extraction is performed on the audio signal after noise reduction processing to obtain energy features including subband energy and differential energy. Dynamic subband GMM modeling is used to perform Bayesian decision optimization based on the energy features to obtain the audio signal after noise reduction processing, including: The sound activation detection algorithm is used to determine whether the audio signal after noise reduction processing contains human voice. If it does not contain human voice, it needs to be muted. If it does contain human voice, it directly performs automatic gain processing; When squelch processing is required, the audio signal after noise reduction is subjected to frame processing and feature extraction to obtain a 48-dimensional energy feature. The 48-dimensional energy feature includes a 24-dimensional Mel subband energy feature and a 24-dimensional differential energy feature. The 48-dimensional energy features are divided into 6 groups according to the frequency band. Each group of energy features is independently modeled by GMM to obtain speech GMM and noise GMM. Each GMM contains four Gaussian components. The speech GMM corresponding to the group energy feature is recorded as , will The noise GMM corresponding to the group energy feature is recorded as , ; Use the silent period data to train the noise GMM, and train the speech GMM based on the clean speech library to obtain the model parameters of the noise GMM and speech GMM; The subband likelihood ratio of each set of energy features is calculated using the following formula: ; in, represents the subband characteristics, Indicates the Subband likelihood ratio of group energy features, represents a Gaussian distribution, 、 、 express For the The model parameters of Gaussian components, 、 、 express For the The model parameters of the Gaussian components; The following formula is used to update the speech prior probability and noise prior probability of the current frame according to the squelch decision result of the previous frame: ; ; in, Indicates the current frame, Indicates the previous frame, represents the speech prior probability of the current frame, represents the noise prior probability of the current frame, represents the prior probability of speech in the previous frame, is the preset smoothing factor, Indicates the squelch decision result of the previous frame, Indicates that the previous frame is a speech frame, Indicates that the previous frame is a noise frame; The following formula is used to perform a squelch decision on the current frame based on the speech prior probability and noise prior probability of the current frame and the subband likelihood ratio of each set of energy features to obtain the squelch decision result of the current frame: ; ; in, Indicates the squelch decision threshold that is dynamically adjusted with the signal-to-noise ratio. Indicates the initial squelch decision threshold, represents the threshold adjustment factor, represents the signal-to-noise ratio, Indicates the squelch decision result of the current frame. Indicates that the current frame is a speech frame. Indicates that the current frame is a noise frame; The noise frame is attenuated in the frequency domain, and a 10ms fade-in and fade-out window is applied at the switching boundary between the speech frame and the noise frame to achieve a smooth transition. Finally, the audio signal after the squelch process is output.

5. An airborne high-reliability intelligent voice call noise reduction method according to any one of claims 1 to 4, characterized in that: Perform DAC conversion on the audio signal after automatic gain and output a noise-reduced analog audio signal, including: Performing DAC conversion on the audio signal after automatic gain to obtain a noise reduction analog audio signal, filtering the noise reduction analog audio signal before output, and finally outputting the filtered noise reduction analog audio signal; Among them, the pre-output filtering adopts a multi-order high-pass filter and low-pass filter cascade architecture, including two groups of high-pass filters and two groups of low-pass filters. The two groups of high-pass filters are a 300Hz high-pass filter and a 290Hz high-pass filter, and the two groups of low-pass filters are a 4KHz low-pass filter and a 6.48KHz low-pass filter.

6. An airborne high-reliability intelligent voice call noise reduction method according to any one of claims 1 to 4, characterized in that: The analog audio signal obtained is input by four microphones, namely the first microphone, the second microphone, the third microphone, and the fourth microphone. The analog audio signal with noise reduction is output by three headphone outputs, namely the first headphone, the second headphone, and the third headphone. Before acquiring the analog audio signal, first determine the current working mode, which is one of the normal mode, the emergency mode or the follower mode; If the current working mode is normal mode, input selection is performed in the first microphone, the second microphone, and the third microphone, and the analog audio signal of the connected microphone is received. After being converted into a digital audio signal by the ADC, the onboard signal processing module performs mixing, noise reduction, squelch, and automatic gain processing. The audio signal after automatic gain is transmitted to the host computer for confirmation via the SPI protocol. The audio signal received from the host computer via the SPI protocol is converted by DAC to obtain a noise-reduced analog audio signal and output through the first earphone; If the current working mode is emergency mode, the analog audio signals from the first microphone, the second microphone, and the third microphone are received simultaneously, converted into digital audio signals by ADC, and input into the onboard signal processing module for mixing, noise reduction, squelch, and automatic gain processing. The audio signal after automatic gain is converted by DAC to obtain a noise-reduced and squelched analog audio signal and output through the second earphone; If the current working mode is follower mode, the analog audio signal of the fourth microphone is received, converted into a digital audio signal by ADC, and transmitted to the host computer through the SPI protocol. The host computer performs mixing, noise reduction, squelch and automatic gain processing, and returns the audio signal after automatic gain processing through the SPI protocol. After conversion by DAC, the squelch analog audio signal is obtained and output through the third earphone.

7. An airborne high-reliability intelligent voice call noise reduction system, implemented based on an airborne high-reliability intelligent voice call noise reduction method according to any one of claims 1 to 6, characterized in that: include: The receiving and ADC conversion module is used to obtain analog audio signals, perform ADC conversion on the obtained analog audio signals, and obtain digital audio signals; A mixing module is used to perform framing and FFT transformation on the digital audio signal, determine the main spectrum in each complex spectrum, optimize the phase of each complex spectrum based on the main spectrum, calculate the spectral entropy of each complex spectrum, assign dynamic spectral entropy weights to each complex spectrum based on the spectral entropy, and then perform spectrum synthesis on each complex spectrum after phase optimization in combination with the auditory masking compensation factor. After IFFT transformation and overlap elimination, the mixed audio signal is obtained; The noise reduction module is used to pre-emphasize, frame, and perform FFT transformation on the mixed audio signal. It extracts features based on the amplitude spectrum and phase spectrum of each frame signal to obtain the multimodal features of each frame signal. The multimodal features of each frame signal are input into the pre-trained noise reduction model for noise reduction processing to obtain the noise-reduced audio signal. The squelch module is used to extract features from the audio signal after noise reduction processing, obtain energy features including subband energy and differential energy, and use dynamic subband GMM modeling to perform Bayesian decision optimization based on the energy features to obtain the squelch-processed audio signal; An automatic gain module is used to perform long-term static gain, medium-term dynamic compression, and short-term transient protection on the audio signal after squelch processing to achieve automatic gain processing and obtain an audio signal after automatic gain; The DAC conversion and output module is used to perform DAC conversion on the audio signal after automatic gain, obtain the noise-reduced analog audio signal and output it.

8. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an onboard high-reliability intelligent voice call noise reduction method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement an airborne high-reliability intelligent voice call noise reduction method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio system with high sensitivity and adjusting method thereof

    CN111163399A

  • Dynamic eq

    CN112384976A