Noise interference compensation method and system applied to audio transmission system

By extracting the time and frequency domain characteristics of noisy audio signals in the audio transmission system, dynamic feature fusion and training the noise compensation model, the problem of poor noise interference compensation in the prior art is solved, efficient and accurate noise interference compensation is achieved, and audio transmission quality and stability are improved.

CN120148538BActive Publication Date: 2025-08-01SHENZHEN ZIDOO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510610177.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-01
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The prior art noise interference compensation scheme in audio transmission systems fails to fully and accurately characterize the noise interference characteristics, and lacks effective utilization of noise-free audio signals and pure audio signals, resulting in poor noise suppression effect.

Method used

By obtaining the set of noise-free audio signals, extracting the time-domain and frequency-domain noise characteristics, performing dynamic feature fusion processing, training the noise compensation model to generate noise compensation parameters, and loading it in real time to the audio transmission system for noise interference compensation.

Benefits of technology

It achieves efficient and accurate compensation for noise interference during audio transmission, significantly improves audio transmission quality and stability, and provides a clearer audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148538B_ABST
    Figure CN120148538B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio transmission systems, and provides a method for compensating noise interference applied to an audio transmission system. First, a set of noisy audio signals of a target audio transmission channel is obtained, which includes multiple noisy audio segments and corresponding clean audio segments. Then, noise interference characteristics are extracted from the set of noisy audio signals to obtain the time-domain noise characteristics and frequency-domain noise characteristics of each noisy audio segment. Next, based on a preset noise compensation model, dynamic feature fusion is performed on the time-domain and frequency-domain noise characteristics to generate a set of noise interference characteristics, and the noise compensation model is trained according to the mapping relationship between the set of noise interference characteristics and the corresponding clean audio segments to generate a set of noise compensation parameters. Finally, the set of noise compensation parameters is loaded into a real-time processing module to perform noise interference compensation operations on the noisy audio signals transmitted in real time, thereby effectively reducing the impact of noise on the audio signals and improving the audio transmission quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio transmission optimization, and in particular to a noise interference compensation system applied to an audio transmission system. Background Art

[0002] In the technical field of audio transmission, noise interference has always been a key factor affecting the quality of audio transmission. With the continuous development of communication technologies, audio transmission systems have been widely used in various application scenarios, such as voice calls, video conferences, audio broadcasts, etc. However, during the actual transmission process, audio signals will inevitably be interfered by various noises. The above noises may come from the transmission environment, the device itself, or external electromagnetic interference, etc., resulting in a decline in the quality of the audio signals received at the receiving end, such as distortion, blurring, or even unrecognizability, seriously affecting the user's auditory experience and the accuracy of information transmission.

[0003] The noise interference compensation schemes of related technologies usually extract noise features only from one dimension of the time domain or the frequency domain, ignoring the mutual correlation and complementarity between the time-domain and frequency-domain features. The time-domain features mainly reflect the variation characteristics of audio signals over time, while the frequency-domain features reflect the distribution of audio signals in terms of frequency. Relying only on the features of a single dimension cannot comprehensively and accurately characterize the characteristics of noise interference, thus affecting the effect of noise suppression. During the noise compensation process, related technologies usually adopt fixed compensation strategies and do not dynamically adjust according to the actual noise situation of the audio signal. Different audio segments may be interfered by different degrees of noise, and their noise characteristics may also vary. Fixed compensation strategies cannot adapt to this change, resulting in poor compensation effects and being unable to effectively restore the quality of the original audio signal.

[0004] In addition, related technologies often lack the effective utilization of a large number of noisy audio signals and corresponding clean audio signals. In practical applications, although some noisy audio signals can be obtained, there are often no corresponding clean audio signals as references, which limits model training and optimization and makes it difficult to generate accurate and effective noise compensation parameters. Summary of the Invention

[0005] In view of the above, it is aimed to at least partially solve the deficiencies existing in the prior art and bring a new solution to the audio transmission field. In a first aspect, an embodiment of this application provides a method for compensating noise interference applied to an audio transmission system, and the method includes:

[0006] Obtain a set of noisy audio signals for a target audio transmission channel, where the set of noisy audio signals includes a plurality of noisy audio segments and corresponding clean audio segments for each noisy audio segment;

[0007] Perform noise interference feature extraction processing on the set of noisy audio signals to obtain the time-domain noise features and frequency-domain noise features of each noisy audio segment;

[0008] Based on a preset noise compensation model, perform dynamic feature fusion processing on the time-domain noise features and the frequency-domain noise features to generate a noise interference feature set of the noisy audio segment;

[0009] According to the mapping relationship between the noise interference feature set and the corresponding clean audio segment, train the noise compensation model to generate a set of noise compensation parameters;

[0010] Load the set of noise compensation parameters into the real-time processing module of the audio transmission system to perform noise interference compensation operations on the real-time transmitted noisy audio signals.

[0011] Second method, an embodiment of the present application further provides a noise interference compensation system applied to an audio transmission system, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the noise interference compensation method applied to the audio transmission system.

[0012] In summary, the noise interference compensation method and system provided by the embodiments of the present application for an audio transmission system achieve efficient and accurate compensation for noise interference during audio transmission, significantly improving the quality and stability of audio transmission. Specifically, first, a set of noisy audio signals of the target audio transmission channel is obtained. When performing noise interference feature extraction processing on the set of noisy audio signals, the time-domain noise features and frequency-domain noise features of each noisy audio segment can be respectively obtained, enabling a comprehensive characterization of the characteristics of noise interference from both the time domain and the frequency domain. Then, based on a preset noise compensation model, dynamic feature fusion processing is performed on the time-domain noise features and frequency-domain noise features to generate a set of noise interference features for the noisy audio segment, which can adaptively adjust the feature fusion strategy according to the noise characteristics of different audio segments, fully exploiting the correlation and complementarity between the time-domain and frequency-domain noise features, and improving the representation ability and accuracy of the set of noise interference features. Next, by training the noise compensation model according to the mapping relationship between the set of noise interference features and the corresponding clean audio segments to generate a set of noise compensation parameters, the noise compensation model can learn the complex mapping relationship between the noise interference features and the clean audio, thereby generating a targeted and effective set of noise compensation parameters. Finally, the set of noise compensation parameters is loaded into the real-time processing module of the audio transmission system to perform noise interference compensation operations on the real-time transmitted noisy audio signals, realizing real-time and dynamic compensation for noise interference during audio transmission, effectively reducing the impact of noise on the audio signal, and improving the clarity and intelligibility of the audio signal. Thus, the anti-noise interference ability of the audio transmission system is significantly improved, effectively enhancing the quality of audio transmission.

[0013] Other features and advantages of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on the above drawings without creative efforts.

[0015] To more fully understand the present application and its beneficial effects, the following will be described in conjunction with the drawings, where the same reference numerals in the following description represent the same parts.

[0016] Figure 1 It is a flowchart showing a noise interference compensation method for an audio transmission system provided by an embodiment of the present application.

[0017] Figure 2It is a schematic diagram of an audio transmission system provided by an embodiment of the present application.

[0018] Figure 3 It is a schematic diagram of a noise interference compensation system provided by an embodiment of the present application. Specific embodiments

[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application.

[0020] Please refer to Figure 1 and Figure 2 as shown, Figure 1 It is a schematic flowchart of a noise interference compensation method applied to an audio transmission system provided by an embodiment of the present application. Figure 2 It is a schematic architecture diagram of an audio transmission system. The audio transmission system may include at least two audio terminals for audio transmission, and a noise interference compensation system for compensating the noise interference of the audio signals transmitted by the at least two audio terminals. Among them, in this embodiment, the method may be implemented by the noise interference compensation system. For example, Figure 2 as shown, the audio terminal may be a device with data transmission, analysis, and processing capabilities such as a computer terminal, a mobile phone, a tablet computer, etc., and the noise interference compensation system may be a server, a server cluster, a computer device, etc. This embodiment does not specifically limit this.

[0021] As Figure 1 shown, the method includes steps S10 - S50, which will be introduced in detail below.

[0022] S10. Obtain a set of noisy audio signals of a target audio transmission channel, where the set of noisy audio signals includes a plurality of noisy audio segments and corresponding clean audio segments for each noisy audio segment.

[0023] In this application, the target audio transmission channel can be understood as a specific channel in the entire audio transmission system that requires noise interference compensation. Among them, the set of noisy audio signals includes multiple noisy audio segments and the corresponding clean audio segments. Specifically, the noisy audio segment can be the part of the audio signal that is interfered by noise during the actual audio transmission process. For example, in an audio transmission system for a voice call, if there is background noise in the call environment, such as the conversation sounds of people around, the running sounds of machines, etc., then the received voice signal is the noisy audio signal. Correspondingly, the clean audio segment can be the part of the original audio signal that is not interfered by any noise in the ideal state. For example, in actual situations, it may be obtained by pre-recording in a noise-free environment or by separating the noise signal from the known noisy audio signal through special signal processing means and then restoring it.

[0024] In the process of obtaining the set of noisy audio signals, a signal acquisition module can be set at the receiving end of the audio transmission system to collect the transmitted audio signals according to the set sampling frequency. When enough audio signals are collected, the set of noisy audio signals is then aggregated and generated.

[0025] Step S20: Perform noise interference feature extraction processing on the set of noisy audio signals to obtain the time-domain noise features and frequency-domain noise features of each noisy audio segment.

[0026] In this embodiment, after obtaining the set of noisy audio signals, it is necessary to perform noise interference feature extraction processing on it to obtain the time-domain noise features and frequency-domain noise features of each noisy audio segment.

[0027] Specifically, the time-domain noise features reflect the characteristics of the noise in the time dimension. For example, the noise may be sudden, with a large amplitude suddenly appearing at a certain moment and then disappearing quickly; for another example, it may also be continuous, maintaining a relatively stable amplitude within a certain period of time. By analyzing the noisy audio segment in the time domain, features such as the amplitude change trend and duration of the noise can be extracted. Exemplarily, a time-domain analysis algorithm can be used, such as calculating the amplitude difference of the noisy audio segment at different time points, calculating the average value and standard deviation of the amplitude within a certain period of time, etc., to obtain the time-domain noise features.

[0028] Furthermore, the frequency-domain noise features reflect the distribution characteristics of noise in the frequency domain. Specifically, in an audio signal, different sound components correspond to different frequency ranges, so noise also has its specific frequency distribution. For example, the noise generated by electrical equipment may be concentrated in a certain specific high-frequency band, while the wind noise may have a large energy distribution in the low-frequency band. Based on this, in order to obtain the frequency-domain noise features, frequency-domain analysis methods such as Fourier transform can be performed on the noisy audio segment. Through Fourier transform, the audio signal can be converted from the time domain to the frequency domain, so that the energy distribution of noise at different frequencies can be determined, such as the energy amplitude at a certain frequency point or frequency interval.

[0029] By extracting the above-mentioned time-domain and frequency-domain noise features from the set of noisy audio signals, the characteristics of noise can be comprehensively described from different dimensions, enabling subsequent noise compensation to more specifically handle different types of noise. For example, in the subsequent noise compensation process, both the burst noise in the time domain and the noise with a specific frequency distribution in the frequency domain can be effectively identified and processed, thereby improving the accuracy of noise compensation.

[0030] Step S30: Based on a preset noise compensation model, perform dynamic feature fusion processing on the time-domain noise features and the frequency-domain noise features to generate a noise interference feature set of the noisy audio segment.

[0031] In this application, based on a preset noise compensation model, dynamic feature fusion processing can be performed on the obtained time-domain noise features and frequency-domain noise features, thereby generating a noise interference feature set of the noisy audio segment.

[0032] Specifically, the preset noise compensation model can be a pre-constructed mathematical model that can perform corresponding calculations and processing based on the input feature information. For the dynamic feature fusion processing process, since the time-domain noise features and the frequency-domain noise features describe the characteristics of noise from different perspectives, in order to more comprehensively represent the interference of noise on the audio signal, the time-domain noise features and the frequency-domain noise features need to be fused.

[0033] For example, in an audio scenario, there may be a type of noise that appears periodic in the time domain and is concentrated at several specific frequencies in the frequency domain. Through dynamic feature fusion processing, the above-mentioned time-domain and frequency-domain information can be integrated. Among them, the specific fusion method can adopt the weighted summation method, and different weights are assigned according to the importance of different features. For example, for some frequency-domain features that have a greater impact on audio quality, higher weights can be given, while for relatively less important time-domain features, lower weights can be given. Among them, the determination of the weights can be pre-determined according to a large amount of experimental data or prior knowledge of the audio transmission system, and no detailed restrictions are imposed here. During the fusion process, the noise compensation model can generate a set of noise interference features containing comprehensive noise interference information according to the input time-domain and frequency-domain noise features and the set fusion rules. Each element in the set of noise interference features can represent the noise interference features of a noisy audio segment after comprehensively considering the time-domain and frequency-domain noise effects.

[0034] Through the above dynamic feature fusion processing process, the scattered time-domain and frequency-domain noise features can be integrated into a unified and more representative set of noise interference features, enabling the subsequent noise compensation model to be trained based on this more comprehensive set of noise interference features, thereby improving the effect of noise compensation and better restoring the pure audio signal.

[0035] Step S40, according to the mapping relationship between the set of noise interference features and the corresponding pure audio segments, train the noise compensation model to generate a set of noise compensation parameters.

[0036] In this embodiment, according to the mapping relationship between the set of noise interference features and the corresponding pure audio segments, the noise compensation model is trained to generate a set of noise compensation parameters. Among them, this mapping relationship reflects the internal connection between the noise interference features and the pure audio. For example, when an element in the set of noise interference features indicates that there is significant noise interference at a certain frequency in a noisy audio segment, the corresponding pure audio segment should have no such noise interference at that frequency. Through a large number of noisy audio segments and their corresponding pure audio segments, the above mapping relationship can be established. Training the noise compensation model can be to use the above mapping relationship to adjust the parameters of the model.

[0037] For example, during the training process, the noise compensation model can predict an audio signal close to the clean audio segment based on the input set of noise interference characteristics, and then compare the predicted audio signal with the actual clean audio segment. For example, error metrics such as the mean square error between the two can be calculated, and based on this error metric, its own parameters can be adjusted to make the next prediction more accurate. This adjustment process is an iterative process that will continue to repeat until a preset training stop condition is reached, such as the error metric being less than a certain set value or the number of training rounds reaching a certain quantity.

[0038] Through the above training process, the noise compensation model can generate a set of noise compensation parameters, which contains various parameter values determined after the model is trained, and which reflects how to compensate the noisy audio to obtain a signal close to the clean audio under different noise interference conditions. Based on this, by using the mapping relationship between the noise interference characteristics and the clean audio to train the noise compensation model, the noise compensation model can accurately learn how to compensate for noise interference, thereby generating an effective set of noise compensation parameters, providing a reliable basis for subsequent noise compensation operations in the actual audio transmission system, greatly improving the noise compensation ability in audio transmission, and enhancing the audio quality.

[0039] Step S50: Load the set of noise compensation parameters into the real-time processing module of the audio transmission system, and perform noise interference compensation operations on the noisy audio signal transmitted in real time.

[0040] In this application, the set of noise compensation parameters can be loaded into the real-time processing module of the audio transmission system to perform noise interference compensation operations on the noisy audio signal transmitted in real time. The real-time processing module of the audio transmission system is the part that processes the audio signal in real time during the audio transmission process, and it can quickly receive, process, and transmit the audio signal. When the set of noise compensation parameters is loaded into this real-time processing module, the real-time processing module can process the noisy audio signal transmitted in real time according to the above set of noise compensation parameters. For example, when the noisy audio signal transmitted in real time enters the real-time processing module, the real-time processing module can identify and quantify the noise components in the noisy audio signal according to the parameters in the set of noise compensation parameters. If the noise compensation parameters indicate that there is significant noise interference in a certain frequency range, then the real-time processing module can adjust the signal in that frequency range according to the corresponding parameters, such as attenuating the amplitude of the noise signal or enhancing the amplitude of the clean audio signal. This adjustment process is carried out in real time and will not have a great impact on the real-time nature of the audio transmission, because the real-time processing module has a high processing speed and efficiency and can quickly complete the noise interference compensation operation while the audio signal is being transmitted.

[0041] Through the above steps, this embodiment can effectively remove or reduce noise interference during the actual process of audio transmission, making the audio signal received at the receiving end closer to a pure audio signal, thereby improving the quality of audio transmission. Whether in application scenarios such as voice calls or audio playback, it can provide users with a clearer and higher-quality audio experience.

[0042] In a possible implementation manner, in step S20, when performing noise interference feature extraction processing on the set of noisy audio signals to obtain the time-domain noise features and frequency-domain noise features of each noisy audio segment, it can be achieved through the following steps S201 - S204, and the specific description is as follows.

[0043] Step S201: Perform frame division and windowing processing on the noisy audio segment to obtain multiple time-domain audio frames.

[0044] In this application, performing frame division and windowing processing on a noisy audio segment to obtain multiple time-domain audio frames can be to divide a continuous noisy audio segment into multiple small segments according to a certain time length, and each small segment can be a time-domain audio frame. Windowing processing is an operation performed on each time-domain audio frame, and its purpose is to reduce adverse effects such as spectral leakage caused by signal truncation due to frame division. For example, in an audio transmission system for speech recognition, the noisy audio may be the speaking voice mixed with background noise. When performing frame division and windowing processing on such a noisy audio segment, the time length of frame division can be determined according to the characteristics of the speech signal and the requirements of subsequent analysis. Generally, a time length that can reflect the short-term stationarity of the speech and does not cause information loss due to too long a frame length will be selected, such as 20 - 30 milliseconds. Window functions can be common window functions such as the Hanning window and Hamming window, and this embodiment does not limit this. Through the above frame division and windowing processing, the multiple time-domain audio frames obtained provide suitable data units for subsequent operations such as calculating short-time energy analysis and zero-crossing rate calculation. In this way, the subsequent extraction of time-domain noise features can be performed with smaller and more representative audio units, improving the accuracy and reliability of feature extraction, and further providing more accurate time-domain noise information for the entire noise interference compensation method.

[0045] Step S202: Perform short-time energy analysis and zero-crossing rate calculation on each time-domain audio frame to generate time-domain noise features, where the time-domain noise features include an energy fluctuation sequence and a zero-crossing rate distribution sequence.

[0046] In this embodiment, after obtaining multiple time-domain audio frames, short-time energy analysis and zero-crossing rate calculation can be performed on each time-domain audio frame to generate time-domain noise features. Among them, short-time energy analysis is a method for measuring the change of audio signal energy in a short period. For each time-domain audio frame, calculating its short-time energy can reflect the intensity of the audio signal within this frame. For example, in an audio transmission system for music playback, if the short-time energy of a certain frame in a noisy audio segment suddenly increases, it may mean that there is a sudden noise interference during this short period or the strong sound part in the music is affected by noise. By calculating the short-time energy of all time-domain audio frames, an energy fluctuation sequence can be obtained, and this energy fluctuation sequence can clearly show the fluctuation of the audio signal energy in the time dimension, which can be used as an important part of the time-domain noise feature.

[0047] Zero-crossing rate calculation is to count the number of times the audio signal crosses the zero axis within each time-domain audio frame. Therefore, the zero-crossing rate can reflect the frequency characteristics of the audio signal, and different types of audio signals have different zero-crossing rate distributions. For example, the zero-crossing rate of high-frequency signals is usually higher than that of low-frequency signals. In noisy audio, the presence of noise may change the zero-crossing rate distribution of the audio signal. By calculating the zero-crossing rate of each time-domain audio frame, a zero-crossing rate distribution sequence can be obtained. This zero-crossing rate distribution sequence and the energy fluctuation sequence together constitute the time-domain noise feature. In this way, by generating the time-domain noise feature through short-time energy analysis and zero-crossing rate calculation, the noise characteristics in the time domain can be accurately described from two aspects related to energy and frequency, providing rich information for subsequent comprehensive analysis of noise interference in noisy audio and helping to perform more accurate noise compensation.

[0048] Step S203: Perform fast Fourier transform processing on each time-domain audio frame to obtain frequency-domain audio frames, extract the spectral envelope feature and sub-band energy ratio feature of the frequency-domain audio frames, and generate frequency-domain noise features.

[0049] In this application, the fast Fourier transform can be performed on each time-domain audio frame to obtain a frequency-domain audio frame, and then the spectral envelope feature and sub-band energy ratio feature of the frequency-domain audio frame can be extracted to generate a frequency-domain noise feature. The fast Fourier transform is an algorithm that converts a time-domain signal into a frequency-domain signal. Specifically, in audio signal processing, it can convert each time-domain audio frame from the time domain to the frequency domain to obtain a frequency-domain audio frame. For example, in a broadcast audio transmission system, the noisy audio may contain various frequency noise interferences. The frequency-domain audio frame obtained through the fast Fourier transform can clearly show the energy distribution of the audio signal at different frequencies. The spectral envelope feature reflects the distribution and change trend of the main frequency components in the frequency-domain audio frame, and it can depict the curve of the main energy distribution contour of the audio signal in the frequency domain. For example, for a noisy audio containing music and noise, the main melody part of the music and the frequency distribution of the noise will be reflected in the spectral envelope feature. The sub-band energy ratio feature divides the frequency-domain audio frame into multiple sub-bands, and then calculates the proportion of the energy in each sub-band to the total energy. Different types of noise may have different energy distributions in different sub-bands. For example, the noise generated by electrical equipment may have a higher energy ratio in the high-frequency sub-band, while the low-frequency noise in the environment may have a higher energy ratio in the low-frequency sub-band. By extracting the spectral envelope feature and sub-band energy ratio feature, the noise characteristics in the frequency domain can be comprehensively described, thereby generating a frequency-domain noise feature.

[0050] In this way, the characteristics of the noise can be accurately characterized from the frequency-domain perspective, providing an important basis for subsequent noise compensation, especially playing a significant role in targeted compensation for noises of different frequencies.

[0051] Step S204: Perform time alignment processing on the time-domain noise feature and the frequency-domain noise feature to form a noise interference feature set corresponding to the noisy audio segment.

[0052] Among them, step S204 may be:

[0053] Step S2041: According to the original time stamp of the noisy audio segment, perform segmented marking on the energy fluctuation sequence and zero-crossing rate distribution sequence in the time-domain noise feature, and at the same time perform synchronous segmented marking on the spectral envelope feature and sub-band energy ratio feature in the frequency-domain noise feature.

[0054] For example, in an audio transmission system of a video conference, the noisy audio segment has its own time sequence. Performing segmented marking according to the original time stamp can ensure the corresponding relationship between the time-domain and frequency-domain noise features in time. Then, the frame offset between the time-domain noise feature and the frequency-domain noise feature is adjusted through the dynamic time warping algorithm to make the time-domain noise feature and the frequency-domain noise feature in the same time interval completely aligned in the time dimension.

[0055] Step S2042, adjust the frame offset between the time-domain noise feature and the frequency-domain noise feature through the dynamic time warping algorithm, so that the time-domain noise feature and the frequency-domain noise feature within the same time interval are completely aligned in the time dimension;

[0056] In this embodiment, the dynamic time warping algorithm can handle the possible time offset problem between the time-domain and frequency-domain features. For example, due to operations such as frame division and windowing, the time-domain and frequency-domain features may not be completely matched in time, and this dynamic time warping algorithm can effectively make adjustments.

[0057] Step S2043, splice the aligned time-domain noise feature and frequency-domain noise feature in chronological order into a multi-dimensional feature vector, which is used as a component of the noise interference feature set. [[ID=**10**]]

[0058] In this embodiment, the spliced multi-dimensional feature vector contains comprehensive noise information in the time domain and frequency domain. For example, in the audio transmission system of an audio editing software, such a noise interference feature set can provide comprehensive input information for the subsequent noise compensation model.

[0059] In this way, through time alignment processing, the noise features in the time domain and frequency domain can be effectively integrated, so that the noise interference feature set can accurately reflect the noise interference situation of the noisy audio segment at different times and frequencies, providing a more comprehensive and accurate basis for subsequent noise compensation operations, thereby improving the effect of noise compensation.

[0060] In a possible implementation manner, the above step S30 can be implemented through the following S301 - S304, and the specific description is as follows.

[0061] S301, call the residual convolutional network in the noise compensation model to perform multi-layer convolutional processing on the time-domain noise feature, and extract the deep time-domain features. The residual convolutional network includes a three-level convolutional structure, and each level of convolutional structure is composed of convolutional layers with decreasing convolutional kernel sizes, and a skip connection is introduced after the output of each level of convolutional layer to retain the low-order information of the original time-domain noise feature.

[0062] For example, in an audio enhancement system, for the input time-domain noise characteristics, the design with decreasing convolutional kernel size can be understood as a gradually focusing process. The large convolutional kernel can capture more macroscopic time-domain noise characteristic information in the initial stage. As the convolutional kernel size decreases, subsequent convolutional layers can gradually focus on more detailed features. And a skip connection is introduced after the output of each convolutional layer to retain the low-order information of the original time-domain noise characteristics. The skip connection mechanism is similar to retaining some basic building blocks when constructing a complex structure. For example, in an audio signal, the original time-domain noise characteristics may contain some simple but key low-order information, such as certain persistent weak noise fluctuation patterns. Through the skip connection, the above-mentioned low-order information will not be lost during the multi-layer convolution process. When performing multi-layer convolution processing, each layer of convolution will perform operations such as weighted summation on the input time-domain noise characteristics according to the characteristics of its convolutional kernel, thereby gradually extracting more complex and more abstract deep time-domain features. Thus, the complex information in the time-domain noise characteristics can be deeply mined, providing richer and deeper time-domain features for the subsequent fusion with frequency-domain features and the generation of the final noise interference feature set, which helps to improve the understanding and processing ability of time-domain noise.

[0063] S302. Call the frequency-domain attention network in the noise compensation model to perform frequency-band weight assignment processing on the frequency-domain noise characteristics and extract frequency-domain attention features. Among them, the frequency-domain attention network includes a sub-band division module and a weight calculation module. The sub-band division module divides the frequency-domain noise characteristics into multiple non-uniform sub-bands, and the weight calculation module generates frequency-band weight coefficients according to the energy ratio and spectral flatness of each sub-band.

[0064] In this embodiment, in a scenario for processing music audio, considering that the energy distribution of the audio signal in the frequency domain is not uniform, the sub-band division module divides the frequency-domain noise characteristics into multiple non-uniform sub-bands. For example, the sounds emitted by different musical instruments in music have their own characteristics in the frequency-domain distribution. The sound energy of some musical instruments is concentrated in the low-frequency band, while that of others is in the high-frequency band. The non-uniform sub-band division can better adapt to the above characteristics. The weight calculation module generates a frequency-band weight coefficient according to the energy proportion and spectral flatness of each sub-band. The energy proportion can reflect the relative importance of each sub-band in the entire frequency-domain noise characteristics. For example, if a certain sub-band has a large energy proportion, it indicates that the noise in this sub-band may have a greater impact on the audio quality. The spectral flatness can reflect the flat or fluctuating situation of the spectrum within the sub-band. A flat spectrum may indicate that the noise is evenly distributed, while a spectrum with large fluctuations may imply noise interference at specific frequencies. The frequency-band weight coefficient generated according to the above factors can enable the frequency-domain attention network to focus more specifically on important sub-bands in subsequent processing. In this way, by allocating different attention weights according to the characteristics of different sub-bands from the frequency-domain perspective, the important information in the frequency-domain noise characteristics can be captured more accurately, providing more targeted frequency-domain characteristics for subsequent cross-domain fusion, which helps to improve the analysis and processing efficiency of frequency-domain noise.

[0065] S303. Perform cross-domain splicing on the time-domain deep features and the frequency-domain attention features to obtain initial fusion features.

[0066] In this application, the matrix transpose operation can be first performed on the channel dimension of the time-domain deep features and the frequency-band dimension of the frequency-domain attention features. For example, the time-domain deep features and the frequency-domain attention features originally have specific structures in their respective dimensions. The matrix transpose operation makes it more convenient to splice the two in terms of dimensions. Then, considering that the resolutions of the time-domain deep features and the frequency-domain attention features on the time axis may be different, the interpolation algorithm can be used to adjust the resolutions of the two on the time axis. The interpolation algorithm can generate new data points based on the existing data points, so that the two have the same resolution on the time axis. In this way, the initial fusion features after splicing have a consistent data structure in the time and frequency-domain dimensions. Through the method of cross-domain splicing, the deep features and attention features extracted from the time domain and the frequency domain can be effectively integrated together to form an initial fusion feature containing comprehensive time-domain and frequency-domain information, providing a unified feature basis for subsequent further processing, which helps to more comprehensively describe the noise interference situation in the audio signal.

[0067] S304. Process the initial fusion features through a bidirectional long short-term memory network to capture the temporal dependence relationship and generate dynamic fusion features as the noise interference feature set.

[0068] In this embodiment, the bidirectional long short-term memory network is a neural network structure suitable for processing data with temporal information. As an example, the initial fusion feature contains comprehensive information in the time domain and frequency domain, but there may be complex temporal dependencies between the above information. For example, the noise interference in the audio signal may have different manifestations at different time points, and the noise characteristics at the previous and subsequent moments may be correlated with each other. The bidirectional long short-term memory network can process forward and backward temporal information simultaneously through its unique structure. It can remember the feature information at past moments and can make predictive associations with the features at future moments. When processing the initial fusion feature, each unit in the network processes the input feature according to its internal gating mechanism. For example, the input gate determines which information needs to be updated into the cell state, the forget gate determines which information needs to be forgotten, and the output gate determines which information can be output. In the above way, the bidirectional long short-term memory network can capture the temporal dependencies in the initial fusion feature, thereby generating a dynamic fusion feature. This dynamic fusion feature, as a set of noise interference features, can more comprehensively and accurately reflect the noise interference characteristics in the audio signal, including information such as the dynamic changes of noise in the time series. Based on this, by capturing the temporal dependencies, the generated set of noise interference features can better adapt to the dynamic change characteristics of the noise in the audio signal, provide more effective feature input for subsequent operations such as training the noise compensation model, and thus improve the performance of the entire noise compensation method.

[0069] Among them, for step S304, it may include the following steps S3041 - S3044, which are specifically described as follows.

[0070] S3041, input the initial fusion feature into the forward long short-term memory unit in chronological order to generate a forward hidden state sequence. Among them, the gating mechanism of the forward long short-term memory unit calculates the activation values of the input gate, forget gate, and output gate according to the initial fusion feature at the current time step and the hidden state at the previous time step to update the cell state and output the forward hidden state.

[0071] In this application, as an example, in an audio event analysis system, the initial fusion feature contains comprehensive information about noise integrated from the time domain and the frequency domain. When this feature is input into the forward long short-term memory unit in chronological order, at each time step, the gating mechanism starts to operate. For example, for the initial fusion feature at the current time step and the hidden state at the previous time step, the activation value of the input gate is calculated. The input gate determines how much of the current input information can enter the cell state. The calculation of the activation value of the forget gate determines which information in the cell state needs to be forgotten. The output gate determines which information to output as the forward hidden state based on the current cell state. In the above manner, the forward long short-term memory unit can gradually update the cell state over time and output a sequence of forward hidden states. The sequence of forward hidden states can reflect the state changes of the initial fusion feature in the forward time order, providing a forward information basis for capturing the complete temporal dependence relationship subsequently. Thus, through the gating mechanism of the forward long short-term memory unit, the temporal information in the initial fusion feature can be effectively processed, and a hidden state sequence related to the forward time order can be gradually generated, which helps to discover the feature change rules of the initial fusion feature in the forward time dimension and provides an important basis for comprehensively understanding the temporal characteristics of noise interference.

[0072] S3042, Input the initial fusion feature into the backward long short-term memory unit in reverse chronological order to generate a sequence of backward hidden states. Among them, the backward long short-term memory unit adopts a structure symmetric to the forward unit, but its time processing direction is opposite to capture the reverse temporal dependence relationship.

[0073] In this embodiment, in order to capture the reverse temporal dependence relationship, the backward long short-term memory unit adopts a structure symmetric to the forward unit, but the time processing direction is opposite. For example, in an audio signal restoration system, although the initial fusion feature has integrated various information, unique information can also be obtained from the perspective of reverse time. When the initial fusion feature is input in reverse chronological order, the gating mechanism of the backward long short-term memory unit operates in the reverse time order. For example, it also calculates the activation values of the input gate, the forget gate, and the output gate according to the initial fusion feature at the current reverse time step and the hidden state at the previous reverse time step. Due to the opposite time order, the information it captures is different from that of the forward long short-term memory unit. The sequence of backward hidden states can reflect the state changes of the initial fusion feature in the reverse time order. Thus, the feature relationships that are difficult to detect from the forward time order can be discovered. For example, some noise interferences in the audio signal may present specific patterns in the reverse time order, and the sequence of backward hidden states generated by the backward long short-term memory unit can capture the above patterns, providing a reverse information supplement for comprehensively analyzing the temporal dependence relationship of noise interference.

[0074] S3043, perform frame-by-frame superposition processing on the forward hidden state sequence and the backward hidden state sequence to obtain a bidirectional fusion feature. Among them, the frame-by-frame superposition processing includes: splicing the channel dimensions of the forward hidden state and the backward hidden state at the same time point, and reducing the dimension to the original feature dimension through a fully connected layer.

[0075] In this application, the forward hidden state sequence and the backward hidden state sequence analyze the initial fusion feature from the forward and backward time orders respectively. The frame-by-frame superposition processing is the key to integrating the analysis results in these two different directions. Specifically, splicing the channel dimensions of the forward hidden state and the backward hidden state at the same time point is similar to combining the features of the same thing observed from two different perspectives. Then, the dimension is reduced to the original feature dimension through a fully connected layer. The fully connected layer plays a role in integrating and optimizing the features here. For example, in audio signal processing, the forward and backward hidden states may contain a lot of information. Through the dimension reduction operation of the fully connected layer, while retaining the key information, the obtained bidirectional fusion feature can have a dimension structure similar to the original initial fusion feature, which is convenient for subsequent processing. In this way, the information in the forward and backward hidden state sequences is effectively fused to generate a bidirectional fusion feature, which contains both the feature change information in the forward time order and the feature change information in the backward time order, thus more comprehensively describing the characteristics of the initial fusion feature in the time dimension and providing richer feature information for more accurate analysis of noise interference in audio.

[0076] S3044, perform a non-linear activation function mapping on the bidirectional fusion feature to generate a dynamic fusion feature; among them, the non-linear activation function mapping includes: using a gated linear unit to perform inter-channel information screening on the bidirectional fusion feature, retaining the feature components strongly related to noise interference, and suppressing irrelevant components.

[0077] In this embodiment, a gated linear unit is used to perform inter-channel information screening on the bidirectional fusion features, which can retain the feature components strongly related to noise interference and suppress the irrelevant components. For example, although the bidirectional fusion features have already fused the information in the forward and reverse time orders, they may still contain some information that is irrelevant to the description of noise interference. The gated linear unit performs inter-channel information screening. For example, according to certain calculation rules, the feature components in each channel are evaluated to determine their correlation with noise interference. The feature components strongly related to noise interference will be retained, while the components with weak correlation will be suppressed. Based on this, by screening and optimizing the information in the bidirectional fusion features, the generated dynamic fusion features are more focused on the features related to noise interference, so that the feature set can more accurately reflect the noise interference situation in the audio signal, providing a more targeted and effective feature input for subsequent operations such as training the noise compensation model based on the noise interference feature set, and further improving the performance of the entire noise compensation method in dealing with audio noise interference.

[0078] In a possible implementation manner, the above step S40 can be implemented by the following S401 - S405, and the specific description is as follows.

[0079] S401, input the noise interference feature set into the parameter prediction layer of the noise compensation model, and output the initial compensation parameters. The parameter prediction layer is composed of a fully connected network and a regressor. The fully connected network reduces the dimension of the noise interference feature set, and the regressor outputs the initial compensation parameters corresponding to the band gain and the time-domain filtering coefficient.

[0080] In this application, the noise interference feature set contains comprehensive description information of audio noise interference from multiple aspects and may have a relatively high dimension. Therefore, the fully connected network first performs dimension reduction processing on the noise interference feature set. For example, through its numerous connection weights, linear combination and non-linear transformation are performed on the input high-dimensional features, and they are mapped to a low-dimensional space. Then, based on the low-dimensional features output by the fully connected network, the regressor outputs the initial compensation parameters corresponding to the band gain and the time-domain filtering coefficient. The band gain and the time-domain filtering coefficient are key parameters for noise compensation of noisy audio. For example, the band gain can adjust the energy of different frequency sub-bands. For the frequency bands with noise interference, the energy can be adjusted through an appropriate band gain to achieve the purpose of suppressing noise. The time-domain filtering coefficient can be used to filter the audio signal in the time domain, for example, to eliminate some periodic noise interferences. In this way, the initial compensation parameters can be automatically generated according to the noise interference feature set, providing a starting adjustment strategy for subsequent noise compensation operations, and thus starting the process of denoising the noisy audio segment.

[0081] S402. Perform frequency-domain filtering on the noisy audio segment based on the initial compensation parameters to generate a denoised audio segment.

[0082] In this embodiment, the band gain coefficient in the initial compensation parameters is a key factor in the frequency-domain filtering process. The frequency-domain filtering process includes dynamically attenuating the energies of the subbands of the noisy audio segment according to the band gain coefficient in the initial compensation parameters. Since an audio signal is composed of subbands of different frequencies in the frequency domain, different subbands may be affected by different degrees of noise interference. For example, in a speech audio containing background noise, the high-frequency subbands may be interfered by the high-frequency noise generated by electronic devices, while the low-frequency subbands may be affected by environmental low-frequency noise. Through the band gain coefficient, the energies of the above-mentioned interfered subbands can be adjusted specifically. If the gain coefficient of a certain subband is less than 1, the energy of that subband will be attenuated, thereby suppressing the noise energy in that subband. At the same time, perform a convolution operation on the time-domain filtering coefficients to eliminate periodic noise interference. Periodic noise exhibits a certain periodic pattern in the time domain. By performing a convolution operation with appropriate time-domain filtering coefficients, the above-mentioned periodic noise components can be effectively identified and removed. After the frequency-domain filtering process, a denoised audio segment can be obtained. Compared with the original noisy audio segment, the noise interference in the frequency domain and the time domain of this denoised audio segment is suppressed to a certain extent, making it closer to a clean audio segment, providing a basis for calculating the difference loss between it and the clean audio segment in the subsequent process.

[0083] As an example, step S402 may include the following S4021 - S4024, which are specifically described as follows.

[0084] S4021. Perform frame division and windowing on the noisy audio segment to obtain a plurality of noisy audio frames.

[0085] Among them, the noisy audio segment is a continuous audio signal. The frame division operation can be to cut this continuous audio segment into multiple small segments according to a set time length, and each small segment is a noisy audio frame. The above frame division method enables subsequent processing to be carried out in smaller units, facilitating local analysis of the audio signal. For example, when processing a long music audio mixed with environmental noise, frame division can observe the noise situation in each small part of the audio. Each noisy audio frame after frame division contains local audio information, providing a suitable data unit for subsequent frequency-domain transformation and filtering operations, thus helping to more precisely process the noise components in the noisy audio.

[0086] Among them, the framed windowing process can adopt a Hamming window with an overlapping rate of the set overlapping rate. In this embodiment, the purpose of adopting a Hamming window with an overlapping rate of the set overlapping rate for the framed windowing process is to reduce the spectral leakage effect. For example, when framing an audio, if an appropriate window function is not used, spectral leakage will occur. For example, in a speech recognition system, if the spectral leakage is severe, it may cause the spectral characteristics of the speech to be blurred, affecting the subsequent recognition effect. The Hamming window is a commonly used window function, which has good spectral characteristics. Setting a certain overlapping rate and using the Hamming window is because during the framing process, if there is a certain overlap between adjacent frames and the Hamming window is used for weighting processing, it can make the transition between frames smoother and reduce the spectral leakage caused by the signal discontinuity due to framing. The reduction of the above spectral leakage effect can ensure that in the subsequent frequency-domain processing, more accurate spectral information can be obtained, thus laying a foundation for effectively performing noise filtering.

[0087] S4022. Perform fast Fourier transform processing on each noisy audio frame to obtain a noisy spectral frame.

[0088] Among them, for each noisy audio frame after framing, it is not easy to directly observe the distribution of noise in the frequency domain from the information in the time domain. Through the fast Fourier transform, the noisy audio frame is transformed from the time domain to the frequency domain to obtain a noisy spectral frame. For example, when processing an audio signal containing machine noise, the noisy audio frame looks like a chaotic waveform in the time domain, but the noisy spectral frame after the fast Fourier transform can clearly show the energy distribution of the noise at different frequencies. This helps to determine the specific position and intensity of the noise in the frequency domain, providing the necessary frequency-domain information for subsequent amplitude adjustment according to the band gain coefficient in the initial compensation parameter.

[0089] S4023. Adjust the amplitude of the noisy spectral frame according to the band gain coefficient in the initial compensation parameter to generate a denoised spectral frame. Among them, the amplitude adjustment includes: multiplying the band gain coefficient point by point with the amplitude spectrum of the noisy spectral frame and retaining the original phase information to ensure the continuity of the audio signal.

[0090] In this embodiment, the amplitude spectrum reflects the energy magnitudes of different frequency components. For example, in an audio restoration system, the band gain coefficients in the initial compensation parameters are noise suppression strategies for different frequency bands. For a noisy spectral frame, the band gain coefficients are multiplied point by point with the amplitude spectrum of the noisy spectral frame. For example, if the band gain coefficient of a certain frequency band is less than 1, then the energy of that frequency band in the noisy spectral frame will be attenuated, thereby achieving the purpose of suppressing the noise in that frequency band. Meanwhile, the original phase information is retained during this process to ensure the continuity of the audio signal. Because the phase information also plays an important role in the audio signal, changing the phase may cause distortion of the waveform of the audio signal in the time domain. Through the above amplitude adjustment method, the noise in the generated denoised spectral frame is suppressed to a certain extent in the frequency domain, and the basic characteristics of the audio signal are maintained.

[0091] S4024, perform an inverse fast Fourier transform on the denoised spectral frame to obtain a denoised audio frame, and perform a superposition and synthesis process on multiple denoised audio frames to generate a denoised audio segment. Among them, the superposition and synthesis process includes: using a linear weighted superposition algorithm for the overlapping denoised audio frames to eliminate the signal distortion introduced by frame division.

[0092] Among them, the inverse fast Fourier transform is the inverse process of the fast Fourier transform, which converts a frequency-domain signal back to a time-domain signal. For example, the denoised spectral frame obtained through the previous steps has been processed for noise suppression in the frequency domain. Through the inverse fast Fourier transform, it can be converted back to a time-domain denoised audio frame. However, due to the overlapping part in the previous frame division and windowing process, directly splicing the above denoised audio frames simply will cause signal distortion. Therefore, a superposition and synthesis process is adopted, and a linear weighted superposition algorithm is used for the overlapping denoised audio frames. For example, when processing a speech audio containing background noise, the above linear weighted superposition algorithm can assign different weights according to different positions of the overlapping part, thereby eliminating the signal distortion introduced by frame division, and finally generating a complete denoised audio segment. The noise in this denoised audio segment is effectively suppressed in the time domain and is closer to a pure audio segment.

[0093] S403, calculate the frequency-domain difference loss and time-domain waveform loss between the denoised audio segment and the corresponding pure audio segment.

[0094] In this application, the frequency-domain difference loss is an index that measures the difference degree between the denoised audio segment and the clean audio segment in the frequency domain. Since the energy distribution of the audio signal in the frequency domain is one of its important characteristics, the frequency-domain difference loss can be calculated by comparing the energy distribution differences between the two in different sub-bands. For example, methods such as mean square error (MSE) can be used to calculate the sum of the squares of the energy differences between the denoised audio segment and the clean audio segment in each frequency band, and the sum value reflects the difference degree in the frequency domain. If the energy of the denoised audio segment in a certain frequency band differs greatly from that of the clean audio segment, then the difference in the frequency band will contribute greatly to the frequency-domain difference loss. The time-domain waveform loss focuses on the waveform differences between the two in the time domain. The time-domain waveform of the audio signal reflects the change of sound over time, and the time-domain waveform loss can be determined by calculating the amplitude differences between the two at different time points. For example, the mean square error method can also be used to calculate the sum of the squares of the amplitude differences between the denoised audio segment and the clean audio segment at each time step. The calculation of the frequency-domain difference loss and the time-domain waveform loss can accurately quantify the differences between the denoised audio segment and the clean audio segment from two different important dimensions, providing a necessary basis for constructing a joint optimization objective and further optimizing the noise compensation model in the subsequent steps.

[0095] As an example, in step S403, it can be implemented through the following steps S4031 - S4035, which are specifically described as follows.

[0096] S4031, perform synchronous frame division processing on the denoised audio segment and the clean audio segment to obtain multiple groups of aligned audio frame pairs. Among them, the synchronous frame division processing uses the same frame division parameters as the noisy audio segment to ensure strict alignment of time-domain and frequency-domain features.

[0097] In this application, in order to accurately compare the differences between the denoised audio segment and the clean audio segment at different times and frequencies, it is necessary to ensure strict alignment of their time-domain and frequency-domain features. Among them, the synchronous frame division processing can use the same frame division parameters as the noisy audio segment. For example, when processing an audio containing speech information, frame division parameters such as frame length and frame shift determine the division method of audio frames. Using the same frame division parameters to frame the denoised audio segment and the clean audio segment makes the obtained audio frame pairs correspond one by one in time, so that in subsequent frequency-domain and time-domain feature comparisons, it can be ensured that the comparisons are carried out at the same time position and within the same frequency range, thus providing an accurate alignment basis for accurately calculating the difference loss.

[0098] S4032, calculate the spectral amplitude difference between the denoised audio frame and the clean audio frame in each group of audio frame pairs to generate the frame-level frequency-domain difference loss. Among them, the spectral amplitude difference calculation uses the mean square error of the logarithmic mel spectrum to highlight the error weight of the frequency bands sensitive to the human ear.

[0099] In this embodiment, the mean square error of the logarithmic Mel spectrum is used to calculate the spectral amplitude difference in order to highlight the error weight of the frequency bands sensitive to the human ear. In the field of audio processing, the human ear has different sensitivities to sounds of different frequencies. For example, in an audio scene of music playback, the human ear may be more sensitive to the sounds in the mid-low frequency band and relatively less sensitive to the high frequency band. The logarithmic Mel spectrum is a spectral representation method that can simulate the auditory characteristics of the human ear. By calculating the mean square error of the logarithmic Mel spectrum, when calculating the frequency-domain difference loss, the errors in the frequency bands sensitive to the human ear can account for a more important proportion in the overall loss. The frame-level frequency-domain difference loss calculated in this way can be more in line with the human ear's perception of audio quality, providing a more targeted frame-level basis for generating a more reasonable global frequency-domain difference loss later.

[0100] S4033. Calculate the mean square error between the time-domain waveform of the denoised audio segment and the time-domain waveform of the clean audio segment to generate a time-domain waveform loss. Among them, the time-domain waveform loss calculation uses the windowed segmented energy comparison method to reduce the influence of instantaneous noise spikes on the loss calculation.

[0101] In this application, the windowed segmented energy comparison method can be used to calculate the time-domain waveform loss to reduce the influence of instantaneous noise spikes on the loss calculation. In an audio signal, there may be some instantaneous noise spikes, and the above spikes may cause great interference to the calculation of the mean square error, resulting in the time-domain waveform loss not being able to accurately reflect the true difference of the audio signal in the time domain. For example, in the processing of an audio signal containing burst noise interference, using the windowed segmented energy comparison method is similar to dividing the time-domain waveform into several small segments and performing energy comparison within each small segment. By smoothing out the influence of instantaneous noise spikes to a certain extent, the calculated time-domain waveform loss can more accurately reflect the true difference between the denoised audio segment and the clean audio segment in the time-domain waveform, providing accurate loss information in the time domain aspect for constructing the joint optimization objective.

[0102] S4034. Perform weighted summation on all frame-level frequency-domain difference losses to obtain a global frequency-domain difference loss; among them, the weight coefficients corresponding to the weighted summation are dynamically adjusted according to the signal-to-noise ratio of each frame.

[0103] In this embodiment, the signal-to-noise ratios of different audio frames may vary, and the contribution of the error in the frequency domain to the overall frequency-domain difference loss also differs. For example, in an audio segment that contains both low-signal-to-noise-ratio and high-signal-to-noise-ratio audio frames, the frequency-domain error of the low-signal-to-noise-ratio audio frames may have a greater impact on the overall audio quality. By dynamically adjusting the weight coefficients according to the signal-to-noise ratio of each frame, the audio frames with low signal-to-noise ratios can be given a greater weight in the weighted summation, enabling the global frequency-domain difference loss to more accurately reflect the overall difference in the frequency domain of the entire audio segment, and providing accurate loss information in the frequency domain for subsequent reasonable combination of the global frequency-domain difference loss and the time-domain waveform loss.

[0104] S4035, combine the global frequency-domain difference loss and the time-domain waveform loss according to a preset ratio to generate a joint optimization objective. Wherein, the preset ratio is dynamically set according to the cross-validation result of the validation set.

[0105] In this application, the preset ratio is dynamically set according to the cross-validation result of the validation set to find an optimal combination method. During the optimization process of the audio processing algorithm, the cross-validation result of the validation set can reflect the performance of the model under different ratio combinations. For example, during the development of an audio noise reduction algorithm, by performing multiple cross-validations on the validation set and trying different preset ratios, a ratio that enables the model to perform best in terms of noise reduction effect and audio quality preservation can be found. The preset ratio dynamically set according to the cross-validation result can ensure a balance between the losses in the frequency domain and the time domain of the joint optimization objective, thereby providing a reasonable objective function for subsequent optimization of the noise compensation model based on the joint optimization objective.

[0106] S404, construct a joint optimization objective based on the frequency-domain difference loss and the time-domain waveform loss, and reversely update the network weights of the noise compensation model.

[0107] Among them, the frequency-domain difference loss reflects the degree of difference between the denoised audio segment and the clean audio segment in the frequency domain. For example, when the energy distribution of the denoised audio segment is different from that of the clean audio segment at certain frequencies, the frequency-domain difference loss will reflect the magnitude of the above differences. The time-domain waveform loss reflects the difference between the two in the time-domain waveform. These two losses measure the gap between the denoising effect and the ideal clean audio from different dimensions. Constructing a joint optimization objective is to combine these two losses in a reasonable way. For example, the weighted summation method can be used to assign different weights according to the importance of the frequency domain and the time domain in the entire audio feature. Then, based on this joint optimization objective, the network weights of the noise compensation model are updated in the reverse direction. The backpropagation algorithm is a commonly used method to implement this process. During the backpropagation process, the error propagates from the output layer (i.e., the calculated loss) to the input layer, and the weights are adjusted according to the partial derivatives of the loss with respect to the network weights. By continuously learning how to better adjust the weights in the above way, the value of the joint optimization objective is reduced. In this way, by considering the losses in both the frequency domain and the time domain to update the network weights, the noise compensation model can comprehensively optimize its own parameters, thereby improving the compensation ability for different types of noise, and better handling both the energy distribution problem in the frequency domain and the waveform difference problem in the time domain.

[0108] S405, when the joint optimization objective converges, lock the network parameters of the noise compensation model and export the set of noise compensation parameters. Among them, the frequency-domain filtering process includes: dynamically attenuating the energies of each sub-band of the noisy audio segment according to the band gain coefficient in the initial compensation parameters, and at the same time performing a convolution operation on the time-domain filtering coefficient to eliminate periodic noise interference.

[0109] In one embodiment, the convergence of the joint optimization objective means that as the noise compensation model is continuously trained, the value of the joint optimization objective composed of the frequency-domain difference loss and the time-domain waveform loss no longer decreases significantly or has reached a pre-set convergence criterion. For example, after a certain number of training iterations, the change in the loss value may be very small and approach a stable value. At this time, the noise compensation model has learned a relatively stable parameter configuration and can effectively compensate for the input noisy audio segment. Locking the network parameters of the noise compensation model is to fix the optimized parameter state and prevent the above parameters from being accidentally modified during subsequent operations. Then, a set of noise compensation parameters is exported. This set of parameters includes various parameters of the model in the converged state, such as band gain coefficients, time-domain filtering coefficients, etc. The above parameters reflect the compensation strategies of the model for different types of noise after sufficient training. For example, the band gain coefficient determines the amplitude of adjusting the energy of different sub-bands in the frequency domain, and the time-domain filtering coefficient determines how to eliminate periodic noise interference in the time domain. In this way, by locking the network parameters and exporting the set of noise compensation parameters, an effective set of parameters for a specific audio noise compensation problem can be obtained. This effective set of parameters can be directly applied to the real-time processing module in the actual audio transmission system, thereby improving the noise compensation effect during the audio transmission process and providing users with a higher-quality audio experience.

[0110] In a possible implementation manner, the above step S50 can be implemented by the following S501 - S505, and the specific description is as follows.

[0111] S510, Deploy the trained noise compensation model in the real-time processing module and receive the real-time input noisy audio signal. Among them, the real-time processing module adopts a multi-thread buffer mechanism.

[0112] In this application, the real-time processing module is used in the audio transmission system to process the continuously input audio signals in a timely manner. The multi-thread buffer mechanism is adopted to improve the processing efficiency and stability. For example, in a real-time voice call audio transmission system, multiple audio data may flow into the real-time processing module simultaneously. The multi-thread mechanism allows the above data to be processed in parallel, thereby improving the processing speed. The buffer mechanism can play a buffering role when data flows in and out of the module, avoiding data loss or untimely processing. When the trained noise compensation model is deployed in this real-time processing module, it can process the real-time input noisy audio signal at any time.

[0113] S520, Perform noise interference feature extraction processing on the real-time input noisy audio signal to generate real-time noise features. Among them, the noise interference feature extraction processing uses the same frame segmentation parameters and feature extraction algorithms as in the training stage.

[0114] In this embodiment, the noise interference features of the real-time input noisy audio signal are extracted to generate real-time noise features. It is very important to use the same frame segmentation parameters and feature extraction algorithms as in the training stage here. In an audio live broadcast system, the frame segmentation parameters determine how to divide the continuous noisy audio signal into small segments that are easy to process, and the feature extraction algorithm is used to extract the information that can represent the noise features from the above small segments. Since in the training stage, the noise compensation model is trained based on the noise interference features obtained by specific frame segmentation parameters and feature extraction algorithms, using the same method in the real-time processing stage can ensure that the extracted real-time noise features are consistent with those in the training stage. In this way, the noise compensation model can perform accurate processing based on the familiar feature representation, thereby improving the accuracy of noise compensation.

[0115] S530. Invoke the noise compensation model to perform dynamic feature fusion processing on the real-time noise features to generate real-time noise interference features. Among them, the dynamic feature fusion processing performs the same residual convolution and frequency domain attention operations on the real-time noise features through the pre-loaded model weights as in the training stage.

[0116] In this application, the pre-loaded model weights contain the important information about noise feature fusion learned by the noise compensation model during training. The residual convolution operation can deeply explore the time-domain information in the real-time noise features and extract the deep time-domain features related to noise from the features at different levels. The frequency domain attention operation can perform weight allocation according to the importance of different frequency bands of the frequency domain noise features, and focus on the frequency bands that have a greater impact on the noise. By performing the above operations same as in the training stage, the generated real-time noise interference features can accurately reflect the noise interference situation in the real-time input noisy audio signal, providing a suitable input for subsequent compensation parameter matching.

[0117] S540. Match the target compensation parameter in the noise compensation parameter set according to the real-time noise interference features. Among them, the matching process uses the nearest neighbor search algorithm to select the historical parameter with the smallest Euclidean distance from the noise compensation parameter set to the real-time noise interference features.

[0118] In this embodiment, the noise compensation parameter set contains the compensation parameters in various different situations obtained from previous training. The real-time noise interference features represent the noise situation of the currently input noisy audio signal. By using the nearest neighbor search algorithm to find the historical parameter with the smallest Euclidean distance, the compensation parameter corresponding to the historical processing situation most similar to the current noise situation can be found. The selected target compensation parameter can adapt to the noise characteristics of the real-time input noisy audio signal to the greatest extent, providing a suitable parameter basis for accurate noise compensation.

[0119] Among them, step S540 may include the following steps S5401 - S5404, which are specifically described as follows.

[0120] S5401, calculate the cosine similarity between the real - time noise interference feature and each historical noise interference feature in the noise compensation parameter set. Among them, before calculating the cosine similarity, L2 normalization processing is performed on the real - time noise interference feature and the historical noise interference feature to eliminate the influence of amplitude difference.

[0121] Based on the above content, for the real - time noise interference feature and the historical noise interference feature, they are both noise - related features represented in vector form. For example, when processing an audio stream, the real - time noise interference feature reflects the noise situation in the audio at the current moment, while the historical noise interference feature is a record of the noise situation in the previously processed audio. By calculating the cosine similarity, the degree of closeness in direction between the real - time noise interference feature and each historical noise interference feature can be determined, which helps to find the historical situation most similar to the current noise situation, thus providing a basis for selecting appropriate compensation parameters.

[0122] Considering that in audio feature representation, amplitude differences may interfere with the accurate calculation of cosine similarity. For example, different audio signals may have different amplitudes due to various reasons (such as signal source strength, transmission distance, etc.). Without normalization processing, even if two noise interference features are very similar in direction, the calculated cosine similarity may be inaccurate due to amplitude differences. L2 normalization processing divides each element of the feature vector by the Euclidean norm of the vector, making the processed vector have a unit length. In this way, when calculating the cosine similarity, it can more accurately reflect the directional similarity between features and avoid the misguidance of amplitude differences on the matching process.

[0123] S5402, select the compensation parameters corresponding to the historical noise interference features with cosine similarity greater than a preset threshold as the candidate parameter set.

[0124] During the audio processing process, the preset threshold is a judgment standard set according to experience or system requirements. For example, in an audio noise reduction application scenario, if the cosine similarity is greater than the preset threshold, it indicates that the historical noise interference feature is similar enough to the real - time noise interference feature, and its corresponding compensation parameter may be applicable to the current real - time audio situation. By setting such a threshold and screening out the compensation parameters corresponding to the historical noise interference features that meet the conditions to form a candidate parameter set, the compensation parameters in it are all matched to a certain extent with the current real - time noise interference feature, providing a preliminary screening range for further determining the target compensation parameter.

[0125] S5403. Perform weighted average processing on the compensation parameters in the candidate parameter set to generate a target compensation parameter. The weight coefficients of the weighted average are dynamically allocated according to the cosine similarity values. The higher the similarity, the greater the weight of the corresponding parameter.

[0126] Although each compensation parameter in the candidate parameter set has a certain similarity to the real-time noise interference feature, the degree of similarity is different. Dynamically allocating weight coefficients according to the cosine similarity values enables the compensation parameters that are more similar to the real-time noise interference feature (i.e., have a higher cosine similarity) to have a greater influence when generating the target compensation parameter. For example, if the cosine similarity between the historical noise interference feature corresponding to a compensation parameter and the real-time noise interference feature is very high, then its weight in the weighted average will be larger, and its contribution to the finally generated target compensation parameter will also be greater. Through weighted average processing, the information of each compensation parameter in the candidate parameter set can be comprehensively considered to generate a target compensation parameter that better conforms to the current real-time noise interference feature.

[0127] S5404. When the cosine similarities are all less than a preset threshold, call the noise compensation model to generate new compensation parameters and add them to the noise compensation parameter set. Among them, the new compensation parameters are generated using the model inference mode, and directly output the compensation parameters that have not undergone historical parameter matching according to the real-time noise interference feature.

[0128] In this application, when the cosine similarities are all less than the preset threshold, it indicates that the existing historical noise interference features are quite different from the real-time noise interference feature, and it is impossible to find suitable historical compensation parameters to handle the noise situation in the current real-time audio. The model inference mode of the noise compensation model can directly output the compensation parameters that have not undergone historical parameter matching according to the real-time noise interference feature. The new compensation parameters are generated based on the noise compensation model's understanding of the real-time noise interference feature. Adding the new compensation parameters to the noise compensation parameter set helps to enrich the content of the noise compensation parameter set, so that when encountering similar real-time noise interference features in the future, there will be more reference bases, improving the processing ability of the entire system for different noise situations.

[0129] S550. Perform frequency-domain gain adjustment on the real-time input noisy audio signal based on the target compensation parameter, and output the denoised real-time audio signal. Among them, the frequency-domain gain adjustment uses the overlap-and-save method to preserve the integrity of the edge frames when restoring the time-domain signal through inverse transformation after frequency-domain filtering.

[0130] In this application, the overlapping save method is adopted for frequency-domain gain adjustment to preserve the integrity of the edge frames when restoring the time-domain signal through inverse transformation after frequency-domain filtering. For example, frequency-domain gain adjustment is a process of adjusting the energy of the noisy audio signal in the frequency domain according to the target compensation parameter. After the frequency-domain filtering operation, the signal needs to be converted back from the frequency domain to the time domain. In this conversion process, the processing of the edge frames is relatively special. If not handled properly, it may lead to signal incompleteness or generate additional noise. The overlapping save method ensures the integrity of the edge frames is preserved when restoring the time-domain signal by reasonably processing the overlapping parts of the edge frames, thereby outputting a high-quality denoised real-time audio signal and improving the quality of audio playback.

[0131] Based on the above, the method of the embodiment of this application further includes steps such as S60-S90 described below, and the specific description is as follows.

[0132] S60, periodically collect the environmental noise samples of the audio transmission system to generate an incremental noise data set.

[0133] In this application, periodically collecting the environmental noise samples of the audio transmission system to generate an incremental noise data set is to continuously update the noise compensation model to adapt to new noise situations. For example, in an audio monitoring environment, such as the audio transmission system in an airport waiting hall, the surrounding environmental noise situation is constantly changing. It may vary with the change of passenger flow at different time periods, and the noise sources and intensities will also be different. Periodic collection can capture the environmental noise that changes over time. By collecting samples at different times to form an incremental noise data set, an information library about the environmental noise that is continuously updated is formed, providing a new data source for subsequent improvement of the noise compensation model to cope with possible new types of noise interference.

[0134] S70, perform noise interference feature extraction processing on the incremental noise data set to generate an incremental noise interference feature set.

[0135] In this embodiment, performing noise interference feature extraction processing on the incremental noise data set to generate an incremental noise interference feature set requires extracting the features that can represent noise interference from the incremental noise data set. The incremental noise data set may contain abnormal noise or noise changes caused by environmental factor changes. Through specific feature extraction processing, the above complex noise information can be converted into an incremental noise interference feature set, which contains the key information extracted from the newly collected noise samples, such as the special performance of the new noise in the time domain and frequency domain, etc., providing suitable input features for subsequent incremental training.

[0136] S80. Input the incremental noise interference feature set into the noise compensation model for incremental training to generate an updated set of noise compensation parameters. Among them, the incremental training uses the Elastic Weight Consolidation algorithm to retain the memory ability for historical data when updating the model parameters.

[0137] In this application, inputting the incremental noise interference feature set into the noise compensation model for incremental training to generate an updated set of noise compensation parameters can use the Elastic Weight Consolidation algorithm to retain the memory ability for historical data when updating the model parameters. For example, in a complex audio processing scenario, such as the audio transmission system of a large concert, the previous noise compensation model was trained based on previous noise data and was already able to handle some common noise situations well. When a new incremental noise interference feature set is input, using the Elastic Weight Consolidation algorithm can, while adjusting the model parameters to adapt to the new noise situation, avoid over-forgetting the processing ability for the previous old noise situations. For example, the noise compensation model will not completely change the previous compensation strategy for some common noises due to the new noise data, but on the basis of retaining the historical memory, reasonably update the parameters to adapt to the new noise interference, thereby generating an updated set of noise compensation parameters.

[0138] S90. Dynamically load the updated set of noise compensation parameters into the real-time processing module to replace the original set of noise compensation parameters.

[0139] In this embodiment, the real-time processing module needs to use the latest and most environment-noise-situation-adaptive noise compensation parameters. When the updated set of noise compensation parameters is generated, the dynamic loading process can achieve the upgrade of the real-time processing module. By replacing the original set of noise compensation parameters, the real-time processing module can immediately start using the new parameters to process the noisy audio signals transmitted in real time. In this way, the audio transmission system can better cope with newly emerging noise interference, improve the quality of audio transmission, and provide a clearer audio experience for listeners.

[0140] As Figure 3 shown, it is a schematic diagram of a noise interference compensation system applied to an audio transmission system provided by an embodiment of this application. The noise interference compensation system includes components such as a processor, a machine-readable storage medium, and input / output devices. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes. The processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above-mentioned noise interference compensation method applied to the audio transmission system.

[0141] Among them, the machine-readable storage medium may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. Among them, the machine-readable storage medium is used to store a program, and the processor executes the program after receiving an execution instruction.

[0142] The processor may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be, but is not limited to, a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.

[0143] In summary, the noise interference compensation method and system provided by the embodiments of the present application for an audio transmission system first comprehensively extract the time-domain and frequency-domain noise characteristics from the set of noisy audio signals. Compared with the traditional single-domain feature extraction, it can depict the noise characteristics more meticulously and completely. Then, a preset noise compensation model is used for dynamic feature fusion, breaking the conventional fixed-mode feature processing method. It can flexibly adjust the fusion strategy according to the complex noise conditions of different audio segments, and generate a feature set that accurately reflects the noise interference situation. The set of noise compensation parameters obtained by training the model using the mapping relationship between this set and the clean audio segments can accurately adapt to the noisy audio signals in real-time transmission. Loading it into the real-time processing module for noise interference compensation operations effectively overcomes the problem of variable noise interference during real-time transmission, greatly reducing the impact of noise on audio quality. Compared with the prior art, it has achieved a qualitative leap in improving the clarity, fidelity, and stability of audio transmission, providing a strong guarantee for the reliable operation of the audio transmission system in a complex noise environment. Therefore, the anti-interference ability of the audio transmission system is significantly improved.

[0144] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Among the embodiments, implementation manners and related technical features of the present application, they can be combined and replaced with each other without conflict. The above are only the preferred embodiments of the present application, and do not impose any formal restrictions on the present application. However, any simple modifications, equivalent changes and decorations made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application still fall within the scope of the technical solution of the present application.

Claims

1. A method for compensating noise interference in an audio transmission system, characterized in that The method includes: Obtaining a set of noisy audio signals for a target audio transmission channel, where the set of noisy audio signals includes multiple noisy audio segments and corresponding clean audio segments for each noisy audio segment; Performing noise interference feature extraction processing on the set of noisy audio signals to obtain the time-domain noise features and frequency-domain noise features of each noisy audio segment; Based on a preset noise compensation model, performing dynamic feature fusion processing on the time-domain noise features and the frequency-domain noise features to generate a set of noise interference features for the noisy audio segment, including: calling a residual convolutional network in the noise compensation model to perform multi-layer convolutional processing on the time-domain noise features to extract time-domain deep features; where the residual convolutional network includes a three-level convolutional structure, and each level of convolutional structure consists of convolutional layers with decreasing convolutional kernel sizes, and a skip connection is introduced after the output of each level of convolutional layer to retain the low-order information of the original time-domain noise features; calling a frequency-domain attention network in the noise compensation model to perform frequency band weight allocation processing on the frequency-domain noise features to extract frequency-domain attention features; where the frequency-domain attention network includes a sub-band division module and a weight calculation module, the sub-band division module divides the frequency-domain noise features into multiple non-uniform sub-bands, and the weight calculation module generates frequency band weight coefficients according to the energy ratio and spectral flatness of each sub-band; performing cross-domain splicing on the time-domain deep features and the frequency-domain attention features to obtain an initial fusion feature; performing bidirectional long short-term memory network processing on the initial fusion feature to capture temporal dependencies and generate a dynamic fusion feature as the set of noise interference features; According to the mapping relationship between the set of noise interference features and the corresponding clean audio segments, training the noise compensation model to generate a set of noise compensation parameters, including: inputting the set of noise interference features into the parameter prediction layer of the noise compensation model to output initial compensation parameters; where the parameter prediction layer consists of a fully connected network and a regressor, the fully connected network reduces the dimension of the set of noise interference features, and the regressor outputs initial compensation parameters corresponding to the frequency band gain and time-domain filtering coefficients; performing frequency-domain filtering processing on the noisy audio segment based on the initial compensation parameters to generate a denoised audio segment; calculating the frequency-domain difference loss and time-domain waveform loss between the denoised audio segment and the corresponding clean audio segment; constructing a joint optimization objective according to the frequency-domain difference loss and the time-domain waveform loss, and reversely updating the network weights of the noise compensation model; when the joint optimization objective converges, locking the network parameters of the noise compensation model and exporting the set of noise compensation parameters; where the frequency-domain filtering processing includes: dynamically attenuating the energy of each sub-band of the noisy audio segment according to the frequency band gain coefficient in the initial compensation parameters, and performing a convolutional operation on the time-domain filtering coefficient to eliminate periodic noise interference; Loading the set of noise compensation parameters into the real-time processing module of the audio transmission system to perform noise interference compensation operations on the real-time transmitted noisy audio signals.

2. The method according to claim 1, characterized in that Perform noise interference feature extraction processing on the set of noisy audio signals to obtain the time-domain noise features and frequency-domain noise features of each noisy audio segment, including: Perform frame segmentation and windowing processing on the noisy audio segment to obtain multiple time-domain audio frames; Perform short-time energy analysis and zero-crossing rate calculation on each time-domain audio frame to generate time-domain noise features, where the time-domain noise features include an energy fluctuation sequence and a zero-crossing rate distribution sequence; Perform fast Fourier transform processing on each time-domain audio frame to obtain frequency-domain audio frames, extract the spectral envelope features and sub-band energy ratio features of the frequency-domain audio frames, and generate frequency-domain noise features; Perform time alignment processing on the time-domain noise features and the frequency-domain noise features to form a noise interference feature set corresponding to the noisy audio segment.

3. The method according to claim 1, wherein The frequency-domain filtering processing of the noisy audio segment based on the initial compensation parameter to generate a denoised audio segment includes: Perform frame segmentation and windowing processing on the noisy audio segment to obtain multiple noisy audio frames; Perform fast Fourier transform processing on each noisy audio frame to obtain a noisy spectral frame; Adjust the amplitude of the noisy spectral frame according to the band gain coefficient in the initial compensation parameter to generate a denoised spectral frame; where the amplitude adjustment includes: multiplying the band gain coefficient by the amplitude spectrum of the noisy spectral frame point by point and retaining the original phase information to ensure the continuity of the audio signal; Perform inverse fast Fourier transform processing on the denoised spectral frame to obtain denoised audio frames, and perform superposition synthesis processing on multiple denoised audio frames to generate a denoised audio segment; where the superposition synthesis processing includes: using a linear weighted superposition algorithm for the overlapping denoised audio frames to eliminate the signal distortion introduced by frame segmentation.

4. The method according to claim 1, characterized in that, The calculation of the frequency-domain difference loss and the time-domain waveform loss between the denoised audio segment and the corresponding clean audio segment includes: Perform synchronous frame segmentation processing on the denoised audio segment and the clean audio segment to obtain multiple groups of aligned audio frame pairs; where the synchronous frame segmentation processing uses the same frame segmentation parameters as the noisy audio segment to ensure strict alignment of time-domain and frequency-domain features; Calculate the spectral amplitude difference between the denoised audio frame and the clean audio frame in each group of audio frame pairs to generate a frame-level frequency-domain difference loss; where the spectral amplitude difference calculation uses the mean square error of the logarithmic Mel spectrum to highlight the error weight of the frequency bands sensitive to the human ear; Calculate the mean square error between the time-domain waveform of the denoised audio segment and the time-domain waveform of the clean audio segment to generate a time-domain waveform loss; where the time-domain waveform loss calculation uses a windowed segmented energy comparison method to reduce the influence of instantaneous noise spikes on the loss calculation; Perform weighted summation on all frame-level frequency-domain difference losses to obtain a global frequency-domain difference loss; where the weight coefficient corresponding to the weighted summation is dynamically adjusted according to the signal-to-noise ratio of each frame; Combine the global frequency-domain difference loss and the time-domain waveform loss according to a preset ratio to generate a joint optimization target; where the preset ratio is dynamically set according to the cross-validation results of the validation set.

5. The method according to claim 1, wherein Loading the set of noise compensation parameters into the real-time processing module of the audio transmission system to perform noise interference compensation operations on the real-time transmitted noisy audio signal, including: Deploying the trained noise compensation model in the real-time processing module to receive the real-time input noisy audio signal; wherein, the real-time processing module adopts a multi-thread buffer mechanism; Performing noise interference feature extraction processing on the real-time input noisy audio signal to generate real-time noise features; wherein, the noise interference feature extraction processing adopts the same frame division parameters and feature extraction algorithm as in the training stage; Invoking the noise compensation model to perform dynamic feature fusion processing on the real-time noise features to generate real-time noise interference features; wherein, the dynamic feature fusion processing performs the same residual convolution and frequency domain attention operations on the real-time noise features through the pre-loaded model weights as in the training stage; Matching the real-time noise interference features with the target compensation parameters in the set of noise compensation parameters; wherein, the matching process adopts the nearest neighbor search algorithm to select the historical parameter with the smallest Euclidean distance from the real-time noise interference features from the set of noise compensation parameters; Performing frequency domain gain adjustment on the real-time input noisy audio signal based on the target compensation parameters and outputting the denoised real-time audio signal; wherein, the frequency domain gain adjustment adopts the overlap-save method to preserve the integrity of the edge frames when restoring the time domain signal through inverse transformation after frequency domain filtering.

6. The method according to claim 5, wherein The matching the target compensation parameters in the set of noise compensation parameters according to the real-time noise interference features includes: Calculating the cosine similarity between the real-time noise interference features and each historical noise interference feature in the set of noise compensation parameters; wherein, before calculating the cosine similarity, the real-time noise interference features and historical noise interference features are subjected to L2 normalization processing to eliminate the influence of amplitude differences; Selecting the compensation parameters corresponding to the historical noise interference features with cosine similarity greater than the preset threshold as the candidate parameter set; Performing weighted average processing on the compensation parameters in the candidate parameter set to generate target compensation parameters; the weight coefficients of the weighted average are dynamically allocated according to the cosine similarity values, and the higher the similarity, the greater the weight of the corresponding parameter; When the cosine similarities are all less than the preset threshold, invoking the noise compensation model to generate new compensation parameters and adding them to the set of noise compensation parameters; wherein, the generation of the new compensation parameters adopts the model inference mode and directly outputs the compensation parameters that have not been matched with historical parameters according to the real-time noise interference features.

7. The method according to claim 1, wherein The method further includes: Periodically collecting the environmental noise samples of the audio transmission system to generate an incremental noise data set; Performing noise interference feature extraction processing on the incremental noise data set to generate an incremental noise interference feature set; Inputting the incremental noise interference feature set into the noise compensation model for incremental training to generate an updated set of noise compensation parameters; wherein, the incremental training adopts the elastic weight consolidation algorithm to retain the memory ability for historical data when updating the model parameters. Dynamically load the updated set of noise compensation parameters into the real-time processing module to replace the original set of noise compensation parameters.

8. A noise interference compensation system applied to an audio transmission system, characterized in that, It includes a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the noise interference compensation method applied to the audio transmission system according to any one of claims 1-7.

Citation Information

Patent Citations

  • Speech enhancement algorithm based on improved phase spectrum compensation and full convolutional neural network

    CN114242099A

  • Recording de-noising method, system and equipment based on sound characteristics and medium

    CN119694328A