Complex sound field live broadcast audio anti-interference optimization method and system
By using time-frequency transformation and spatial positioning technology of multi-channel audio data, interference sources in live teaching can be identified and suppressed, solving the problem of inaccurate interference source identification in traditional methods. This achieves a balance between interference suppression and voice fidelity, thus improving the audio quality of remote teaching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING HONGSHENG YANXUE EDUCATION CONSULTING CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-05-05
AI Technical Summary
In existing live teaching systems, traditional audio processing methods struggle to track and suppress dynamic and spatially uncertain interference sources in real time, resulting in residual interference components in the purified audio, which affects the clarity and immersion of the lesson.
By collecting multi-channel audio data and performing time-frequency transformation, calculating phase difference and spatial orientation angle, statistically analyzing cumulative energy distribution, identifying the orientation of interference sources, establishing the time-frequency evolution trajectory of interference sources, generating time-varying filter coefficients for adaptive attenuation, and achieving precise suppression of interference frequency bands.
It achieves real-time and intelligent suppression of sudden and mobile interference noise, ensuring the clarity and fidelity of teaching audio and improving the remote learning experience.
Smart Images

Figure CN121983018A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to audio signal processing technology, and more particularly to a method and system for optimizing live audio in complex sound fields to prevent interference. Background Technology
[0002] In existing live-streaming teaching systems, audio processing typically employs traditional noise reduction techniques. A common approach is to use single-channel or fixed-beamform microphone arrays to capture the teacher's voice, combined with noise suppression algorithms such as spectral subtraction or Wiener filtering based on statistical models to reduce ambient noise. These methods primarily rely on the differences in the spectral characteristics of the audio signals, estimating the noise spectrum and subtracting it from the mixed signal to enhance the speech. Some more advanced solutions utilize the spatial information from dual microphones for simple sound source localization, distinguishing the target sound source in front from some ambient noise, and then implementing directional pickup or fixed-direction filtering.
[0003] However, existing conventional methods have significant drawbacks. Interference sources in live-streamed teaching scenarios, such as sudden questions from students, temporary noise outside the classroom, and noise from moving objects, often exhibit dynamic characteristics and uncertain spatial orientation. Traditional methods based on fixed noise spectrum estimation or static beam pointing struggle to track the continuous spatial and temporal evolution of these interference sources in real time. This results in either excessive suppression of interfering sounds similar to the teacher's speech spectrum, causing speech distortion, or an inability to effectively track and suppress moving or intermittent interference sources, leaving noticeable interference components in the purified audio and affecting the clarity and immersion of the live-streamed lesson. Summary of the Invention
[0004] This invention provides a method and system for optimizing live audio anti-interference in complex sound fields, which can solve the problems in the prior art.
[0005] A first aspect of this invention provides a method for optimizing live audio playback in complex sound fields to prevent interference, comprising:
[0006] Collect multi-channel audio data from a live teaching session, and perform time-frequency transformation on the multi-channel audio data to obtain time-spectrum data;
[0007] The phase difference of each channel in the time-frequency data is calculated, and the spatial orientation angle of each time-frequency point is determined based on the phase difference to obtain the time-frequency spatial distribution data.
[0008] The cumulative energy distribution of each spatial orientation interval in the statistical time-frequency spatial distribution data is compared with the preset teaching sound source spatial area to identify the orientation of the interference source outside the teaching sound source spatial area and obtain the spatial positioning result of the interference source.
[0009] Based on the spatial location results of the interference source, the set of time and frequency points belonging to the location of the interference source is extracted from the time and frequency spatial distribution data, the time and frequency evolution trajectory of the interference source is established, and the time and frequency points in the time and frequency spectrum data are dynamically classified and labeled according to the time and frequency evolution trajectory of the interference source to obtain time and frequency classification labels.
[0010] Based on the time-frequency classification labels and the time-varying characteristics of the time-frequency evolution trajectory of the interference source, time-varying filter coefficients that are synchronized with the evolution of the interference source are generated. The time-varying filter coefficients are applied to the time-spectrum data to adaptively attenuate the interference frequency band, and the filtered time-spectrum data is obtained.
[0011] The filtered time-frequency data is subjected to inverse time-frequency transformation to generate an audio output signal after interference is eliminated.
[0012] Collect multi-channel audio data from a live teaching session, perform time-frequency transformation on the multi-channel audio data, and obtain time-spectrum data including:
[0013] The audio signals of the teaching live broadcast scene are collected synchronously by multiple spatially distributed audio acquisition units. The spatial position coordinates of each audio acquisition unit are recorded, and the correspondence between the audio acquisition unit and the spatial position coordinates is established to obtain multi-channel audio data with spatial position identifiers.
[0014] The multi-channel audio data is subjected to windowing and frame segmentation processing to divide the continuous audio data into multiple audio frames that overlap in time. Short-time Fourier transform is performed on each audio frame to obtain the frequency domain complex spectrum corresponding to each audio frame.
[0015] The amplitude and phase values of each frequency point in the frequency domain complex spectrum are extracted, and the amplitude and phase values are fused with the time position, frequency position, and spatial position coordinates of the corresponding audio frame and the channel phase space data to obtain time-frequency spectrum data containing time-frequency amplitude, time-frequency phase, and spatial position information.
[0016] Calculate the phase difference of each channel in the time-frequency data, determine the spatial orientation angle of each time-frequency point based on the phase difference, and obtain the time-frequency spatial distribution data including:
[0017] The complex spectrum values of each channel at each time frequency point are extracted from the time spectrum data, and the phase is calculated from the complex spectrum values to obtain the phase angle values of each channel at each time frequency point.
[0018] Select a channel with known spatial coordinates as the phase reference channel, calculate the difference in phase angle values between each non-reference channel and the phase reference channel at the same time frequency point, and obtain the inter-channel phase difference data including the time frequency point identifier and the phase difference value;
[0019] The phase difference values at each time frequency point in the phase difference data between channels are extracted. Combined with the spatial baseline distance between audio acquisition units and the sound wave propagation speed, the propagation path difference of the sound wave from the sound source to different audio acquisition units is calculated. Based on the spatial geometric position relationship between the propagation path difference and the audio acquisition unit array, the spatial incident angle value of the sound wave relative to the normal of the audio acquisition unit array is calculated.
[0020] By associating and binding the spatial incident angle values of each time and frequency point with the time and frequency point identifiers in the inter-channel phase difference data, a correspondence between the time and frequency point identifiers and the spatial incident angle values is established, thus obtaining the time and frequency spatial distribution data.
[0021] The cumulative energy distribution of each spatial location interval in the statistical time-frequency spatial distribution data is compared with the preset teaching sound source spatial area to identify the location of interference sources outside the teaching sound source spatial area, and the spatial localization results of the interference sources are obtained, including:
[0022] Extract the spatial incident angle value and time spectrum amplitude value of each time frequency point from the time frequency spatial distribution data. Divide the spatial incident angle value into multiple spatial azimuth intervals according to the preset angle interval. Accumulate the time spectrum amplitude value of all time frequency points in each spatial azimuth interval and statistically obtain the cumulative energy distribution of each spatial azimuth interval.
[0023] A time-frequency energy matrix is constructed for the energy accumulation distribution of each spatial orientation interval. The energy variance in the time dimension and the energy entropy in the frequency dimension of the time-frequency energy matrix are calculated to generate the time-frequency stability index for each spatial orientation interval. The time-frequency stability index values are marked for each spatial orientation interval in the energy accumulation distribution.
[0024] Obtain the spatial orientation angle range corresponding to the preset teaching sound source spatial region, compare the energy accumulation distribution after labeling the time-frequency stability index value with the preset teaching sound source spatial region in terms of spatial position, and identify the spatial orientation intervals in the energy accumulation distribution that are outside the spatial orientation angle range and whose time-frequency stability index value is lower than the preset stability threshold as the orientation of the interference source.
[0025] The spatial orientation identifier and time-frequency stability index values of the interference source are extracted to obtain the spatial positioning result of the interference source.
[0026] Based on the spatial location results of the interference source, a set of time-frequency points belonging to the location of the interference source is extracted from the time-frequency spatial distribution data. The time-frequency evolution trajectory of the interference source is established. Based on the time-frequency evolution trajectory of the interference source, the time-frequency points in the time-frequency spectrum data are dynamically classified and labeled to obtain time-frequency classification labels, including:
[0027] Based on the spatial orientation identifier in the spatial location result of the interference source, retrieve the time and frequency points that match the spatial orientation identifier in the time and frequency spatial distribution data to obtain the set of time and frequency points belonging to the orientation of the interference source.
[0028] The time-frequency points in the time-frequency point set are sorted by time coordinates. The frequency coordinate difference between adjacent time-frequency points is calculated. Adjacent time-frequency points with frequency coordinate differences less than a preset continuity threshold are connected in sequence to form a time-frequency evolution path. The frequency change rate of each time-frequency evolution path is counted. All time-frequency evolution paths are summarized to establish the time-frequency evolution trajectory of the interference source.
[0029] Extract the termination frequency coordinates and frequency change rate of each time-frequency evolution path in the time-frequency evolution trajectory of the interference source, calculate the frequency extension, add the termination frequency coordinates and the frequency extension to obtain the predicted frequency coordinates, locate the time-frequency point corresponding to the predicted frequency coordinates in the time-frequency data and mark it as the predicted interference time-frequency point.
[0030] Extract the time-frequency points covered by the time-frequency evolution trajectory of the interference source as historical interference time-frequency points, merge the historical interference time-frequency points with the predicted interference time-frequency points, mark the merged time-frequency points as interference time-frequency points, and mark the remaining time-frequency points as valid time-frequency points to obtain time-frequency classification labels.
[0031] Based on the time-frequency classification labels and the time-varying characteristics of the time-frequency evolution trajectory of the interference source, time-varying filter coefficients synchronized with the evolution of the interference source are generated. These time-varying filter coefficients are then applied to the time-spectrum data to adaptively attenuate the interference frequency band, resulting in filtered time-spectrum data including:
[0032] The frequency change rate of the time-frequency evolution path within a continuous time window is extracted from the time-frequency evolution trajectory of the interference source as a time-varying feature. Time-frequency evolution paths whose time-varying features exceed a preset rate threshold are marked as abrupt interference paths.
[0033] The frequency coordinates of the interference time-frequency points within the time window corresponding to the abrupt interference path are extracted from the time-frequency classification labels. The frequency drift corresponding to the frequency coordinates is calculated based on the time-varying characteristics of the abrupt interference path. The frequency drift is then superimposed on the frequency coordinates to obtain the predicted interference frequency range.
[0034] Extract the first amplitude value sequence of time-frequency points within the predicted interference frequency range and the second amplitude value sequence of interference time-frequency points on the abrupt interference path. Calculate the correlation value between the first amplitude value sequence and the second amplitude value sequence. Based on the correlation value, select time-frequency points with consistent amplitude evolution trends to form the confirmed interference coverage area.
[0035] Extract the spectral feature differences between time-frequency points within the confirmed interference coverage area and adjacent time-frequency points outside the boundary, set the suppression intensity for time-frequency points within the confirmed interference coverage area based on the spectral feature differences, and generate time-varying filter coefficients that are synchronized with the evolution of the interference source;
[0036] Multiply the time-varying filter coefficients by the complex values of the corresponding time-frequency points in the time-spectrum data to obtain the filtered time-spectrum data.
[0037] Performing an inverse time-frequency transform on the filtered time-spectrum data to generate the interference-free audio output signal includes:
[0038] Extract the complex values of each time-frequency point from the filtered time-frequency data, and organize the complex values into a two-dimensional complex matrix according to the frequency dimension and the time dimension.
[0039] Perform an inverse Fourier transform on the two-dimensional complex matrix along the frequency dimension to convert the frequency domain complex values into time domain real values, generating multiple time-domain audio frames.
[0040] Extract the length of the overlapping region between adjacent temporal audio frames, and assign weighting coefficients to each sample point in the overlapping region according to the length of the overlapping region for weighted superposition.
[0041] The weighted overlapping region sample points are then concatenated with the non-overlapping region sample points of each time-domain audio frame in chronological order to generate the audio output signal after interference elimination.
[0042] A second aspect of the present invention provides a complex sound field live audio anti-interference optimization system, comprising:
[0043] An audio acquisition unit is used to acquire multi-channel audio data from a live teaching scenario, and to perform time-frequency transformation on the multi-channel audio data to obtain time-spectrum data.
[0044] The time-frequency conversion unit is used to calculate the phase difference of each channel in the time-frequency data, determine the spatial orientation angle of each time-frequency point based on the phase difference, and obtain the time-frequency spatial distribution data.
[0045] The spatial positioning unit is used to statistically analyze the cumulative energy distribution of each spatial orientation interval in the time-frequency spatial distribution data, compare the cumulative energy distribution with the preset teaching sound source spatial area, identify the orientation of the interference source outside the teaching sound source spatial area, and obtain the spatial positioning result of the interference source.
[0046] The interference identification unit is used to extract the set of time-frequency points belonging to the location of the interference source in the time-frequency spatial distribution data based on the spatial positioning result of the interference source, establish the time-frequency evolution trajectory of the interference source, and dynamically classify and label the time-frequency points in the time-frequency spectrum data according to the time-frequency evolution trajectory of the interference source to obtain time-frequency classification labels.
[0047] The trajectory analysis unit is used to generate time-varying filter coefficients that are synchronized with the evolution of the interference source based on the time-frequency classification label and the time-varying characteristics of the time-frequency evolution trajectory of the interference source. The time-varying filter coefficients are then applied to the time-spectrum data to adaptively attenuate the interference frequency band, resulting in filtered time-spectrum data.
[0048] The time-frequency marking unit is used to perform inverse time-frequency transformation on the filtered time-spectrum data to generate an audio output signal after interference elimination.
[0049] A third aspect of the present invention provides an electronic device, comprising:
[0050] processor;
[0051] Memory used to store processor-executable instructions;
[0052] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0053] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0054] In this embodiment, based on time-frequency spatial distribution data, by statistically analyzing the cumulative energy distribution in different spatial orientation intervals and comparing it with a preset teaching sound source area, interference sources from outside the teaching area can be effectively identified and located. This method not only determines the existence of interference but also accurately outputs the spatial orientation of the interference source, achieving a leap from energy statistics to spatial positioning, and accurately distinguishing the time-frequency components of the main teaching sound source and the interference source. This step ensures the dynamism and accuracy of interference identification, enabling the tracking of changes in the interference sound. It can achieve adaptive and precise attenuation of the interference frequency band, synchronized with the evolution of the interference. This filtering method avoids the damage to the main teaching sound source caused by traditional fixed filtering, achieving a balance between interference suppression and speech fidelity.
[0055] Finally, the time-frequency spectrum after adaptive attenuation processing is inversely transformed to output a clean audio signal free of interference. The entire process forms a complete closed loop from spatial perception, interference localization, trajectory tracking to dynamic filtering, realizing real-time and intelligent suppression of sudden and mobile interference noise during live teaching, greatly improving the remote learning experience. Attached Figure Description
[0056] Figure 1 This is a flowchart illustrating the method for optimizing live audio anti-interference in complex sound fields according to an embodiment of the present invention.
[0057] Figure 2 This is a flowchart illustrating the spatial identification and localization process for audio interference sources in an embodiment of the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0059] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0060] Figure 1 This is a flowchart illustrating the method for optimizing live audio anti-interference in complex sound fields according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0061] Collect multi-channel audio data from a live teaching session, and perform time-frequency transformation on the multi-channel audio data to obtain time-spectrum data;
[0062] The phase difference of each channel in the time-frequency data is calculated, and the spatial orientation angle of each time-frequency point is determined based on the phase difference to obtain the time-frequency spatial distribution data.
[0063] The cumulative energy distribution of each spatial orientation interval in the statistical time-frequency spatial distribution data is compared with the preset teaching sound source spatial area to identify the orientation of the interference source outside the teaching sound source spatial area and obtain the spatial positioning result of the interference source.
[0064] Based on the spatial location results of the interference source, the set of time and frequency points belonging to the location of the interference source is extracted from the time and frequency spatial distribution data, the time and frequency evolution trajectory of the interference source is established, and the time and frequency points in the time and frequency spectrum data are dynamically classified and labeled according to the time and frequency evolution trajectory of the interference source to obtain time and frequency classification labels.
[0065] Based on the time-frequency classification labels and the time-varying characteristics of the time-frequency evolution trajectory of the interference source, time-varying filter coefficients that are synchronized with the evolution of the interference source are generated. The time-varying filter coefficients are applied to the time-spectrum data to adaptively attenuate the interference frequency band, and the filtered time-spectrum data is obtained.
[0066] The filtered time-frequency data is subjected to inverse time-frequency transformation to generate an audio output signal after interference is eliminated.
[0067] Collect multi-channel audio data from a live teaching session, perform time-frequency transformation on the multi-channel audio data, and obtain time-spectrum data including:
[0068] The audio signals of the teaching live broadcast scene are collected synchronously by multiple spatially distributed audio acquisition units. The spatial position coordinates of each audio acquisition unit are recorded, and the correspondence between the audio acquisition unit and the spatial position coordinates is established to obtain multi-channel audio data with spatial position identifiers.
[0069] The multi-channel audio data is subjected to windowing and frame segmentation processing to divide the continuous audio data into multiple audio frames that overlap in time. Short-time Fourier transform is performed on each audio frame to obtain the frequency domain complex spectrum corresponding to each audio frame.
[0070] The amplitude and phase values of each frequency point in the frequency domain complex spectrum are extracted, and the amplitude and phase values are fused with the time position, frequency position, and spatial position coordinates of the corresponding audio frame and the channel phase space data to obtain time-frequency spectrum data containing time-frequency amplitude, time-frequency phase, and spatial position information.
[0071] In live-streaming teaching scenarios, a multi-channel audio acquisition system consists of at least three audio acquisition units, which are deployed in a uniform or non-uniform distribution within the live-streaming space. The audio acquisition units can employ omnidirectional condenser microphones or directional dynamic microphones, with a sampling rate set to 48kHz and a quantization depth of 24 bits. During system initialization, the three-dimensional coordinates (x, y, y) of each audio acquisition unit relative to the spatial reference origin are obtained using a ranging device or a preset configuration. i y i , z i ), where i represents the audio acquisition unit number. Spatial coordinate information is written into the metadata field of the audio data, and a mapping table between channel index and spatial location is established to ensure that subsequent processing steps can directly retrieve the corresponding spatial location based on the channel number.
[0072] The multi-channel audio data acquisition process employs a hardware clock synchronization mechanism, with each acquisition unit sharing a unified time base and clock deviation controlled within 1 microsecond. The acquired raw audio data is stored in a multi-channel interleaved format, with each sampling point containing synchronized sample values from all channels. When performing windowing and framing operations on the raw audio data, the frame length is chosen to be 2048 sampling points, corresponding to a time length of approximately 42.7 milliseconds. The overlap rate between adjacent audio frames is set to 50%, i.e., the frame shift is 1024 sampling points. The windowing function uses either a Hanning window or a Hamming window, with the window function coefficients pre-calculated and stored in a lookup table. The window function is multiplied point-by-point by the audio frame data to suppress spectral leakage effects at frame boundaries.
[0073] A 2048-point short-time Fourier transform is performed on each windowed audio frame, and a fast Fourier transform algorithm is used to improve computational efficiency. The transform yields a complex spectrum with 1025 frequency points, with a frequency resolution of approximately 23.4 Hz. Each frequency point in the frequency domain complex spectrum consists of a real part and an imaginary part. The amplitude value is obtained by calculating the square root of the sum of the squares of the real and imaginary parts, and the phase value is obtained by calculating the ratio of the real to the imaginary part using the arctangent function. The phase value range is normalized to the [-π, π] interval. For frequency points with small amplitude values, the reliability of the phase value is low; therefore, frequency points with amplitude values below a preset threshold are marked, and their weight is reduced in subsequent spatial localization stages.
[0074] The extracted amplitude and phase values are organized into a three-dimensional data structure according to time, frequency, and channel dimensions. The time dimension corresponds to the audio frame number, the frequency dimension to the frequency point index, and the channel dimension to the audio acquisition unit number. Amplitude and phase data for a specific time frame, frequency point, and channel are accessed through index operations. The audio frame number is multiplied by the frame shift and sampling period to convert it into a physical time label. The frequency point index is multiplied by the frequency resolution to convert it into a physical frequency label. The mapping table is retrieved based on the channel number to obtain the corresponding spatial coordinates. A complete description of the time-spectrum data is constructed; each data unit contains a five-tuple structure of time label, frequency label, amplitude value, phase value, and spatial coordinates, providing basic data support for subsequent spatial orientation calculations and interference source localization.
[0075] Calculate the phase difference of each channel in the time-frequency data, determine the spatial orientation angle of each time-frequency point based on the phase difference, and obtain the time-frequency spatial distribution data including:
[0076] The complex spectrum values of each channel at each time frequency point are extracted from the time spectrum data, and the phase is calculated from the complex spectrum values to obtain the phase angle values of each channel at each time frequency point.
[0077] Select a channel with known spatial coordinates as the phase reference channel, calculate the difference in phase angle values between each non-reference channel and the phase reference channel at the same time frequency point, and obtain the inter-channel phase difference data including the time frequency point identifier and the phase difference value;
[0078] The phase difference values at each time frequency point in the phase difference data between channels are extracted. Combined with the spatial baseline distance between audio acquisition units and the sound wave propagation speed, the propagation path difference of the sound wave from the sound source to different audio acquisition units is calculated. Based on the spatial geometric position relationship between the propagation path difference and the audio acquisition unit array, the spatial incident angle value of the sound wave relative to the normal of the audio acquisition unit array is calculated.
[0079] By associating and binding the spatial incident angle values of each time and frequency point with the time and frequency point identifiers in the inter-channel phase difference data, a correspondence between the time and frequency point identifiers and the spatial incident angle values is established, thus obtaining the time and frequency spatial distribution data.
[0080] Calculating the phase difference of each channel in time-frequency data first requires extracting the complex spectral values of each channel at each time-frequency point from the time-frequency data. These complex spectral values typically exist in the form of real and imaginary parts, and can be represented as a + bi, where a is the real part and b is the imaginary part. When performing phase calculation on these complex spectral values, the arctangent function is used to calculate the complex phase angle, i.e., θ = arctan(b / a). It is important to note that to avoid calculation errors due to the real part being zero, the two-parameter arctangent function arctan2(b, a) should be used for calculation. In this way, the phase angle values of each channel at each time-frequency point can be obtained, and these values are typically in the range of -π to π. In practical applications, when performing a 2048-point short-time Fourier transform on an audio signal with a sampling rate of 48kHz, 1025 complex spectral values can be obtained, each containing phase information.
[0081] After obtaining the phase angle values at each time frequency point of each channel, a channel with known spatial coordinates needs to be selected as the phase reference channel. The selection of the reference channel should consider its stability and representativeness in the array; typically, the channel at the center of the array can be chosen as the reference. Calculating the difference between the phase angle values of each non-reference channel and the phase reference channel at the same time frequency point yields the inter-channel phase difference data, which includes the time frequency point identifier and the phase difference value. When calculating the phase difference, since the phase angle values are in the range of -π to π, if the result of subtracting two phase angles exceeds this range, phase folding processing is required to ensure that the phase difference value remains within the range of -π to π. For a four-channel microphone array, if the first channel is used as the reference, the phase differences of the second, third, and fourth channels relative to the first channel need to be calculated, forming three sets of inter-channel phase difference data.
[0082] After extracting the phase difference values at various time-frequency points from the inter-channel phase difference data, and combining this with the spatial baseline distance between the audio acquisition units and the speed of sound propagation, the propagation path difference of the sound wave from the sound source to different audio acquisition units can be calculated. Assuming the speed of sound in air is 343 m / s, and the distance between the two microphones is 0.1 m, if the phase difference measured at a certain time-frequency point is 0.5π radians and the frequency is 1000 Hz, then the propagation path difference can be calculated as 0.5π × 343 / (2π × 1000) ≈ 0.08575 meters. Based on the spatial geometric relationship between the propagation path difference and the audio acquisition unit array, the spatial incident angle of the sound wave relative to the normal of the audio acquisition unit array can be calculated. For a linear microphone array, if the distance between the two microphones is d, the phase difference is φ, the frequency is f, and the speed of sound is c, then the incident angle θ of the sound wave can be calculated using the formula sinθ = φc / (2πfd).
[0083] The spatial incident angle values at each time-frequency point are associated and bound to the time-frequency point identifiers in the inter-channel phase difference data to establish a correspondence between time-frequency point identifiers and spatial incident angle values, thereby obtaining the time-frequency spatial distribution data. The time-frequency spatial distribution data can be represented as a three-dimensional matrix, where two dimensions represent time and frequency, and the third dimension represents the spatial incident angle. For each time-frequency point, there is a corresponding spatial incident angle value, which can be calculated through the aforementioned steps. To improve calculation accuracy, various spatial scanning algorithms can be used, such as the minimum variance distortionless response algorithm or the multi-signal classification algorithm, to accurately estimate the spatial incident angle. In practical applications, to reduce computational load, angle calculations can be performed only on time-frequency points with energy exceeding a preset threshold, while a default angle value can be set for time-frequency points with lower energy, or they can be ignored directly.
[0084] To reduce noise interference during phase difference calculations, the time-spectrum data can be preprocessed, such as by applying noise suppression techniques like spectral subtraction or Wiener filtering, to improve the signal-to-noise ratio. Outliers in the phase difference data can be smoothed using methods like median filtering or weighted averaging. When calculating the spatial incident angle, the complexity of the sound field, such as the effects of reflection and diffraction on sound wave propagation, should be considered, and an appropriate sound field model should be used for angle correction. When multiple sound sources exist simultaneously, clustering algorithms can be used to group the incident angles at different time-frequency points, identifying the time-frequency distribution characteristics of different sound sources.
[0085] In constructing time-frequency spatial distribution data, different processing strategies can be adopted for different frequency bands. For example, for low-frequency bands, the smoothness of angle estimation can be increased, while for high-frequency bands, the focus is on the accuracy of angle estimation. Simultaneously, for time-varying sound field environments, adaptive algorithms can be used to dynamically adjust phase difference calculations and angle estimations to adapt to environmental changes. Regarding data storage, to reduce storage requirements, time-frequency spatial distribution data can be compressed, for example, retaining only the spatial distribution information of time-frequency points with significant energy, or quantizing similar angle values.
[0086] In this embodiment, by acquiring time-frequency spatial distribution data, sound sources from different directions can be effectively distinguished, the signal of the target sound source can be enhanced, and interference signals can be suppressed, thereby improving the quality of live audio in complex sound field environments. This method does not depend on a specific signal model, is applicable to various sound field environments, and has strong robustness.
[0087] The cumulative energy distribution of each spatial location interval in the statistical time-frequency spatial distribution data is compared with the preset teaching sound source spatial area to identify the location of interference sources outside the teaching sound source spatial area, and the spatial localization results of the interference sources are obtained, including:
[0088] Extract the spatial incident angle value and time spectrum amplitude value of each time frequency point from the time frequency spatial distribution data. Divide the spatial incident angle value into multiple spatial azimuth intervals according to the preset angle interval. Accumulate the time spectrum amplitude value of all time frequency points in each spatial azimuth interval and statistically obtain the cumulative energy distribution of each spatial azimuth interval.
[0089] A time-frequency energy matrix is constructed for the energy accumulation distribution of each spatial orientation interval. The energy variance in the time dimension and the energy entropy in the frequency dimension of the time-frequency energy matrix are calculated to generate the time-frequency stability index for each spatial orientation interval. The time-frequency stability index values are marked for each spatial orientation interval in the energy accumulation distribution.
[0090] Obtain the spatial orientation angle range corresponding to the preset teaching sound source spatial region, compare the energy accumulation distribution after labeling the time-frequency stability index value with the preset teaching sound source spatial region in terms of spatial position, and identify the spatial orientation intervals in the energy accumulation distribution that are outside the spatial orientation angle range and whose time-frequency stability index value is lower than the preset stability threshold as the orientation of the interference source.
[0091] The spatial orientation identifier and time-frequency stability index values of the interference source are extracted to obtain the spatial positioning result of the interference source.
[0092] like Figure 2 The diagram shown illustrates the flowchart for spatial identification and localization of audio interference sources in this embodiment.
[0093] Extracting the spatial incident angle and time-frequency amplitude values for each time-frequency point from time-frequency spatial distribution data is fundamental to spatial localization of interference sources. Time-frequency spatial distribution data typically contains information in three dimensions: time, frequency, and space. Each time-frequency point corresponds to a spatial incident angle and a time-frequency amplitude value. The spatial incident angle represents the angle of incidence of the sound wave propagation direction relative to the microphone array, usually expressed in degrees; the time-frequency amplitude value represents the energy level at that time-frequency point, usually expressed in decibels or linear amplitude. When processing time-frequency spatial distribution data, the 360-degree space can be divided into multiple spatial azimuth intervals according to preset angle intervals. The preset angle intervals can be set according to actual application requirements, such as 10 degrees or 5 degrees. When the space is divided into 10-degree intervals, the entire space will be divided into 36 azimuth intervals, namely [0°, 10°), [10°, 20°), ..., [350°, 360°). For each spatial azimuth interval, the cumulative energy distribution of each spatial azimuth interval is obtained by accumulating the time-frequency amplitude values of all time-frequency points within that interval. If there are 100 time-frequency points in the interval [0°, 10°) and the time-frequency amplitude of each time-frequency point is 0.1, then the cumulative energy value of this interval is 10.
[0094] A time-frequency energy matrix is constructed based on the cumulative energy distribution across spatial azimuth intervals to evaluate the time-frequency characteristics of sound sources in each azimuth. The time-frequency energy matrix is a two-dimensional matrix where rows represent time and columns represent frequency. Each element represents the energy level of its corresponding time-frequency point within a specific spatial azimuth interval. The temporal stability of the sound source can be assessed by calculating the energy variance of the time-frequency energy matrix in the time dimension. The energy variance is calculated as: Time Dimension Variance = Σ(Eij - μi)² / T, where Eij represents the energy value at the i-th frequency and j-th time point, μi represents the average energy value of the i-th frequency in the time dimension, and T represents the number of time frames. Simultaneously, the spectral characteristics of the sound source can be evaluated by calculating the energy entropy value in the frequency dimension. The energy entropy is calculated as: Frequency Dimension Entropy = -Σ(pj × log(pj)), where pj represents the proportion of energy at the j-th frequency to the total energy, calculated as pj = Ej / ΣEj, where Ej represents the energy value at the j-th frequency. The time-frequency stability index, composed of time-dimensional variance and frequency-dimensional entropy, is used to determine whether a sound source in a certain direction is an interference source. Generally, teaching sound sources exhibit higher time-frequency stability, while interference sources show lower stability. Therefore, labeling the time-frequency stability index values for each spatial location interval in the energy accumulation distribution can serve as an important basis for identifying interference sources.
[0095] Obtaining the spatial azimuth angle range corresponding to the preset teaching sound source spatial region is a prerequisite for interference source identification. In teaching scenarios, teaching sound sources are usually located in fixed spatial positions, such as the teacher's podium area or the student's speaking area. By measuring or setting in advance, the spatial azimuth angle range corresponding to these areas can be determined. For example, the spatial azimuth angle range corresponding to the teacher's podium area may be [40°, 60°], and the spatial azimuth angle range corresponding to the student's speaking area may be [100°, 160°]. By comparing the spatial positions of these preset teaching sound source spatial regions with the energy accumulation distribution marked with time-frequency stability index values, spatial azimuth intervals located outside the teaching sound source spatial regions and with time-frequency stability index values below the preset stability threshold can be identified. These intervals are likely to be the locations of interference sources.
[0096] The preset stability threshold can be adjusted according to the actual application scenario. Generally, the time dimension variance threshold can be set to 20% of the total energy, and the frequency dimension entropy threshold can be set to 0.8. When the time dimension variance of a certain spatial orientation interval is greater than the threshold or the frequency dimension entropy is less than the threshold, that interval is identified as a potential source of interference.
[0097] The final step is to extract the spatial location identifier and time-frequency stability index values of the interference source to obtain the spatial localization result. The spatial location identifier includes the azimuth range of the interference source, such as [260°, 270°]; the time-frequency stability index values include the time dimension variance and frequency dimension entropy values of this azimuth range, such as a time dimension variance of 0.35 and a frequency dimension entropy of 0.6. This information together constitutes the spatial localization result of the interference source, providing accurate spatial location information for subsequent interference suppression.
[0098] In practical applications, the spatial localization of interference sources may be affected by environmental noise and reflections. To improve localization accuracy, several optimization measures can be taken. Temporal smoothing of the time-frequency energy matrix can reduce the impact of instantaneous noise and improve the stability of interference source localization. Temporal smoothing can employ methods such as moving average or exponential smoothing, with a smoothing window length of 5-10 frames. Simultaneously, spatial smoothing of the energy accumulation distribution can reduce localization errors caused by insufficient spatial resolution. Spatial smoothing can employ methods such as Gaussian smoothing or median smoothing, with a smoothing window size of 3-5 azimuth intervals.
[0099] For situations where multiple interference sources coexist, energy peak detection and cluster analysis can be used for multi-source identification. First, all energy peaks are detected in the cumulative energy distribution; each peak may correspond to a sound source. Then, cluster analysis is performed on these peaks, grouping peaks with similar spatial locations into one category, with each category representing a possible sound source. Finally, based on the time-frequency stability index values of each category, it is determined whether it is an interference source. Energy peak detection can employ a local maximum detection algorithm, and cluster analysis can use a distance-based hierarchical clustering method, with a clustering distance threshold set to 15 degrees.
[0100] In complex acoustic environments, the characteristics of interference sources may change over time. To adapt to these changes, adaptive threshold adjustment technology can be employed. By monitoring the time-frequency stability index of the teaching audio source in real time, the preset stability threshold is dynamically adjusted, making interference source identification more accurate. The adaptive threshold adjustment can be calculated based on the average time-frequency stability index value of the teaching audio source within a sliding time window, with the window length set to 30 seconds and the adjustment step size to 0.05.
[0101] Based on the spatial location results of the interference source, a set of time-frequency points belonging to the location of the interference source is extracted from the time-frequency spatial distribution data. The time-frequency evolution trajectory of the interference source is established. Based on the time-frequency evolution trajectory of the interference source, the time-frequency points in the time-frequency spectrum data are dynamically classified and labeled to obtain time-frequency classification labels, including:
[0102] Based on the spatial orientation identifier in the spatial location result of the interference source, retrieve the time and frequency points that match the spatial orientation identifier in the time and frequency spatial distribution data to obtain the set of time and frequency points belonging to the orientation of the interference source.
[0103] The time-frequency points in the time-frequency point set are sorted by time coordinates. The frequency coordinate difference between adjacent time-frequency points is calculated. Adjacent time-frequency points with frequency coordinate differences less than a preset continuity threshold are connected in sequence to form a time-frequency evolution path. The frequency change rate of each time-frequency evolution path is counted. All time-frequency evolution paths are summarized to establish the time-frequency evolution trajectory of the interference source.
[0104] Extract the termination frequency coordinates and frequency change rate of each time-frequency evolution path in the time-frequency evolution trajectory of the interference source, calculate the frequency extension, add the termination frequency coordinates and the frequency extension to obtain the predicted frequency coordinates, locate the time-frequency point corresponding to the predicted frequency coordinates in the time-frequency data and mark it as the predicted interference time-frequency point.
[0105] Extract the time-frequency points covered by the time-frequency evolution trajectory of the interference source as historical interference time-frequency points, merge the historical interference time-frequency points with the predicted interference time-frequency points, mark the merged time-frequency points as interference time-frequency points, and mark the remaining time-frequency points as valid time-frequency points to obtain time-frequency classification labels.
[0106] Based on the spatial azimuth marker in the spatial localization results of the interference source, the process of retrieving time-frequency points matching the spatial azimuth marker in the time-frequency spatial distribution data requires precise localization of the time-frequency point corresponding to the azimuth of the interference source. The spatial azimuth marker typically includes the azimuth interval of the interference source, such as [260°, 270°]. During the retrieval process, each time-frequency point in the time-frequency spatial distribution data is traversed, and it is determined whether its spatial incident angle falls within the azimuth interval of the interference source. If the spatial incident angle of a time-frequency point is 265°, and the azimuth interval of the interference source is [260°, 270°], then that time-frequency point is determined to belong to the azimuth of the interference source. Considering the potential angle estimation error in practical applications, a tolerance range, such as ±5°, can be introduced to expand the retrieval range and improve the robustness of the retrieval.
[0107] By searching, a series of time-frequency points belonging to the location of the interference source can be obtained. Each time-frequency point contains time coordinates, frequency coordinates, and energy value information, such as (2.5s, 1000Hz, 0.8), which indicates that sound from the location of the interference source was detected at 2.5 seconds and 1000Hz, with an energy value of 0.8. These time-frequency points constitute the set of time-frequency points for the location of the interference source, providing basic data for subsequent analysis of the time-frequency evolution characteristics of the interference source.
[0108] Sort the time-frequency points in the time-frequency point set by time coordinates; this is the first step in establishing the time-frequency evolution trajectory of the interference source. Efficient algorithms such as quicksort or mergesort can be used for sorting to ensure that the time coordinates are arranged in ascending order. After sorting, the frequency coordinate difference between adjacent time-frequency points is calculated to determine whether they are continuous. The formula for calculating the frequency coordinate difference is: difference = |f(t+Δt)-f(t)|, where f(t) represents the frequency coordinate at time t, and f(t+Δt) represents the frequency coordinate at time t+Δt. If the frequency coordinate difference is less than a preset continuity threshold, the two time-frequency points are considered to be continuous in frequency and may belong to the same time-frequency evolution path of the interference source. The preset continuity threshold can be set according to the audio sampling rate and time-frequency analysis parameters. For example, for a short-time Fourier transform parameter with a sampling rate of 48kHz, a frame length of 2048, and a frame shift of 512, the continuity threshold can be set to 50Hz. Adjacent time-frequency points with frequency coordinate differences less than the preset continuity threshold are connected sequentially to form the time-frequency evolution path. A time-frequency evolution path can be represented as a series of time-frequency points ordered by time, such as {(2.5s, 1000Hz), (2.6s, 1050Hz), (2.7s, 1080Hz)}, indicating that the frequency of the interference source changes from 1000Hz at 2.5 seconds to 1080Hz at 2.7 seconds.
[0109] Statistical analysis of the frequency change rate along each time-frequency evolution path is an important method for quantifying the frequency variation characteristics of interference sources. The formula for calculating the frequency change rate is: Change rate = (f(t)) / (t) n )-f(t1)) / (tn -t1), where f(t1) and f(t) n ) represent the frequency coordinates of the starting and ending points of the path, respectively, t n -t1 indicates the time span of the path. The rate of change of frequency is measured in Hz / s, representing how quickly the frequency changes over time. Positive values indicate a frequency increase, negative values indicate a frequency decrease, and the magnitude of the value indicates the rate of change.
[0110] If a time-frequency evolution path starts at (2.5s, 1000Hz) and ends at (2.7s, 1080Hz), then its frequency change rate is (1080-1000) / (2.7-2.5) = 400Hz / s, indicating that the frequency of the interference source increases at a rate of 400Hz / s. Summarizing all time-frequency evolution paths, including their time-frequency point sequences and frequency change rates, constitutes the time-frequency evolution trajectory of the interference source, comprehensively describing its behavioral characteristics in the time-frequency domain.
[0111] By extracting the termination frequency coordinates and frequency change rate of each time-frequency evolution path in the time-frequency evolution trajectory of the interference source, the frequency distribution of the interference source at future times can be predicted, providing a basis for active interference suppression. The frequency extension is calculated based on the frequency change rate and the prediction time window. The formula is: Extension = Change Rate × Prediction Window, where the prediction window represents the length of time to the future, usually set to 10-50 ms. The termination frequency coordinates are added to the frequency extension to obtain the predicted frequency coordinates. For example, if the termination point of a time-frequency evolution path is (2.7s, 1080Hz), the frequency change rate is 400Hz / s, and the prediction window is 20ms, then the frequency extension is 400 × 0.02 = 8Hz, and the predicted frequency coordinates are 1080 + 8 = 1088Hz. The time-frequency point corresponding to the predicted frequency coordinates is located in the time-frequency data, that is, the time-frequency point with the time of 2.7s + 0.02 = 2.72s and the frequency closest to 1088Hz is found and marked as the predicted interference time-frequency point. Predicting the frequency points of interference represents the frequency location where the interference source may appear in the near future, providing the possibility of predictive operation for interference suppression.
[0112] The time-frequency points covered by the time-frequency evolution trajectory of the interference source are extracted as historical interference time-frequency points, recording the time-frequency regions already affected by the interference source. These historical interference time-frequency points are merged with the predicted interference time-frequency points to comprehensively mark the historical and potential impacts of the interference source, improving the coverage and accuracy of interference identification. The merged time-frequency points are marked as interference time-frequency points, indicating that these points are affected by the interference source and require special attention or suppression in subsequent processing. The remaining time-frequency points are marked as valid time-frequency points, indicating that these points mainly contain information from valid sound sources and should be preserved or enhanced.
[0113] This classification labeling provides a precise time-frequency classification tag, which is essential for subsequent interference suppression and sound quality enhancement.
[0114] In practical applications, the time-frequency characteristics of interference sources can be complex and variable. To improve the accuracy of establishing time-frequency evolution trajectories, a multi-path tracking strategy can be adopted. When multiple possible frequency change trends are detected, multiple time-frequency evolution paths can be established simultaneously, and weights can be assigned based on characteristics such as energy intensity and continuity of the paths, prioritizing the path with higher weight for prediction. Multi-path tracking can be implemented using algorithms such as Kalman filtering or particle filtering to improve the tracking capability for nonlinear frequency changes. Simultaneously, for possible frequency jumps from the interference source, a jump detection mechanism can be set. When the frequency difference between adjacent moments exceeds a jump threshold (e.g., 100Hz) but the energy continuity is strong, a frequency jump is considered to have occurred, and a new time-frequency evolution path is established.
[0115] To address the issue of energy fluctuations in interference sources, an energy screening mechanism can be introduced when establishing the time-frequency evolution trajectory. This mechanism selects only time-frequency points with energy exceeding a threshold for trajectory establishment, reducing interference from low-energy noise points. The energy threshold can be set to twice the average energy of the time-frequency spectrum, or adaptively adjusted based on the signal-to-noise ratio. For cases where the time-frequency evolution path is discontinuous, interpolation techniques can be used to fill in the missing time-frequency points, maintaining path continuity. Linear interpolation or cubic spline interpolation can be selected as the interpolation method, with the appropriate method chosen based on the smoothness of the interference source's frequency variation.
[0116] In this embodiment, by establishing the time-frequency evolution trajectory of the interference source and dynamically marking time-frequency points, accurate identification and tracking of interference sources in complex sound fields are achieved. Based on the spatial localization results of the interference source, a set of time-frequency points is extracted. A time-frequency evolution path is constructed through time-frequency continuity analysis, thereby predicting the frequency change trend of the interference source and providing a foundation for accurate time-frequency localization for interference suppression. Adapting to the dynamic characteristics of the interference source, it can effectively handle interference sources with changing frequencies, significantly improving the anti-interference capability and sound quality of live audio in complex sound fields.
[0117] Based on the time-frequency classification labels and the time-varying characteristics of the time-frequency evolution trajectory of the interference source, time-varying filter coefficients synchronized with the evolution of the interference source are generated. These time-varying filter coefficients are then applied to the time-spectrum data to adaptively attenuate the interference frequency band, resulting in filtered time-spectrum data including:
[0118] The frequency change rate of the time-frequency evolution path within a continuous time window is extracted from the time-frequency evolution trajectory of the interference source as a time-varying feature. Time-frequency evolution paths whose time-varying features exceed a preset rate threshold are marked as abrupt interference paths.
[0119] The frequency coordinates of the interference time-frequency points within the time window corresponding to the abrupt interference path are extracted from the time-frequency classification labels. The frequency drift corresponding to the frequency coordinates is calculated based on the time-varying characteristics of the abrupt interference path. The frequency drift is then superimposed on the frequency coordinates to obtain the predicted interference frequency range.
[0120] Extract the first amplitude value sequence of time-frequency points within the predicted interference frequency range and the second amplitude value sequence of interference time-frequency points on the abrupt interference path. Calculate the correlation value between the first amplitude value sequence and the second amplitude value sequence. Based on the correlation value, select time-frequency points with consistent amplitude evolution trends to form the confirmed interference coverage area.
[0121] Extract the spectral feature differences between time-frequency points within the confirmed interference coverage area and adjacent time-frequency points outside the boundary, set the suppression intensity for time-frequency points within the confirmed interference coverage area based on the spectral feature differences, and generate time-varying filter coefficients that are synchronized with the evolution of the interference source;
[0122] Multiply the time-varying filter coefficients by the complex values of the corresponding time-frequency points in the time-spectrum data to obtain the filtered time-spectrum data.
[0123] Extracting the frequency change rate of the time-frequency evolution path within a continuous time window from the time-frequency evolution trajectory of the interference source as a time-varying feature is the foundation for identifying abrupt interference paths. A time-frequency evolution path typically contains a series of time-frequency points, each with time and frequency coordinates. The frequency change rate is calculated as follows: within a continuous time window, select a start time point t1 and an end time point t2, extract the corresponding frequency coordinates f1 and f2, and calculate the frequency change rate as (f2-f1) / (t2-t1), in Hz / s. The time window length is determined based on the audio processing frame length, typically 20-50ms. When the absolute value of the frequency change rate exceeds a preset rate threshold, the time-frequency evolution path segment is marked as an abrupt interference path. The preset rate threshold setting needs to consider the frequency change characteristics of normal speech and music signals, and is typically set to 2000-5000Hz / s. Setting the rate threshold too high may prevent some interference paths from being identified, while setting it too low may misjudge normal audio as interference. In a live teaching scenario, if an interference path is detected where the frequency changes from 1000Hz to 1300Hz within 30ms, the calculated frequency change rate is 10000Hz / s, which significantly exceeds the preset rate threshold of 5000Hz / s. Therefore, this path is marked as an abrupt interference path.
[0124] Extracting the frequency coordinates of the interfering time-frequency points within the time window corresponding to the abrupt interference path from the time-frequency classification labels is a crucial step in determining the interference frequency range. The time-frequency classification labels contain marking information for both the interfering and valid time-frequency points. By filtering the interfering time-frequency points within the time window corresponding to the abrupt interference path, the frequency distribution of the interference band can be obtained. The frequency drift is calculated based on the time-varying characteristics of the abrupt interference path. Considering the continuous change in the interference source frequency, the frequency drift calculation formula is: Drift = Frequency Change Rate × Prediction Time, where the prediction time represents the length of time to predict into the future, typically 10-30 ms. The calculated frequency drift is then superimposed onto the frequency coordinates of the current interfering time-frequency point to obtain the predicted interference frequency range. If the current interfering time-frequency point frequency is 1300 Hz, the frequency change rate is 10000 Hz / s, and the prediction time is 20 ms, then the frequency drift is 10000 × 0.02 = 200 Hz, and the predicted interference frequency range is [1300 Hz, 1500 Hz]. In practical applications, a frequency spreading factor can be introduced to appropriately extend the prediction range and enhance the adaptability to irregular frequency changes. The spreading factor is usually taken as 1.2-1.5.
[0125] The first amplitude value sequence of time-frequency points within the predicted interference frequency range and the second amplitude value sequence of interference time-frequency points on the abrupt interference path are extracted. By comparing the correlation between the two, the interference coverage area can be further confirmed. The first amplitude value sequence contains the amplitude values of all time-frequency points within the predicted interference frequency range, arranged in chronological order; the second amplitude value sequence contains the amplitude values of the confirmed interference time-frequency points on the abrupt interference path, also arranged in chronological order. The correlation value between the two sequences is calculated. Commonly used correlation calculation methods include Pearson correlation coefficient and cosine similarity. The formula for calculating the Pearson correlation coefficient is: Correlation coefficient = Covariance(Sequence 1, Sequence 2) / (Standard deviation(Sequence 1) × Standard deviation(Sequence 2)), with a value range of [-1, 1]. The closer the value is to 1, the stronger the positive correlation. Time-frequency points with consistent amplitude evolution trends are selected based on the correlation values. The correlation coefficient threshold is usually set to 0.7-0.8. When the correlation coefficient is greater than the threshold, the time-frequency point is considered to have a similar amplitude evolution trend to the known interference point and is likely to belong to the same interference source. Therefore, it is included in the confirmed interference coverage area. If the correlation coefficient between a certain time-frequency sequence and a known interference sequence is calculated to be 0.85, which exceeds the threshold of 0.8, then the time-frequency point corresponding to the sequence is confirmed as part of the interference coverage area.
[0126] Extracting the spectral feature differences between time-frequency points within the confirmed interference coverage area and adjacent time-frequency points outside the boundary provides a basis for setting the suppression intensity. Spectral feature differences can be quantified by calculating indicators such as spectral envelope shape, energy distribution, and phase continuity. Specific methods include calculating the amplitude ratio of time-frequency points inside and outside the interference area, and calculating the spectral gradient of adjacent frequency points. Based on the spectral feature differences, a suppression intensity is set for the time-frequency points within the confirmed interference coverage area. The suppression intensity is directly proportional to the spectral feature differences; the greater the difference, the greater the suppression intensity. The suppression intensity calculation formula is: Suppression Intensity = Basic Suppression Coefficient × (1 + α × Spectral Feature Difference), where the basic suppression coefficient is usually set to 0.1-0.3, and α is an adjustment parameter with a value range of 0.5-2. The generated time-varying filter coefficient is actually an attenuation factor, whose value is equal to 1 - suppression intensity, with a value range of [0, 1], where 0 represents complete suppression and 1 represents no suppression. If the calculated difference in spectral characteristics at a certain interference time-frequency point is 0.5, the basic suppression coefficient is 0.2, and α is 1, then the suppression strength is 0.2×(1+1×0.5)=0.3, and the corresponding time-varying filter coefficient is 1-0.3=0.7.
[0127] The time-varying filter coefficients are multiplied by the complex values of the corresponding time-frequency points in the time-spectrum data to obtain the filtered time-spectrum data. Each time-frequency point in the time-spectrum data corresponds to a complex value containing amplitude and phase information. The time-varying filter coefficients only adjust the amplitude of the complex value, preserving the original phase information to reduce phase distortion introduced by the processing.
[0128] If the original complex value is a + bi, its modulus is sqrt(a² + b²), its phase angle is arctan(b / a), and the time-varying filter coefficient is h, then the filtered complex value is h × sqrt(a² + b²) × (cos(arctan(b / a)) + i × sin(arctan(b / a))). Applying the corresponding time-varying filter coefficients to all interference time-frequency points yields the filtered time-spectrum data, achieving precise attenuation of the interference frequency band while preserving the effective audio signal.
[0129] In practical applications, to improve the smoothness and naturalness of time-varying filters, a smoothing transition technique for filter coefficients can be employed. This involves smoothing the time-varying filter coefficients in both the time and frequency dimensions to avoid abrupt changes and artifacts introduced by the filtering operation. Time-dimensional smoothing can use exponential smoothing: smoothed coefficient = α × current coefficient + (1-α) × previous frame coefficient, where α is the smoothing factor, typically between 0.2 and 0.5. Frequency-dimensional smoothing can use Gaussian smoothing, which involves a weighted average of the filter coefficients at adjacent frequency points. The weights are distributed according to a Gaussian function, and the standard deviation is typically set to a width of 2-5 frequency points.
[0130] For situations with strong interference amplitude, a nonlinear suppression strategy can be introduced to dynamically adjust the suppression intensity based on the interference amplitude. The relationship between suppression intensity and interference amplitude can be set as a nonlinear mapping, such as suppression intensity = min(maximum suppression intensity, β × (interference amplitude / reference amplitude)). γ The nonlinear mapping typically uses β, which is usually 0.2-0.5, and γ, which is 1.5-2.5, with a maximum suppression strength of 0.9-0.95. This nonlinear mapping ensures sufficient suppression of strong interferences while maintaining moderate suppression of weak interferences, avoiding over-processing.
[0131] In this embodiment, the time-frequency variation characteristics of the interference source can be tracked in real time, and time-varying filter coefficients synchronized with the evolution of the interference source can be generated to achieve precise location and suppression of the interference signal. By analyzing the frequency jump characteristics of the interference source, predicting the interference frequency drift, and verifying the consistency of amplitude evolution trends, the accuracy of interference identification is improved. Adaptive suppression strength setting ensures effective suppression of interference while preserving effective audio information to the maximum extent and reducing processing distortion.
[0132] Performing an inverse time-frequency transform on the filtered time-spectrum data to generate the interference-free audio output signal includes:
[0133] Extract the complex values of each time-frequency point from the filtered time-frequency data, and organize the complex values into a two-dimensional complex matrix according to the frequency dimension and the time dimension.
[0134] Perform an inverse Fourier transform on the two-dimensional complex matrix along the frequency dimension to convert the frequency domain complex values into time domain real values, generating multiple time-domain audio frames.
[0135] Extract the length of the overlapping region between adjacent temporal audio frames, and assign weighting coefficients to each sample point in the overlapping region according to the length of the overlapping region for weighted superposition.
[0136] The weighted overlapping region sample points are then concatenated with the non-overlapping region sample points of each time-domain audio frame in chronological order to generate the audio output signal after interference elimination.
[0137] After adaptive filtering of the time-spectrum data, it is necessary to restore the processed frequency domain data to a time-domain audio signal. Time-spectrum data typically exists in complex form, containing both amplitude and phase information. These complex values need to be organized logically according to the frequency and time dimensions to form a two-dimensional complex matrix structure. Specifically, the rows of the matrix represent different frequency components, the columns represent different time frames, and each matrix element represents the complex value of a specific frequency at a specific time frame. For a typical audio signal, the frequency dimension may contain hundreds to thousands of frequency points, while the time dimension depends on the audio duration and frame shift. In practical applications, this complex matrix can be represented as R(f, t) + jI(f, t), where R represents the real part, I represents the imaginary part, f represents the frequency index, and t represents the time frame index.
[0138] After obtaining the two-dimensional complex matrix, an inverse Fourier transform needs to be performed along the frequency dimension to convert the complex values in the frequency domain into real values in the time domain. This transformation process is implemented using the inverse fast Fourier transform algorithm, processing each column of the matrix (representing all frequency components of a time frame) separately. For each time frame t, the complex values of all frequency points corresponding to that time frame are taken to form a complex vector V(f,t). Then, an inverse Fourier transform is performed on V(f,t) to obtain the corresponding time-domain audio frame A(n,t), where n represents the index of the time-domain sample point. The calculation of the inverse Fourier transform can be expressed as: A(n,t) = IFFT{V(f,t)}. Since audio signals are usually real signals, after performing the inverse transform on the spectrum of the real signal, only the real part of the result should be taken as the time-domain signal. Each time-domain audio frame obtained after the transformation contains the same number of sample points as the input frame length, usually several hundred to several thousand sample points.
[0139] Overlapping regions typically exist between adjacent temporal audio frames. This is due to the use of frame overlay techniques during time-frequency analysis to improve temporal resolution. The length of the overlapping region depends on the frame shift parameter set during the initial time-frequency analysis and can be represented by an overlap factor. For example, when the overlap factor is 75%, the number of overlapping sample points between adjacent frames is 75% of the frame length. To avoid discontinuities caused by direct splicing, the sample points within the overlapping region need to be weighted and overlaid. Commonly used weighted window functions include triangular windows, Hanning windows, or cosine windows. For a sample point at position i within the overlapping region, if linearly gradient weights are used, the weight coefficient can be expressed as w(i) = i / L, where L is the length of the overlapping region. If the overlapping region sample points of frame t and frame t+1 are A(n, t) and A(n, t+1) respectively, then the weighted and overlaid sample point value is: S(i) = (1 - w(i)) × A(i, t) + w(i) × A(i, t+1), where i represents the index of the sample point within the overlapping region.
[0140] After weighted superposition of the overlapping regions, the non-overlapping sample points of each temporal audio frame are concatenated with the weighted superposition of the overlapping sample points in chronological order. The concatenation process requires precise calculation of the position of each sample point in the final output signal. Assuming each frame has a length of L and a frame shift of S, the first (LS) sample points of frame t constitute the non-overlapping region of that frame, and the last S sample points, along with the first S sample points of frame t+1, constitute the overlapping region. During concatenation, the first (LS) sample points of frame t are directly placed into their corresponding positions in the output signal, followed by the S sample points of the weighted superposition of the overlapping region. For the last frame, all its sample points must be included in the output signal. This method ensures a smooth and natural transition between frames, avoiding discontinuities at the concatenation points, thus generating a complete, interference-free audio output signal.
[0141] In this embodiment, precise inverse time-frequency transformation and smooth frame stitching technology effectively eliminate various interferences in complex sound field live streaming environments, significantly improving the clarity and intelligibility of the audio signal. This method preserves the effective information of the original audio while maximally suppressing background noise, reverberation, and other non-target sound sources, making the target sound stand out more in noisy environments. Seamless frame stitching technology avoids artificial breaks and discontinuities in the audio signal, ensuring a natural and smooth output audio.
[0142] A second aspect of the present invention provides a complex sound field live audio anti-interference optimization system, the system comprising:
[0143] An audio acquisition unit is used to acquire multi-channel audio data from a live teaching scenario, and to perform time-frequency transformation on the multi-channel audio data to obtain time-spectrum data.
[0144] The time-frequency conversion unit is used to calculate the phase difference of each channel in the time-frequency data, determine the spatial orientation angle of each time-frequency point based on the phase difference, and obtain the time-frequency spatial distribution data.
[0145] The spatial positioning unit is used to statistically analyze the cumulative energy distribution of each spatial orientation interval in the time-frequency spatial distribution data, compare the cumulative energy distribution with the preset teaching sound source spatial area, identify the orientation of the interference source outside the teaching sound source spatial area, and obtain the spatial positioning result of the interference source.
[0146] The interference identification unit is used to extract the set of time-frequency points belonging to the location of the interference source in the time-frequency spatial distribution data based on the spatial positioning result of the interference source, establish the time-frequency evolution trajectory of the interference source, and dynamically classify and label the time-frequency points in the time-frequency spectrum data according to the time-frequency evolution trajectory of the interference source to obtain time-frequency classification labels.
[0147] The trajectory analysis unit is used to generate time-varying filter coefficients that are synchronized with the evolution of the interference source based on the time-frequency classification label and the time-varying characteristics of the time-frequency evolution trajectory of the interference source. The time-varying filter coefficients are then applied to the time-spectrum data to adaptively attenuate the interference frequency band, resulting in filtered time-spectrum data.
[0148] The time-frequency marking unit is used to perform inverse time-frequency transformation on the filtered time-spectrum data to generate an audio output signal after interference elimination.
[0149] A third aspect of the present invention provides an electronic device, comprising:
[0150] processor;
[0151] Memory used to store processor-executable instructions;
[0152] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0153] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0154] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing live audio playback in complex sound fields to prevent interference, characterized in that, include: Collect multi-channel audio data from a live teaching session, and perform time-frequency transformation on the multi-channel audio data to obtain time-spectrum data; The phase difference of each channel in the time-frequency data is calculated, and the spatial orientation angle of each time-frequency point is determined based on the phase difference to obtain the time-frequency spatial distribution data. The cumulative energy distribution of each spatial orientation interval in the statistical time-frequency spatial distribution data is compared with the preset teaching sound source spatial area to identify the orientation of the interference source outside the teaching sound source spatial area and obtain the spatial positioning result of the interference source. Based on the spatial location results of the interference source, the set of time and frequency points belonging to the location of the interference source is extracted from the time and frequency spatial distribution data, the time and frequency evolution trajectory of the interference source is established, and the time and frequency points in the time and frequency spectrum data are dynamically classified and labeled according to the time and frequency evolution trajectory of the interference source to obtain time and frequency classification labels. Based on the time-frequency classification labels and the time-varying characteristics of the time-frequency evolution trajectory of the interference source, time-varying filter coefficients that are synchronized with the evolution of the interference source are generated. The time-varying filter coefficients are applied to the time-spectrum data to adaptively attenuate the interference frequency band, and the filtered time-spectrum data is obtained. The filtered time-frequency data is subjected to inverse time-frequency transformation to generate an audio output signal after interference is eliminated.
2. The method according to claim 1, characterized in that, Collect multi-channel audio data from a live teaching session, perform time-frequency transformation on the multi-channel audio data, and obtain time-spectrum data including: The audio signals of the teaching live broadcast scene are collected synchronously by multiple spatially distributed audio acquisition units. The spatial position coordinates of each audio acquisition unit are recorded, and the correspondence between the audio acquisition unit and the spatial position coordinates is established to obtain multi-channel audio data with spatial position identifiers. The multi-channel audio data is subjected to windowing and frame segmentation processing to divide the continuous audio data into multiple audio frames that overlap in time. Short-time Fourier transform is performed on each audio frame to obtain the frequency domain complex spectrum corresponding to each audio frame. The amplitude and phase values of each frequency point in the frequency domain complex spectrum are extracted, and the amplitude and phase values are fused with the time position, frequency position, and spatial position coordinates of the corresponding audio frame and the channel phase space data to obtain time-frequency spectrum data containing time-frequency amplitude, time-frequency phase, and spatial position information.
3. The method according to claim 1, characterized in that, Calculate the phase difference of each channel in the time-frequency data, determine the spatial orientation angle of each time-frequency point based on the phase difference, and obtain the time-frequency spatial distribution data including: The complex spectrum values of each channel at each time frequency point are extracted from the time spectrum data, and the phase is calculated from the complex spectrum values to obtain the phase angle values of each channel at each time frequency point. Select a channel with known spatial coordinates as the phase reference channel, calculate the difference in phase angle values between each non-reference channel and the phase reference channel at the same time frequency point, and obtain the inter-channel phase difference data including the time frequency point identifier and the phase difference value; The phase difference values at each time frequency point in the phase difference data between channels are extracted. Combined with the spatial baseline distance between audio acquisition units and the sound wave propagation speed, the propagation path difference of the sound wave from the sound source to different audio acquisition units is calculated. Based on the spatial geometric position relationship between the propagation path difference and the audio acquisition unit array, the spatial incident angle value of the sound wave relative to the normal of the audio acquisition unit array is calculated. By associating and binding the spatial incident angle values of each time and frequency point with the time and frequency point identifiers in the inter-channel phase difference data, a correspondence between the time and frequency point identifiers and the spatial incident angle values is established, thus obtaining the time and frequency spatial distribution data.
4. The method according to claim 1, characterized in that, The cumulative energy distribution of each spatial location interval in the statistical time-frequency spatial distribution data is compared with the preset teaching sound source spatial area to identify the location of interference sources outside the teaching sound source spatial area, and the spatial localization results of the interference sources are obtained, including: Extract the spatial incident angle value and time spectrum amplitude value of each time frequency point from the time frequency spatial distribution data. Divide the spatial incident angle value into multiple spatial azimuth intervals according to the preset angle interval. Accumulate the time spectrum amplitude value of all time frequency points in each spatial azimuth interval and statistically obtain the cumulative energy distribution of each spatial azimuth interval. A time-frequency energy matrix is constructed for the energy accumulation distribution of each spatial orientation interval. The energy variance in the time dimension and the energy entropy in the frequency dimension of the time-frequency energy matrix are calculated to generate the time-frequency stability index for each spatial orientation interval. The time-frequency stability index values are marked for each spatial orientation interval in the energy accumulation distribution. Obtain the spatial orientation angle range corresponding to the preset teaching sound source spatial region, compare the energy accumulation distribution after labeling the time-frequency stability index value with the preset teaching sound source spatial region in terms of spatial position, and identify the spatial orientation intervals in the energy accumulation distribution that are outside the spatial orientation angle range and whose time-frequency stability index value is lower than the preset stability threshold as the orientation of the interference source. The spatial orientation identifier and time-frequency stability index values of the interference source are extracted to obtain the spatial positioning result of the interference source.
5. The method according to claim 1, characterized in that, Based on the spatial location results of the interference source, a set of time-frequency points belonging to the location of the interference source is extracted from the time-frequency spatial distribution data. The time-frequency evolution trajectory of the interference source is established. Based on the time-frequency evolution trajectory of the interference source, the time-frequency points in the time-frequency spectrum data are dynamically classified and labeled to obtain time-frequency classification labels, including: Based on the spatial orientation identifier in the spatial location result of the interference source, retrieve the time and frequency points that match the spatial orientation identifier in the time and frequency spatial distribution data to obtain the set of time and frequency points belonging to the orientation of the interference source. The time-frequency points in the time-frequency point set are sorted by time coordinates. The frequency coordinate difference between adjacent time-frequency points is calculated. Adjacent time-frequency points with frequency coordinate differences less than a preset continuity threshold are connected in sequence to form a time-frequency evolution path. The frequency change rate of each time-frequency evolution path is counted. All time-frequency evolution paths are summarized to establish the time-frequency evolution trajectory of the interference source. Extract the termination frequency coordinates and frequency change rate of each time-frequency evolution path in the time-frequency evolution trajectory of the interference source, calculate the frequency extension, add the termination frequency coordinates and the frequency extension to obtain the predicted frequency coordinates, locate the time-frequency point corresponding to the predicted frequency coordinates in the time-frequency data and mark it as the predicted interference time-frequency point. Extract the time-frequency points covered by the time-frequency evolution trajectory of the interference source as historical interference time-frequency points, merge the historical interference time-frequency points with the predicted interference time-frequency points, mark the merged time-frequency points as interference time-frequency points, and mark the remaining time-frequency points as valid time-frequency points to obtain time-frequency classification labels.
6. The method according to claim 1, characterized in that, Based on the time-frequency classification labels and the time-varying characteristics of the time-frequency evolution trajectory of the interference source, time-varying filter coefficients synchronized with the evolution of the interference source are generated. These time-varying filter coefficients are then applied to the time-spectrum data to adaptively attenuate the interference frequency band, resulting in filtered time-spectrum data including: The frequency change rate of the time-frequency evolution path within a continuous time window is extracted from the time-frequency evolution trajectory of the interference source as a time-varying feature. Time-frequency evolution paths whose time-varying features exceed a preset rate threshold are marked as abrupt interference paths. The frequency coordinates of the interference time-frequency points within the time window corresponding to the abrupt interference path are extracted from the time-frequency classification labels. The frequency drift corresponding to the frequency coordinates is calculated based on the time-varying characteristics of the abrupt interference path. The frequency drift is then superimposed on the frequency coordinates to obtain the predicted interference frequency range. Extract the first amplitude value sequence of time-frequency points within the predicted interference frequency range and the second amplitude value sequence of interference time-frequency points on the abrupt interference path. Calculate the correlation value between the first amplitude value sequence and the second amplitude value sequence. Based on the correlation value, select time-frequency points with consistent amplitude evolution trends to form the confirmed interference coverage area. Extract the spectral feature differences between time-frequency points within the confirmed interference coverage area and adjacent time-frequency points outside the boundary, set the suppression intensity for time-frequency points within the confirmed interference coverage area based on the spectral feature differences, and generate time-varying filter coefficients that are synchronized with the evolution of the interference source; Multiply the time-varying filter coefficients by the complex values of the corresponding time-frequency points in the time-spectrum data to obtain the filtered time-spectrum data.
7. The method according to claim 1, characterized in that, Performing an inverse time-frequency transform on the filtered time-spectrum data to generate the interference-free audio output signal includes: Extract the complex values of each time-frequency point from the filtered time-frequency data, and organize the complex values into a two-dimensional complex matrix according to the frequency dimension and the time dimension. Perform an inverse Fourier transform on the two-dimensional complex matrix along the frequency dimension to convert the frequency domain complex values into time domain real values, generating multiple time-domain audio frames. Extract the length of the overlapping region between adjacent temporal audio frames, and assign weighting coefficients to each sample point in the overlapping region according to the length of the overlapping region for weighted superposition. The weighted overlapping region sample points are then concatenated with the non-overlapping region sample points of each time-domain audio frame in chronological order to generate the audio output signal after interference elimination.
8. A complex sound field live audio anti-interference optimization system, used to implement the method of any one of claims 1-7, characterized in that, include: An audio acquisition unit is used to acquire multi-channel audio data from a live teaching scenario, and to perform time-frequency transformation on the multi-channel audio data to obtain time-spectrum data. The time-frequency conversion unit is used to calculate the phase difference of each channel in the time-frequency data, determine the spatial orientation angle of each time-frequency point based on the phase difference, and obtain the time-frequency spatial distribution data. The spatial positioning unit is used to statistically analyze the cumulative energy distribution of each spatial orientation interval in the time-frequency spatial distribution data, compare the cumulative energy distribution with the preset teaching sound source spatial area, identify the orientation of the interference source outside the teaching sound source spatial area, and obtain the spatial positioning result of the interference source. The interference identification unit is used to extract the set of time-frequency points belonging to the location of the interference source in the time-frequency spatial distribution data based on the spatial positioning result of the interference source, establish the time-frequency evolution trajectory of the interference source, and dynamically classify and label the time-frequency points in the time-frequency spectrum data according to the time-frequency evolution trajectory of the interference source to obtain time-frequency classification labels. The trajectory analysis unit is used to generate time-varying filter coefficients that are synchronized with the evolution of the interference source based on the time-frequency classification label and the time-varying characteristics of the time-frequency evolution trajectory of the interference source. The time-varying filter coefficients are then applied to the time-spectrum data to adaptively attenuate the interference frequency band, resulting in filtered time-spectrum data. The time-frequency marking unit is used to perform inverse time-frequency transformation on the filtered time-spectrum data to generate an audio output signal after interference elimination.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.