Electric power scene audio enhancement method and system based on deep learning

By using frame-segmented short-time Fourier transform and deep learning network processing, frequency domain suppression weight maps and transient restoration maps are generated, solving the problem of suppressing power frequency harmonic hum and transient impulse noise in power scenarios, achieving effective speech enhancement, and maintaining speech naturalness and auditory coherence.

CN122024752APending Publication Date: 2026-05-12STATE GRID GANSU ELECTRIC POWER CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID GANSU ELECTRIC POWER CORP
Filing Date
2026-01-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing deep learning-based audio enhancement methods struggle to effectively suppress power frequency harmonic hum and transient switching impulse noise in power scenarios, while maintaining the naturalness and auditory coherence of speech, especially when low-frequency steady-state pure tones coexist with transient impulse noise.

Method used

After employing frame-by-frame short-time Fourier transform, the frequency domain suppression weight map and transient indicator map are generated through parallel processing of the harmonic suppression branch network and the transient repair branch network. The enhanced spectrum is then generated by dynamically fusing the weights, and finally, the enhanced time-domain audio is obtained through inverse short-time Fourier transform.

Benefits of technology

It effectively suppresses power frequency harmonic noise and transient impulse noise, reduces the attenuation of the fundamental frequency and low-order harmonics of speech, reduces noise residue and trailing effects, and improves speech fidelity and auditory coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024752A_ABST
    Figure CN122024752A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power scene audio enhancement method and system based on deep learning, and relates to the technical field of audio signal processing and voice enhancement. Candidate harmonic frequency bands are constructed around power grid power frequency and multi-order harmonic waves of the power grid power frequency, and a continuous soft suppression weight map is generated; a smooth recess is formed in a harmonic frequency band, and energy coherence is kept in a non-harmonic frequency band, so that the perceptibility of harmonic buzzing is reduced, and meanwhile, accidental damage to voice low-frequency information is reduced; besides, the candidate frequency band bandwidth is adaptively determined according to the spectrum peak broadening characteristic, the local noise floor and the frequency resolution, and the network outputs the adjustment amount of the suppression intensity, the bandwidth and the center position, so that the suppression range and intensity are updated along with the change of the intra-frame spectrum shape, and the adaptability to different stations and different equipment noise forms is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing and speech enhancement technology, and in particular to a method and system for audio enhancement in power scenarios based on deep learning. Background Technology

[0002] Remote conferencing for power dispatching / station control is a crucial communication scenario for daily dispatching, emergency command, and collaborative operations in the power industry. However, in conferencing / duty environments within substation control rooms or adjacent equipment areas, voice recording is easily affected by industrial noise. Common types of background noise include: (1) The continuous low-frequency humming sound caused by the magnetostriction and structural vibration of the iron core of equipment such as power transformers and reactors usually exhibits a discrete spectrum dominated by the second harmonic of the power frequency and its harmonics (and may be accompanied by the power frequency and other low-frequency components), with obvious low-frequency and pure tone characteristics. (2) Transient impact noise caused by the opening and closing of circuit breakers / disconnecting switches, relay protection and related actuators (the operation of control panel buttons / cabinet may also be superimposed indoors). This type of noise usually has the characteristics of instantaneous high energy and wide spectrum.

[0003] Existing technologies mostly employ deep learning-based speech / audio enhancement methods that perform masking estimation or mapping reconstruction in the time and frequency domains. Some schemes also use a dual-branch structure in the time and frequency domains to fuse different representations. For the aforementioned persistent low-frequency pure tone interference, if a strong low-frequency global suppression or fixed masking strategy is adopted, it may weaken the fundamental frequency and its low-order harmonic components while reducing buzzing, thus affecting the naturalness and fullness of the speech. For transient impulse noise, enhancement methods based on time-frequency masking estimation may produce residual noise and processing artifacts (such as time spread / tailing, musical noise, etc.) in some cases, affecting auditory coherence and intelligibility.

[0004] Some solutions introduce attention mechanisms or gated loop structures to model long-range dependencies and temporal dynamics. However, in power scenarios, where low-frequency steady-state pure tones and transient impulse noise coexist and their energy and statistical characteristics differ significantly, there is still room for further optimization in how to achieve adaptive suppression of the two types of noise while maintaining speech fidelity. Summary of the Invention

[0005] In view of the aforementioned existing problems, the present invention is proposed.

[0006] This invention provides a deep learning-based audio enhancement method and system for power scenarios to solve the problem that existing enhancements are prone to damaging low-frequency speech and producing trailing distortion in power scenarios containing strong power frequency harmonic hum and transient switching pulses.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide a deep learning-based audio enhancement method for power scenarios, comprising: Step S1: Segment the noisy time-domain audio into frames and perform a short-time Fourier transform to obtain the complex time spectrum; Step S2: Input the complex time-spectrum parallel harmonic suppression branch network and the transient repair branch network; The harmonic suppression branch network generates a continuous frequency domain suppression weight map based on the preset or estimated power grid frequency and its harmonic frequencies, and weights the complex time spectrum to suppress steady-state harmonic noise to obtain the first enhancement spectrum. The transient repair branch network outputs a transient indicator map, and the time-frequency region marked in the transient indicator map is used by a gated convolutional repair network to generate a replacement spectrum based on the neighborhood context to obtain a second enhanced spectrum; Step S3: Calculate the dynamic fusion weights based on the confidence levels output by the two branches, and perform weighted fusion of the first enhancement spectrum and the second enhancement spectrum to obtain the fusion spectrum; perform inverse short-time Fourier transform on the fusion spectrum to obtain the enhanced time-domain audio.

[0008] As a preferred embodiment of the deep learning-based audio enhancement method for power scenarios described in this invention, the estimation of the power grid frequency includes: performing peak search or autocorrelation analysis on the amplitude spectrum of the complex time spectrum within a preset power frequency candidate range to obtain power frequency candidate values, and selecting the one with the largest energy as the power grid frequency.

[0009] As a preferred embodiment of the deep learning-based audio enhancement method for power scenarios described in this invention, the generation of the continuous frequency domain suppression weights includes: determining candidate harmonic frequency bands with the power grid frequency and its harmonic frequencies as the center, and using a neural network to adaptively determine the frequency band boundary and suppression intensity based on the spectral characteristics inside and outside the candidate harmonic frequency bands, and outputting a frequency domain suppression weight map with continuous values. During the generation of the continuous frequency domain suppression weight map, when constructing candidate harmonic frequency bands around the power frequency and its multi-order harmonic centers, the bandwidth of the candidate harmonic frequency bands is adaptively determined based on the broadening characteristics of the harmonic spectral peaks, the local noise floor, and the frequency resolution. Within the candidate harmonic frequency bands, a continuously transitioning soft suppression window is used to form a basic soft mask, and the neural network outputs adaptive adjustment amounts for adjusting the suppression intensity, bandwidth, and center position based on the spectral characteristics inside and outside the candidate harmonic frequency bands, resulting in a frequency domain suppression weight map that continuously changes at the frequency band boundaries. At the same time, a minimum retention constraint is set for the low-frequency band where the speech fundamental frequency is located, and constraints are applied to the change amplitude of the frequency domain suppression weights.

[0010] As a preferred embodiment of the deep learning-based audio enhancement method for power scenarios described in this invention, the transient indicator map is a continuous value map with the same resolution as the complex time-frequency spectrum, representing the confidence level of each time-frequency unit as transient impulse noise, and the transient region is obtained by thresholding.

[0011] As a preferred embodiment of the deep learning-based audio enhancement method for power scenarios described in this invention, the gated convolutional inpainting network uses the neighborhood time-frequency features of the transient region and its corresponding region mask as conditional inputs, and regulates the contribution of neighborhood information to the inpainting result through a gating mechanism, outputting a replacement spectrum for filling the transient region. The gated convolutional inpainting network generates candidate inpainting features through main branch convolution and generates a threshold through gated branches in each network layer. The threshold is used to modulate the candidate inpainting features point by point to control the contribution of neighborhood information to the inpainting result. The transient region mask is updated between network layers according to the threshold, so that the region to be repaired gradually shrinks as the number of layers increases. In the output stage, a continuous transition mask is constructed based on the transient region mask, and the replacement spectrum and the original time spectrum are continuously fused together accordingly.

[0012] As a preferred embodiment of the deep learning-based audio enhancement method for power scenarios described in this invention, the confidence level includes: a harmonic residual metric of the output of the harmonic suppression branch and a transient region identification and determination metric of the output of the transient repair branch.

[0013] As a preferred embodiment of the deep learning-based audio enhancement method for power scenarios described in this invention, the dynamic fusion weights are calculated based on time frames or time-frequency units and obtained by normalization mapping of the confidence scores, so that branches with higher confidence scores occupy higher fusion weights at the corresponding positions.

[0014] Secondly, the present invention provides a deep learning-based audio enhancement system for power scenarios, comprising: The time-frequency conversion module is used to convert noisy time-domain audio into a complex time-frequency spectrum; A harmonic suppression branch network is used to generate frequency domain suppression weights and output the first enhancement spectrum; A transient repair branch network is used to generate a transient indication map and output a second enhancement spectrum; The fusion module is used to calculate dynamic fusion weights based on confidence levels and output the fusion spectrum; The reconstruction module is used to perform an inverse short-time Fourier transform on the fused spectrum to output enhanced time-domain audio.

[0015] As a preferred embodiment of the deep learning-based audio enhancement system for power scenarios described in this invention, the reconstruction module employs an overlapping and additive time-domain synthesis method when performing inverse short-time Fourier transform, and uses the phase information corresponding to the fusion spectrum or the phase information output by the phase recovery network.

[0016] As a preferred embodiment of the deep learning-based power scene audio enhancement system of the present invention, the harmonic suppression branch network, transient repair branch network and fusion module are deployed on the same computing device or edge computing device.

[0017] Through the above technical solution, the present invention can achieve at least the following beneficial effects: To address the issues of overlapping frequency bands between steady-state power frequency harmonic noise and low-frequency components of speech, and the weakening of speech fundamental frequency and low-order harmonics caused by traditional global low-frequency suppression, this invention constructs candidate harmonic frequency bands around the power grid frequency and its multiple harmonics and generates a continuous soft suppression weight map. This creates a smooth concave shape within the harmonic frequency band while maintaining energy continuity in the non-harmonic frequency band, reducing the perceptibility of harmonic humming and minimizing false damage to low-frequency speech information.

[0018] To address the problem of steady-state harmonic spectral peaks broadening and drifting with varying operating conditions, loads, and structures, leading to mismatches in fixed bandstops or fixed masking boundaries, this invention adaptively determines candidate frequency band bandwidth based on spectral peak broadening characteristics, local noise floor, and frequency resolution. The network outputs adjustments to the suppression intensity, bandwidth, and center position, allowing the suppression range and intensity to update with changes in intra-frame spectral shape, thus improving adaptability to noise patterns at different sites and on different devices.

[0019] To address the issues of high energy, wide frequency spectrum, and easy misjudgment as speech transients and their retention, or the presence of residual trails after suppression that affect auditory coherence, this invention locates transient regions using a transient indicator map and uses gated convolution within these regions to repair and generate replacement spectra. The abnormal region content is then supplemented with neighborhood context, shifting the processing from attenuation to repair and replacement, thereby reducing transient noise residue and mitigating the sense of trails after it disappears.

[0020] To address the problem that the repair process easily leads to repeated referencing of abnormal components in transient regions, causing the repair impact to spread to non-transient regions and introducing boundary abrupt changes, this invention updates the transient region mask between network layers according to the threshold, causing the region to be repaired to shrink layer by layer. In the output stage, a continuous transition fusion is constructed based on the mask, thereby achieving limited repair range, smooth boundary connection, and reduced edge ringing and audible trailing.

[0021] To address the inconsistent reliability of steady-state harmonic suppression and transient restoration across different time frames and frequency bands, the difficulty of achieving both suppression depth and speech fidelity with a single path, and the potential for amplifying one type of interference while suppressing another, this invention uses a confidence-driven dynamic fusion weight composed of harmonic residual measurement and transient identification measurement. This allows more reliable branches to have higher weights at corresponding time frames or time-frequency positions, and employs conservative fusion and amplitude change rate constraints under low confidence conditions to reduce unstable artifacts and speech distortion. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.

[0023] Figure 1 This is a flowchart of the audio enhancement method in the embodiment.

[0024] Figure 2 This is a framework diagram of the audio enhancement system in the embodiment. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0026] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way. Example 1:

[0027] like Figure 1 As shown, this application proposes a deep learning-based audio enhancement method for power scenarios, including the following steps: Step S1: Segment the noisy time-domain audio into frames and perform a short-time Fourier transform to obtain the complex time spectrum; Framing involves overlapping framing of noisy time-domain audio and applying a window function; the output of the short-time Fourier transform is a complex time spectrum, which contains real and imaginary part information in each time-frequency unit, or equivalently contains amplitude and phase information; the amplitude spectrum corresponding to the complex time spectrum is obtained from the complex modulus, and the phase is obtained from the complex argument.

[0028] In this embodiment, the noisy time-domain audio is a single-channel digital audio stream collected by a remote conferencing terminal and recorded in chronological order. The sampling rate is set to 16000 as an implementation parameter, with an allowable range of 8000 to 48000. The frame length is set to 32ms as an implementation parameter, with an allowable range of 20ms to 40ms. The frame shift is set to 16ms as an implementation parameter, with an allowable range of 5ms to 20ms. The overlap rate is determined by the frame length and frame shift and is set to 50% by default, with an allowable range of 25% to 75%. The window function is Hanning window as an implementation parameter, and switching between Hanning window and Hamming window is allowed. The number of points in the short-time Fourier transform is set as an implementation parameter to the least power of 2 of the number of sampling points corresponding to the frame length, with a default range of 256 to 2048, to ensure a balance between frequency domain resolution and real-time performance. The time frame order of the complex time spectrum is consistent with the original audio sampling order. The timestamp is generated by accumulating the frame shift and used for frame-by-frame alignment of the subsequent two branch outputs.

[0029] Step S2: Input the complex time spectrum in parallel into the harmonic suppression branch network and the transient repair branch network; the harmonic suppression branch network and the transient repair branch network share the same complex time spectrum input, and output the network results with the same time resolution and frequency resolution as the complex time spectrum; the harmonic suppression branch network outputs a frequency domain suppression weight map and harmonic correlation statistics used to calculate the harmonic residual metric; the transient repair branch network outputs a transient indicator map, a region mask used for repair, and identification correlation statistics used to calculate the transient region identification and determination metric.

[0030] Specifically, a time frame refers to the position of the frame sequence from the first to the last frame after framing; a frequency index refers to the discrete frequency point numbering from low to high frequency after short-time Fourier transform; the region mask is uniformly referred to as the transient region mask in this embodiment, and its value is 0 or 1 with the same resolution as the complex time spectrum. A value of 1 indicates that the corresponding time-frequency unit is determined to be a transient region; in this text, the binary region mask, transient region mask, and initial transient region mask represent the state of the same mask at different steps, where the initial transient region mask refers to the transient indicator. The transient region mask is obtained by thresholding the graph; the harmonic correlation statistics are calculated in each time frame by the local band energy of the amplitude spectrum within the candidate harmonic frequency band, the reference band energy outside the candidate harmonic frequency band, and the corresponding ratio; the identification correlation statistics are calculated in each time frame by the average confidence of the transient indicator map in the transient and non-transient regions, the difference between the two, and the confidence change at the boundary. All of the above statistics share the same time frame index with the complex time spectrum and are directly used for subsequent confidence and fusion weight calculations.

[0031] The harmonic suppression branch network generates a continuous frequency domain suppression weight map based on the preset or estimated power grid frequency and its harmonic frequencies, and weights the complex time spectrum to suppress steady-state harmonic noise to obtain the first enhanced spectrum; the values ​​of the frequency domain suppression weight map are continuous real numbers and are limited to a preset closed interval; the weighting of the complex time spectrum includes applying the same weight to the real part and imaginary part of the complex spectrum, or applying weight to the amplitude spectrum while keeping the phase unchanged; the first enhanced spectrum is the complex time spectrum after harmonic suppression, and its time axis and frequency axis are consistent with the input complex time spectrum.

[0032] Furthermore, the frequency domain suppression weight map, as a continuous real number map, is used to weight the complex time-frequency spectrum unit by unit. The weights are limited to 0 to 1 by default and the allowed range is 0 to 1. To avoid local energy collapse and numerical instability caused by the weights reaching 0, the minimum retention weight is set to 0.2 as an implementation parameter, with an allowed range of 0.05 to 0.4. Values ​​below the minimum retention weight are truncated when outputting the weight map. The weighting of the complex time-frequency spectrum by default adopts the method of applying the same weight to the real part and the imaginary part. It is allowed to switch to the method of applying weight to the amplitude spectrum while keeping the phase unchanged. Both methods maintain the consistency of the output first enhanced spectrum and the input complex time-frequency spectrum on the time axis and frequency axis.

[0033] The transient repair branch network outputs a transient indicator map, and the time-frequency region marked in the transient indicator map is used by a gated convolutional repair network to generate a replacement spectrum based on the neighborhood context, thus obtaining the second enhanced spectrum; The transient indicator map has the same resolution as the complex time-frequency spectrum. Each time-frequency unit of the transient indicator map represents the confidence level that the time-frequency unit belongs to transient impulse noise. The time-frequency region is determined by a binary region mask obtained by thresholding the transient indicator map. The replacement spectrum is a complex time-frequency result with the same resolution as the complex time-frequency spectrum. Within the region indicated by the binary region mask, it is output by a gated convolutional repair network. Within the non-indicated region of the binary region mask, it retains the original content of the input complex time-frequency spectrum or the content of the first enhancement spectrum. The second enhancement spectrum is obtained by combining the replacement spectrum and the content of the non-repaired region according to the binary region mask. The binary region mask is the transient region mask.

[0034] Step S3: Calculate the dynamic fusion weights based on the confidence levels output by the two branches, and perform weighted fusion of the first enhancement spectrum and the second enhancement spectrum to obtain the fusion spectrum; perform inverse short-time Fourier transform on the fusion spectrum to obtain the enhanced time-domain audio. In this embodiment, the estimation of the power grid frequency includes: performing peak search or autocorrelation analysis on the amplitude spectrum of the complex time spectrum within a preset power frequency candidate range to obtain power frequency candidate values, and selecting the one with the largest energy as the power grid frequency; The power frequency candidate range covers the rated power frequency and allowable offset range of different power grid systems; the peak search is performed based on the set of local maxima points of the amplitude spectrum within the power frequency candidate range, and the one with the largest energy is determined by the local frequency band energy of the corresponding frequency point; the autocorrelation analysis obtains period candidates based on the cepstrum or time-domain correlation of the amplitude spectrum and maps them to frequency candidates; when the power frequency candidate values ​​obtained by the peak search and autocorrelation analysis are inconsistent, the power frequency candidate value with higher local frequency band energy is adopted as the power grid frequency; when the power frequency candidate value is lower than the preset confidence threshold, the preset rated power frequency is adopted as the power grid frequency.

[0035] For example, the power frequency candidate range, as an implementation parameter, defaults to cover 45–65, and can be narrowed to 49–51 or 59–61 when the power grid system is known; the local band energy of the peak search is obtained by accumulating the neighborhood on the amplitude spectrum with the power frequency candidate frequency as the center, and the neighborhood half bandwidth, as an implementation parameter, defaults to 3 frequency indices, with an allowed range of 1–10 frequency indices; the periodic candidate of autocorrelation analysis is obtained from the cepstral peak or equivalent correlation peak of the amplitude spectrum and then mapped to the frequency candidate, and time smoothing is performed on the power frequency candidate value of each time frame, with the smoothing length, as an implementation parameter, defaulting to 5 frames, with an allowed range of 1–20 frames; the preset confidence threshold, as an implementation parameter, defaults to 0.6, with an allowed range of 0.3–0.9, and the confidence is determined by the ratio of the local band energy of the power frequency candidate frequency to the median energy of the same candidate range and normalized to 0–1.

[0036] In this embodiment, the generation of continuous frequency domain suppression weights includes: determining candidate harmonic frequency bands with the power grid frequency and its harmonic frequencies as the center, and using a neural network to adaptively determine the frequency band boundary and suppression intensity based on the spectral characteristics inside and outside the candidate harmonic frequency bands, and outputting a frequency domain suppression weight map with continuous values. The candidate harmonic frequency band includes at least the frequency band containing the predetermined order harmonics of the power grid frequency; the initial boundary of the frequency band is determined by the peak width of the input amplitude spectrum near the center frequency of the corresponding harmonic; the neural network jointly models the local spectral morphology and global spectral energy distribution inside and outside the candidate harmonic frequency band, and outputs the boundary offset used to update the frequency band boundary and the intensity coefficient used to update the suppression intensity; the frequency domain suppression weight map has a continuous transition at the boundary of the candidate harmonic frequency band, and the transition bandwidth is determined by the bandwidth control quantity output by the neural network.

[0037] In the process of generating a continuous frequency domain suppression weight map, when constructing candidate harmonic frequency bands around the power frequency and its multi-order harmonic centers, the bandwidth of the candidate harmonic frequency bands is adaptively determined based on the broadening characteristics of the harmonic spectral peaks, the local noise floor, and the frequency resolution. Within the candidate harmonic frequency bands, a continuously transitioning soft suppression window is used to form a basic soft mask, and the neural network outputs adaptive adjustment amounts to adjust the suppression intensity, bandwidth, and center position based on the spectral characteristics inside and outside the candidate harmonic frequency bands, resulting in a frequency domain suppression weight map that continuously changes at the frequency band boundaries. At the same time, a minimum retention constraint is set for the low-frequency band where the speech fundamental frequency is located, and constraints are applied to the change amplitude of the frequency domain suppression weights to reduce the thin speech phenomenon caused by excessive suppression of low-frequency components. Similarly, the noisy complex time-frequency spectrum is used to characterize the complex spectral values ​​at each time frame and each frequency index. The first enhanced spectrum is the complex spectral value after being weighted by a continuous frequency domain suppression weight map. The sampling rate is in Hz and typically ranges from 8000 to 48000. The number of short-time Fourier transform points is an integer and typically ranges from 256 to 2048. The frequency resolution is in Hz and is determined by both the sampling rate and the number of short-time Fourier transform points. The highest harmonic order involved in suppression is set to 10 by default as an implementation parameter, with an allowable range of 3 to 30. The median statistical half-window width of the local noise floor is set to 10 frequency indices by default as an implementation parameter, with an allowable range of 3. ~30 frequency indices; the half-height scaling factor is set to 0.5 by default as an implementation parameter, with an allowable range of 0.3 to 0.8; the half-bandwidth expansion factor is set to 1.5 by default as an implementation parameter, with an allowable range of 1.0 to 3.0; the suppression strength factor is limited to 0 to 1 by default, with an allowable range of 0 to 1; the bandwidth scaling is positive and has a default range of 0.5 to 3.0; the center frequency offset is in Hz and has a default range of -2 times the frequency resolution to +2 times the frequency resolution; the continuous transition window changes continuously at the boundary of the candidate harmonic frequency band, and the boundary transition bandwidth is determined by the bandwidth scaling and the basic half-bandwidth and is variable between different time frames.

[0038] In one implementation, in frequency domain weighted suppression, the frequency domain suppression weight map with continuous values ​​can be represented as outputting continuous weights for each time frame and frequency point, and used for point-by-point weighting of the complex time spectrum, for example, for the th... Frame number Execution at each frequency point: (1) In equation (1), Indicates the spectrum of a noisy complex number in time frame With frequency index Complex values ​​at; Indicates time frame With frequency index The continuous frequency domain suppression weights at the specified points have a value range of 1. ; This indicates the weighted complex enhancement spectrum in time frame. With frequency index Complex values ​​at the location.

[0039] When forming candidate harmonic frequency bands around the power frequency and its several harmonic centers, the frequency index can be mapped to the actual frequency, and an adaptive bandwidth related to the spectral peak half-width, local noise floor, and frequency resolution can be given. For the ... The center frequency of each frequency point can be written as: (2) In equation (2), Frequency index The corresponding center frequency; Indicates frequency index; Indicates the audio sampling rate; This represents the number of frequency domain points in the short-time Fourier transform. Indicates frequency resolution.

[0040] For the estimated power frequency of the power grid and its first The center frequency of the first harmonic can be written as: (3) In equation (3), Indicates the first The center frequency of the first harmonic; Indicates the harmonic order; Indicates the power grid frequency; Indicates the first The frequency index corresponding to the center frequency of the first harmonic; This means rounding down to the nearest integer.

[0041] To correlate candidate harmonic frequency bands with spectral peak shapes, the amplitude spectrum can be adjusted. The local noise floor is estimated and the half-height threshold is determined accordingly, where: (4)

[0042] In equation (4), Represents time frame Frequency Index The amplitude spectrum at the location; Represents the modulus of a complex number; Represents time frame Frequency Index Local noise floor estimation at the location; Indicates the median operation; Indicates the frequency index participating in local statistics; This represents the half-window width of local statistics, expressed in frequency index.

[0043] Based on the half-height threshold of the spectral peak relative to the noise floor, the left and right half-height width index distances of the harmonic peaks can be given: (5)

[0044] (5.1) In equations (5) and (5.1), Represents time frame Next The minimum index distance from the first harmonic peak to the high-frequency side to reach the half-height threshold; Represents time frame Next The minimum index distance from the first harmonic peak to the low-frequency side to reach the half-maximum threshold; This represents the non-negative integer index distance used for the search; This represents the half-height scaling factor, with a value range of [value range missing]. .

[0045] Therefore, the basic half-bandwidth of the candidate harmonic frequency band is set as follows:

[0046] In equation (6), Represents time frame Next The fundamental half-bandwidth of the first harmonic candidate frequency band; This represents the half-bandwidth expansion coefficient, used to cover spectral peak broadening and window function leakage; This indicates the operation of taking the larger value.

[0047] To obtain continuous values ​​for the soft suppression weight shape, a window function with a continuous transition at the band boundary can be constructed within each harmonic frequency band, and adaptive adjustment can be achieved by the neural network outputting the intensity coefficient, bandwidth adjustment, and boundary offset. The neural network is then applied to the... Frame number The three types of adjustment quantities for the first harmonic output are denoted as follows: , , and make them satisfied. , Then a normalized frequency offset can be constructed:

[0048] In equation (7), Represents time frame Frequency Index Relative to the first Normalized frequency offset of the first harmonic center; Indicates the first The frequency offset of the first harmonic center is used to absorb power frequency estimation errors and spectral peak drift; Indicates the bandwidth scaling amount, used to scale the base half bandwidth. Make adaptive adjustments; It is used subsequently to control the intensity of inhibition.

[0049] To ensure weight continuity at frequency band boundaries, a cosine transition window can be used to generate a basic soft mask.

[0050] In equation (8), Represents time frame Frequency Index From the first The basic soft masking value for the generation of first harmonics; Represents the cosine function; Pi is a constant. This represents absolute value operations.

[0051] Then map the soft masking into continuous suppression weights:

[0052] In equation (9), Represents time frame Frequency Index For the first Continuous suppression weights for first harmonics; Indicates the first The suppression intensity coefficient of the first harmonic; the larger the value, the deeper the suppression of that harmonic frequency band.

[0053] The weights of multiple harmonics can be obtained by multiplying each harmonic sequentially to obtain a continuous weighted graph: (10)

[0054] In equation (10), Represents time frame Frequency Index The continuous suppression weight after integrating all harmonics; This represents a series of multiplication operations; Indicates the highest harmonic order involved in suppression; This represents the continuous suppression weight after adding the lower bound constraint; This represents the minimum retention weight lower bound, used to limit the energy collapse caused by excessive suppression; This indicates that the smaller value is selected during the operation.

[0055] In addition, retention constraints can be introduced near low-frequency speech components to keep the weight map at or above a given retention level in that region.

[0056] For time frames The low-frequency speech fundamental frequency estimation is denoted as and use continuous protective windows Mark its neighborhood: (11)

[0057] In equation (11), Represents time frame Frequency Index The voice protection weight at a given location, with a larger value indicating that it is closer to the voice fundamental frequency neighborhood; Represents an exponential function; Represents time frame The estimated low-frequency speech fundamental frequency; This indicates the frequency extension parameter of the voice protection window; This represents the minimum preservation weight of the voice protection region; This represents the final continuous frequency domain suppression weight map output after applying voice protection.

[0058] Incorporating the constraints of the training phase, a penalty can be added to the rate of change of the weights to limit local abrupt changes, written as:

[0059] In equation (12), The regularization term representing the rate of change of weights; Indicates the frequency index of the adjacent low-frequency side; This indicates the square operation. When this regularization term works in conjunction with the speech protection window, the weight map exhibits a continuous dip in the harmonic frequency band and a continuous rise or gradual change in the low-frequency speech region.

[0060] As can be seen, in the frequency domain weighted suppression stage, the spectrum of the noisy complex time is first weighted point by point to form an enhanced spectrum, and the point-by-point weighting relationship is shown in Equation (1); then the frequency index is mapped to the actual frequency and the frequency resolution is obtained, as shown in Equation (2); then the center of each harmonic is derived from the power frequency and its frequency index is located, as shown in Equation (3). In each time frame, a local noise floor is constructed based on the amplitude spectrum and used as a reference for spectral peak discrimination, as shown in Equation (4); then the left and right boundaries are searched with the half-height threshold of the spectral peak relative to the noise floor, so as to obtain the bandwidth information corresponding to the spectral peak broadening, as shown in Equations (5) and (5.1), and this bandwidth and frequency resolution are used together to determine the basic half-bandwidth of the candidate harmonic frequency band, as shown in Equation (6). Subsequently, the branch network outputs an adaptive adjustment amount based on the spectral features inside and outside the candidate harmonic frequency band, which is used to continuously adjust the frequency band center, frequency band width and suppression intensity. Its normalized frequency offset construction method is shown in Equation (7). Then, a basic soft mask is generated using a continuous transition window to avoid boundary abrupt changes, as shown in Equation (8), and the soft mask is mapped to a continuous suppression weight for single-order harmonics, as shown in Equation (9). The continuous suppression weights for multi-order harmonics are combined and a lower limit constraint is added to form a complete continuous suppression weight map, as shown in Equation (10). To reduce the thinning of low-frequency speech, a minimum retention constraint can be introduced in the low-frequency speech region and fused with the continuous weight map, as shown in Equation (11). During the training phase, the change amplitude of the weights in the frequency direction can be constrained to limit local abrupt changes, and its constraint term construction is shown in Equation (12). The final continuous suppression weight map is used to generate the first enhancement spectrum and is used as the output of the harmonic suppression branch in step S2 to enter the subsequent fusion.

[0061] Furthermore, the minimum preservation constraint for the low-frequency speech region is based on the speech fundamental frequency neighborhood. The default search range for speech fundamental frequency estimation is 60–300, with an allowable range of 50–400. The estimation data used is the amplitude spectrum of the same time frame and is frame-aligned with the complex time spectrum. The minimum preservation weight for the speech protection region is set to 0.3 by default as an implementation parameter, with an allowable range of 0.1–0.6. The frequency extension parameter of the speech protection window is set to 50 by default as an implementation parameter, with an allowable range of 20–150, in Hz. The variation amplitude constraint of the frequency domain suppression weight is calculated during the training phase based on the differential amplitude of adjacent frequency indices in the frequency direction. The weight of the differential penalty term is set to 0.1 by default as an implementation parameter, with an allowable range of 0.01–1.0, so that the weight map has a continuous concave shape in the harmonic band and remains gradually changing in the low-frequency speech region.

[0062] Specifically, in the above implementation method, when generating continuous suppression weights around the power frequency and its harmonics, the formation of candidate harmonic frequency bands is related to frequency resolution, spectral peak broadening, and local noise floor, allowing the frequency band range to change with the spectral shape variations of different time frames. The soft suppression shape uses a continuous transition window, maintaining smooth changes at the frequency band boundaries to avoid abrupt distortion caused by binary masking. The neural network outputs suppression strength, bandwidth adjustment, and center offset within the same branch, making the weight map adaptive to harmonic drift, leakage broadening, and different noise floor levels. The weights of multiple harmonics are fused in a superposition manner to form a continuous weight map, which can generate continuous dips at multiple harmonics while maintaining energy continuity in non-harmonic regions. After introducing minimum retention and gradual variation constraints to low-frequency speech components, harmonic suppression will not form excessively deep dips in the speech fundamental frequency neighborhood, thereby reducing phenomena such as thinning of speech and loss of low frequencies after enhancement, allowing the output to maintain the fullness and naturalness of speech sound while suppressing steady-state harmonic noise.

[0063] In this embodiment, the transient indication map is a continuous value map with the same resolution as the complex time spectrum, which represents the confidence level of each time-frequency unit as transient impulse noise, and the transient region is obtained by thresholding; Thresholding employs either a fixed threshold or a quantile threshold rule; transient regions are subject to a minimum duration frame constraint on the time axis, and isolated regions shorter than the minimum duration frame are removed from the transient regions; transient regions are subject to a minimum bandwidth constraint on the frequency axis, and isolated regions narrower than the minimum bandwidth are removed from the transient regions; transient regions are subject to a connected component merging rule on the time-frequency plane, and adjacent connected components that meet the preset interval conditions are merged into the same transient region.

[0064] In this embodiment, the fixed threshold for thresholding is set to 0.6 by default, with an allowable range of 0.3 to 0.8; the quantile threshold rule is set to 90% by default, with an allowable range of 85% to 98%, and the quantile is obtained by statistically analyzing the transient indicator map values ​​of the same time frame; the shortest duration frame constraint is set to 2 frames by default, with an allowable range of 1 to 5 frames, and the duration frame count is the number of frames in which the transient region mask has a consecutive value of 1 on the time axis; the minimum bandwidth constraint is set to 3 frequency indices by default, with an allowable range of 1 to 8 frequency indices; the time interval of the preset interval condition is set to 1 frame by default, with an allowable range of 0 to 3 frames, and the frequency interval is set to 2 frequency indices by default, with an allowable range of 0 to 6 frequency indices. Adjacent connected components that satisfy both the above time interval and frequency interval conditions are merged into the same transient region.

[0065] In this embodiment, the gated convolutional inpainting network uses the neighborhood time-frequency features of the transient region and its corresponding region mask as conditional inputs, and modulates the contribution of neighborhood information to the inpainting result through a gating mechanism, and outputs a replacement spectrum for filling the transient region. The neighborhood time-frequency features include complex time-frequency features corresponding to the set of adjacent time-frequency units of the transient region on the time and frequency axes; the region mask has the same resolution as the complex time-frequency spectrum and indicates the effective context position and the position to be repaired; the gated convolutional inpainting network outputs content features and gated features simultaneously in each convolutional layer, the gated features are passed through a gated activation function to obtain gated weights, and the content features are modulated point by point; the gated convolutional inpainting network performs an effective region update rule on the region mask during multi-layer propagation, and the updated effective region only contains the repaired position and the original context position inferred from the context; the replacement spectrum performs a boundary smoothing rule at the boundary of the transient region, and the boundary smoothing rule continuously transitions and fuses the repaired result and the unrepaired result in the time direction and the frequency direction, respectively.

[0066] The gated convolutional inpainting network generates candidate inpainting features through main branch convolution and generates a threshold through gated branches in each network layer. The threshold is then used to modulate the candidate inpainting features point by point to control the contribution of neighborhood information to the inpainting result. The transient region mask is updated between network layers according to the threshold, so that the region to be inpainted gradually shrinks as the number of layers increases, thereby suppressing the spread of the inpainting effect to non-transient regions. In the output stage, a continuous transition mask is constructed based on the transient region mask, and the replacement spectrum and the original temporal spectrum are continuously fused together to reduce abrupt changes in the inpainting boundary and mitigate audible trailing. Optionally, the gated activation function uses the Sigmoid gate function to ensure that the gate value is between 0 and 1. The closer the gate value is to 0, the stronger the scaling of the candidate repair features at that position; the closer the gate value is to 1, the more candidate repair features are preserved at that position. The normalization stabilization term, as an implementation parameter, is set to 0.000001 by default, with an allowed range of 0.0000001 to 0.00001. The threshold value, as an implementation parameter, is set to 0.7 by default, with an allowed range of 0.4 to 0.9, and is used to determine the transient region mask during inter-layer updates. Whether to retain the markers to be repaired; the temporal and frequency dimensions of the convolutional kernel are set to 3×3 by default as implementation parameters, with an allowed range of 3×3 to 9×9; the number of network layers is set to 8 by default as implementation parameters, with an allowed range of 4 to 16; the smoothing weight coefficients of the continuous transition mask are non-negative and sum to 1 after normalization, and the default setting is equal weight, but adjustments are allowed with a larger weight at the center; the temporal and frequency dimensions of the smoothing neighborhood are set to 3×3 by default as implementation parameters, with an allowed range of 3×3 to 7×7.

[0067] In one implementation, the gating mechanism in the gated convolutional inpainting network process regulates the contribution of neighborhood information through a continuous mapping: the main branch generates candidate inpainting features, the gated branch generates a threshold, and the threshold is used to scale the candidate inpainting features point by point. Simultaneously, a transient region mask is used as a condition to limit the neighborhood sampling range and suppress the diffusion of inpainting effects into non-transient regions. Specifically: When the gated convolutional repair network is at the Layer-to-time-frequency location When aggregating neighborhood context, transient region masks can be used. As a conditional term for neighborhood sampling, the neighborhood participating in the convolution only comes from non-transient regions, thus obtaining the main branch convolution output: (13)

[0068] In equation (13), Indicates the first Layer in time frame Frequency Index The main branch convolution output at the location; Indicates the network layer index; Indicates the time frame index; Indicates frequency index; Indicates the relative displacement index of the convolution kernel in the time and frequency directions; Indicates the first The set of indices for each convolutional kernel; Indicates the first Layer in position The transient region mask at a given location, with a value of 0 or 1. A value of 1 indicates that the location belongs to a transient region. Indicates the first Layer main branch convolution kernel in displacement The convolution coefficients at the specified points; Indicates the first Layer in position Time-frequency characteristics at the location; This represents the stable term for normalization, used to avoid the denominator being zero.

[0069] The gated branch convolution outputs are computed in parallel at the same layer, and then passed through successive gating functions to generate a gate value, which characterizes the effective contribution ratio of neighborhood information at that position. (14)

[0070] In equation (14), Indicates the first Layer in time frame Frequency Index Gated branch convolution output at the point; Indicates the first Layer-gated branched convolution kernels in displacement The convolution coefficients at the specified points; the meanings of the remaining parameters are consistent with the aforementioned formula.

[0071]

[0072] In equation (15), Indicates the first Layer in position The threshold value at the specified location has a range of values. ; This represents the Sigmoid gate function, used to... Mapped to continuous gate values; The meaning of is defined above.

[0073] By scaling the main branch output point by point based on the gate value, the repair features of this layer can be obtained:

[0074] In equation (16), Indicates the first Layer in position Output features at the location; Represents a nonlinear activation function; This represents point-by-point multiplication; and The meaning of is explained in the preceding definition. Threshold When the threshold is close to 0, the contribution of neighborhood information to the repair features is reduced; when the threshold is close to 1, the contribution of neighborhood information to the repair features is amplified. This forms a continuous control relationship that regulates the contribution of neighborhood information to the repair results through a gating mechanism.

[0075] To support repair only in transient regions and avoid tail propagation, the transient region mask can be updated inter-layer by a threshold, causing the region to be repaired to shrink as the layer progresses, thus confining the repair impact within the transient region:

[0076] In equation (17), Indicates the first Layer in position Updated transient region mask; This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true and 0 when it is false. This represents the threshold value, used to determine whether further repair is still needed at this location; and The meaning is explained in the preceding definition. This update method causes the positions where the departmental values ​​gradually increase within the transient region to be masked out in advance, and subsequent layers no longer apply continuous repairs to them, thereby reducing the risk of trailing diffusion in the time direction.

[0077] To ensure smooth boundary transitions and prevent audible trailing after the disappearance of the spectrum, a continuous transition mask can be introduced at the fusion point between the output replacement spectrum and the original spectrum. This allows the transient region boundary to switch from a binary state to a continuous state:

[0078] In equation (18), Indicates the location The continuous transition mask at the point has a value range of . ; Represents the set of neighborhood indices used for boundary smoothing; Represents neighborhood displacement The smoothing weighting coefficients at the point are taken as non-negative values ​​and used to form pairs. Local weighted average; Indicates the initial transient region mask at position The value at; This indicates that the smaller value is selected during the operation.

[0079] Using a continuous transition mask to fuse the replacement spectrum with the original spectrum yields a repaired output with continuous boundaries:

[0080] In equation (19), The second enhancement spectrum representing the output of the transient repair branch is located at... Complex values ​​at; Indicates the position of the spectrum when the complex number contains noise. Complex values ​​at; This indicates the replacement spectrum output by the gated convolutional repair network at position. Complex values ​​at; The meaning is explained in the preceding definition. This fusion method ensures a continuous transition of the transient region boundary in both time and frequency directions, reducing abrupt changes at the boundary and thus lowering the probability of an audible trailing effect after the pulse disappears.

[0081] Here, the neighborhood time-frequency features and region mask corresponding to the transient region are first used as input conditions, so that the repair only extracts reliable context from the non-transient region. In each network layer, the main branch generates candidate repair features based on the neighborhood information after mask filtering. The construction method of the main branch output is shown in Equation (13). The gated branch generates gated features under the same neighborhood conditions. The output of the gated branch is shown in Equation (14), and continuous gate values ​​are obtained through the gate function. The gate value generation is shown in Equation (15). Then, the candidate repair features of the main branch are scaled point by point using the gate value to obtain the repair feature output of the layer. The combination method is shown in Equation (16), thereby controlling the contribution of neighborhood information to the repair result at the time-frequency unit level. In order to avoid the repair effect from spreading to the non-transient region, the transient region mask can be updated between layers according to the gate value, so that the repair range gradually shrinks as the number of layers advances. The mask update rule is shown in Equation (17). In the output stage, to reduce the abrupt boundary changes between the repaired and unrepaired regions and decrease audible trailing, a continuous transition mask can be constructed from the initial transient region mask, as shown in Equation (18). Then, the replacement spectrum and the original complex time spectrum are fused using the continuous transition mask to obtain the second enhanced spectrum, as shown in Equation (19). The second enhanced spectrum is entered into step S3 as the output of the transient repair branch in step S2, and dynamically fused with the result of another branch.

[0082] Specifically, in gated convolutional inpainting, the main branch generates candidate inpainting information, while the gated branch generates continuous thresholds. The thresholds scale the candidate inpainting information point-by-point, allowing the contribution of the neighborhood context to the inpainting result to have an adjustable proportion. With the region mask as a conditional input, neighborhood aggregation is more biased towards reliable information in non-transient regions. Original anomalous components within transient regions are no longer repeatedly referenced, making the inpainting process closer to completion from the boundary inwards. As the mask is updated between layers with the threshold, the area to be inpainted shrinks layer by layer, and the already completed positions are no longer continuously modified, thus weakening the temporal wake diffusion. The output stage employs a continuous transition fusion method, transforming the boundary between the inpainted and uninpainted regions from abrupt to gradual transition, reducing abruptness at the boundary and minimizing audible trailing or edge ringing after impulse noise disappears.

[0083] In this embodiment, the confidence level includes: the harmonic residual measure of the output of the harmonic suppression branch and the transient region identification and determination measure of the output of the transient repair branch; The harmonic residual metric is determined by the ratio of the residual energy of the first enhanced spectrum within the candidate harmonic frequency band to the reference energy outside the candidate harmonic frequency band, and is smoothed on the time axis; the transient region identification metric is determined by the confidence separation of the transient indicator map in the transient and non-transient regions, and the separation is jointly determined by the average confidence within the region, the average confidence outside the region, and the confidence gradient of the region boundary treatment; both the harmonic residual metric and the transient region identification metric are normalized to a unified value range.

[0084] Furthermore, the residual harmonic metric is calculated in each time frame as the ratio of residual energy within the candidate harmonic frequency band to reference energy outside the candidate harmonic frequency band, and then time smoothing is performed. The time smoothing length is set to 5 frames by default, with an allowable range of 1 to 20 frames. The reference energy outside the candidate harmonic frequency band is statistically obtained from the remaining frequency band energy after removing all candidate harmonic frequency bands in the same time frame, and a lower limit is used to truncate the minimum reference energy to avoid ratio divergence. The lower limit is set to 0.1 times the median value of the reference energy by default, with an allowable range of 0.01 to 0.5 times. The boundary handling confidence gradient in the transient region identification and determination metric is obtained by averaging the absolute differences of the transient indicator maps on the boundary neighborhood of the transient region mask. The boundary neighborhood size is set to 2 frequency indices and 1 time frame by default, with an allowable range of 1 to 6 frequency indices and 0 to 3 time frames. The normalization to a unified value range is achieved by linear scaling and truncation of values ​​outside the range to ensure the stability of subsequent fusion weight calculation.

[0085] In this embodiment, the dynamic fusion weight is calculated according to time frame or time frequency unit and obtained by normalization mapping of confidence, so that the branch with higher confidence occupies a higher fusion weight at the corresponding position. The dynamic fusion weights are calculated at the time frame granularity or the time-frequency unit granularity. Within the same time frame or the same time-frequency unit, the fusion weights of the two branches satisfy the preset normalization constraint. When the confidence of both branches is lower than the preset threshold, the fusion weights are switched to the preset conservative weights and the amplitude change rate constraint is applied to the fusion spectrum. When the confidence of any branch is higher than the preset threshold and the confidence of the other branch is lower than the preset threshold, the fusion weights select the high-confidence branch as the main branch and limit the maximum weight of the low-confidence branch.

[0086] Specifically, the normalization mapping uses the summation and normalization of the confidence scores of the two branches to obtain the fusion weight, and adds a normalization stabilization term to the denominator to avoid numerical instability caused by both branches approaching 0 simultaneously. The normalization stabilization term is set to 0.000001 by default as an implementation parameter, with an allowable range of 0.0000001 to 0.00001. The preset conservative weights are set to 0.5 and 0.5 when the confidence scores of the two branches are simultaneously below the threshold, with an allowable range of 0.4 to 0.6 and 0.6 to 0.4. The maximum weight of the low-confidence branch is set to 0.2 by default as an implementation parameter, with an allowable range of 0.05 to 0.35. The amplitude change rate constraint is calculated on the amplitude spectrum of the fusion spectrum according to the change in adjacent time frames, with a default limit of no more than 6dB per frame, an allowable range of 3dB to 12dB, and the amplitude spectrum of the current time frame is limited when the limit is exceeded to maintain auditory continuity.

[0087] Example 2

[0088] Based on Example 1, such as Figure 2 As shown, this application also proposes a deep learning-based audio enhancement system for power scenarios, including: The time-frequency conversion module is used to convert noisy time-domain audio into a complex time-frequency spectrum; A harmonic suppression branch network is used to generate frequency domain suppression weights and output the first enhancement spectrum; A transient repair branch network is used to generate a transient indication map and output a second enhancement spectrum; The fusion module is used to calculate dynamic fusion weights based on confidence levels and output the fusion spectrum; The reconstruction module is used to perform an inverse short-time Fourier transform on the fused spectrum to output enhanced time-domain audio. In this embodiment, the reconstruction module uses an overlapping and adding time-domain synthesis method when performing inverse short-time Fourier transform, and uses the phase information corresponding to the fusion spectrum or the phase information output by the phase recovery network. When the phase information corresponding to the fused spectrum is used, the phase information is determined by the phase of the input complex time spectrum or the phase of the first enhanced spectrum, and phase continuity constraints are applied to the phase change in the transient region; when the phase information output by the phase recovery network is used, the phase recovery network takes the amplitude spectrum of the fused spectrum and its corresponding region mask as input and outputs the phase result with the same resolution as the fused spectrum. The phase result and the amplitude spectrum of the fused spectrum are combined to form a complex spectrum for inverse short-time Fourier transform.

[0089] Furthermore, by default, the phase information corresponding to the fused spectrum is used, and the phase of the input complex time spectrum is used as the reference. Phase continuity constraints are applied to phase abrupt changes in the transient region. The continuity constraint is measured by the absolute value of the phase difference between adjacent time frames at the same frequency index. The default threshold is 1.0 rad, and the allowable range is 0.3 rad to 2.5 rad. When the threshold is exceeded, the phase difference is truncated to reduce phase jumps at the transient repair boundary. When the phase recovery network is enabled, the phase result output by the phase recovery network is combined point by point with the amplitude spectrum of the fused spectrum at the same time frame and the same frequency index. The combined complex spectrum is subjected to amplitude upper limit pruning before entering the inverse short-time Fourier transform. The pruning upper limit is taken as the 95th percentile value of the amplitude spectrum by default, and the allowable range is 90% to 99% to suppress the harsh sound caused by abnormal amplification of individual time-frequency units.

[0090] In this embodiment, the harmonic suppression branch network, transient repair branch network, and fusion module are deployed on the same computing device or edge computing device to meet the real-time or near-real-time processing requirements of remote conferencing. Real-time or near-real-time processing includes performing forward inference on a continuous audio stream in frame order, with the processing delay for each frame during inference being less than the corresponding frame shift duration; the system uses the network weights that have been trained and frozen during the inference phase; the system uses the threshold rules, normalization rules, connected component merging rules, and boundary smoothing rules determined during the training phase during the inference phase.

[0091] In this embodiment, the continuous audio stream is input in a sliding frame manner during the inference phase. The input buffer length is set to 2 frames by default as an implementation parameter, with an allowed range of 1 to 5 frames, which is used to support the acquisition of the neighborhood context of the transient repair branch. When a single frame is missing or the frame number is not continuous, the complex time spectrum of the missing frame is maintained by the complex time spectrum of the previous frame and marked as missing to participate in subsequent processing. The marking is only used to set the confidence of the frame to 0 in the confidence calculation to trigger conservative weights. When the continuous missing values ​​exceed the buffer length, the output enhanced temporal audio is maintained by the enhancement result of the most recent frame until the input is restored. After restoration, forward inference continues according to the frame order and the timestamp is kept consistent with the frame shift.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0093] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

Claims

1. A deep learning-based audio enhancement method for power scenarios, characterized in that, include: Step S1: Segment the noisy time-domain audio into frames and perform a short-time Fourier transform to obtain the complex time spectrum; Step S2: Input the complex time-spectrum parallel harmonic suppression branch network and the transient repair branch network; The harmonic suppression branch network generates a continuous frequency domain suppression weight map based on the preset or estimated power grid frequency and its harmonic frequencies, and weights the complex time spectrum to suppress steady-state harmonic noise to obtain the first enhancement spectrum. The transient repair branch network outputs a transient indicator map, and the time-frequency region marked in the transient indicator map is used by a gated convolutional repair network to generate a replacement spectrum based on the neighborhood context to obtain a second enhanced spectrum; Step S3: Calculate the dynamic fusion weights based on the confidence levels output by the two branches, and perform weighted fusion of the first enhancement spectrum and the second enhancement spectrum to obtain the fusion spectrum; perform inverse short-time Fourier transform on the fusion spectrum to obtain the enhanced time-domain audio.

2. The deep learning-based audio enhancement method for power scenarios according to claim 1, characterized in that, The estimation of the power grid frequency includes: performing peak search or autocorrelation analysis on the amplitude spectrum of the complex time spectrum within a preset power frequency candidate range to obtain power frequency candidate values, and selecting the one with the largest energy as the power grid frequency.

3. The deep learning-based audio enhancement method for power scenarios according to claim 1, characterized in that, The generation of the continuous frequency domain suppression weights includes: determining candidate harmonic frequency bands with the power grid frequency and its harmonic frequencies as the center, and using a neural network to adaptively determine the frequency band boundary and suppression intensity based on the spectral characteristics inside and outside the candidate harmonic frequency bands, and outputting a frequency domain suppression weight map with continuous values. During the generation of the continuous frequency domain suppression weight map, when constructing candidate harmonic frequency bands around the power frequency and its multi-order harmonic centers, the bandwidth of the candidate harmonic frequency bands is adaptively determined based on the broadening characteristics of the harmonic spectral peaks, the local noise floor, and the frequency resolution. Within the candidate harmonic frequency bands, a continuously transitioning soft suppression window is used to form a basic soft mask, and the neural network outputs adaptive adjustment amounts for adjusting the suppression intensity, bandwidth, and center position based on the spectral characteristics inside and outside the candidate harmonic frequency bands, resulting in a frequency domain suppression weight map that continuously changes at the frequency band boundaries. At the same time, a minimum retention constraint is set for the low-frequency band where the speech fundamental frequency is located, and constraints are applied to the change amplitude of the frequency domain suppression weights.

4. The deep learning-based audio enhancement method for power scenarios according to claim 1, characterized in that, The transient indication map is a continuous value map with the same resolution as the complex time spectrum, which represents the confidence level of each time-frequency unit as transient impulse noise, and the transient region is obtained by thresholding.

5. The deep learning-based audio enhancement method for power scenarios according to claim 1, characterized in that, The gated convolutional inpainting network takes the neighborhood time-frequency features of the transient region and its corresponding region mask as conditional inputs, and modulates the contribution of neighborhood information to the inpainting result through a gating mechanism, and outputs a replacement spectrum to fill the transient region. The gated convolutional inpainting network generates candidate inpainting features through main branch convolution and generates a threshold through gated branches in each network layer. The threshold is used to modulate the candidate inpainting features point by point to control the contribution of neighborhood information to the inpainting result. The transient region mask is updated between network layers according to the threshold, so that the region to be repaired gradually shrinks as the number of layers increases. In the output stage, a continuous transition mask is constructed based on the transient region mask, and the replacement spectrum and the original time spectrum are continuously fused accordingly.

6. The deep learning-based audio enhancement method for power scenarios according to claim 1, characterized in that, The confidence level includes: a measure of harmonic residual output from the harmonic suppression branch and a measure of transient region identification from the transient repair branch.

7. The deep learning-based audio enhancement method for power scenarios according to claim 6, characterized in that, The dynamic fusion weights are calculated based on time frames or time-frequency units and obtained by normalization mapping of the confidence scores, so that branches with higher confidence scores have higher fusion weights at the corresponding positions.

8. A deep learning-based audio enhancement system for power scenarios, based on the deep learning-based audio enhancement method for power scenarios according to any one of claims 1 to 7, characterized in that, include: The time-frequency conversion module is used to convert noisy time-domain audio into a complex time-frequency spectrum; Harmonic suppression branch network is used to generate frequency domain suppression weights and output the first enhancement spectrum; A transient repair branch network is used to generate a transient indication map and output a second enhancement spectrum; The fusion module is used to calculate dynamic fusion weights based on confidence levels and output the fusion spectrum; The reconstruction module is used to perform an inverse short-time Fourier transform on the fused spectrum to output enhanced time-domain audio.

9. The deep learning-based audio enhancement system for power scenarios according to claim 8, characterized in that, The reconstruction module employs an overlapping and adding time-domain synthesis method when performing inverse short-time Fourier transform, and uses the phase information corresponding to the fused spectrum or the phase information output by the phase recovery network.

10. The deep learning-based audio enhancement system for power scenarios according to claim 8, characterized in that, The harmonic suppression branch network, transient repair branch network, and fusion module are deployed on the same computing device or edge computing device.