A speech enhancement method, device and medium based on energy spectrum depth modulation

By preprocessing single-channel speech waveforms and training a frequency-time joint enhancement network, a consistent enhanced spectrum sequence is generated, which solves the problem of insufficient recognition of sensitive frequency bands and harmonic peak representation in existing speech enhancement methods, and realizes the stability and continuity of speech enhancement results.

CN122369484APending Publication Date: 2026-07-10BOYIN HEARING EQUIP (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BOYIN HEARING EQUIP (SUZHOU) CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing speech enhancement methods are insufficient in jointly representing sensitive frequency bands, harmonic peaks, and distortion paths. They are difficult to simultaneously consider the preservation strength and enhancement priority of speech recognition-related information, and the enhancement results have limited constraints in terms of cross-frame continuity, harmonic consistency, and residual convergence.

Method used

By preprocessing the single-channel speech waveform, a complex short-time spectral sequence is generated, and a predicted noise spectral sequence is constructed. Sensitivity confidence maps, harmonic anchoring maps, and distortion split maps are extracted, and a continuous enhancement conditional tensor is encoded. A frequency-time joint enhancement network is constructed and trained. Recursion is performed in conjunction with local refinement branches to output a complex enhancement mask sequence and a depth filter coefficient sequence. Complex weighting and multi-tap complex filtering are performed to generate a consistent enhancement spectral sequence. Finally, waveform reconstruction and steady-state constraints are performed.

Benefits of technology

This approach achieves joint representation of speech recognition sensitive information and harmonic fidelity, improving the discernibility and continuity of the enhancement results and ensuring their stability and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369484A_ABST
    Figure CN122369484A_ABST
Patent Text Reader

Abstract

The application discloses a speech enhancement method and device based on energy spectrum depth modulation and a medium, relates to the technical field of speech signal enhancement, and comprises the following steps: constructing a frequency-time joint enhancement network and training the same, inputting a continuous enhancement condition tensor into the trained frequency-time joint enhancement network, combining a local refinement branch to perform recursion, and outputting a complex enhancement mask sequence and a depth filter coefficient sequence; based on the depth filter coefficient sequence and the complex enhancement mask sequence, performing complex weighting and multi-tap complex filtering on a complex short-time frequency spectrum sequence to generate a candidate enhanced frequency spectrum sequence, and through buffer locking and residual suppression, generating a consistent enhanced frequency spectrum sequence; and performing waveform reconstruction and steady-state constraint on the consistent enhanced frequency spectrum sequence to generate a steady-state enhanced speech sequence. The application achieves the effect of giving consideration to the enhancement continuity and local detail fidelity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal enhancement technology, and in particular to a speech enhancement method, device and medium based on energy spectrum depth modulation. Background Technology

[0002] In the fields of speech signal processing and speech recognition, speech enhancement methods based on time-frequency analysis typically begin by framing, transforming, and representing the spectrum of a single-channel speech waveform. This is followed by noise estimation, time-frequency weighting, masking enhancement, or filtering to achieve noise suppression and speech component recovery. These methods focus on complex spectrum modeling, harmonic structure preservation, temporal continuity constraints, and waveform reconstruction, and are applicable to scenarios such as call noise reduction, terminal audio pickup optimization, and voice interaction front-end processing.

[0003] As the types of noise, speech transition patterns, and harmonic distributions in application scenarios continue to increase, conventional methods still need further improvement in two aspects: First, they are insufficient in jointly representing sensitive frequency bands, harmonic peaks, and distortion paths, making it difficult to simultaneously consider the preservation strength and enhancement priority of speech recognition-related information; second, they have limited constraints on the enhancement results in terms of cross-frame continuity, harmonic consistency, and residual convergence, and there is still room for improvement in the stability of the output speech. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a speech enhancement method based on energy spectrum depth modulation to solve the problems of insufficient directional representation of sensitive information and insufficient consistency constraints of enhancement results.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a speech enhancement method based on energy spectrum depth modulation, comprising: preprocessing a single-channel speech waveform to generate a complex short-time spectral sequence and constructing a predicted noise spectral sequence; based on the predicted noise spectral sequence, extracting a sensitivity confidence map, a harmonic anchoring map, and a distortion splitting map from the complex short-time spectral sequence and encoding them to generate a continuous enhancement conditional tensor; constructing and training a frequency-time joint enhancement network, inputting the continuous enhancement conditional tensor into the trained frequency-time joint enhancement network, and performing recursion by combining local refinement branches to output a complex enhancement mask sequence and a depth filter coefficient sequence; based on the depth filter coefficient sequence and the complex enhancement mask sequence, performing complex weighting and multi-tap complex filtering on the complex short-time spectral sequence to generate a candidate enhancement spectral sequence, and generating a consistent enhancement spectral sequence through buffer locking and residual suppression; and performing waveform reconstruction and steady-state constraints on the consistent enhancement spectral sequence to generate a steady-state enhanced speech sequence.

[0007] As a preferred embodiment of the speech enhancement method based on energy spectrum depth modulation described in this invention, the steps of preprocessing the single-channel speech waveform to generate a complex short-time spectrum sequence and constructing a predicted noise spectrum sequence are as follows: Acquire single-channel speech waveforms and perform pre-emphasis, framing, windowing, and short-time Fourier transform to generate complex short-time spectrum sequences; Based on the complex short-time spectral sequence, a noise candidate map is constructed, and the noise floor value of each frequency band is updated according to the noise candidate map. At the same time, interpolation extrapolation is performed on the time frequency point where speech is dominant to generate a predicted noise spectral sequence.

[0008] As a preferred embodiment of the speech enhancement method based on energy spectrum depth modulation described in this invention, the steps of extracting a sensitivity confidence map, a harmonic anchoring map, and a distortion splitting map from a complex short-time spectrum sequence based on a predicted noise spectrum sequence, and encoding them to generate a continuous enhancement conditional tensor, are as follows: Based on the inter-frame energy changes, spectral stability, and transition segment response differences of each frequency band in the complex short-time spectral sequence, the consonant transition frequency band and the steady-state elementary frequency band are marked to generate a sensitivity confidence map; High-confidence time-frequency points are selected from the sensitivity confidence map, and noise pseudo-peaks are removed by combining the predicted noise spectrum sequence to obtain the harmonic anchoring map. At the same time, a distortion shunt map is generated based on the difference in the amount of disturbance decrease at each time-frequency point. Causal convolutional recursive encoding is performed on complex short-time spectral sequences, sensitivity confidence maps, harmonic anchoring maps, and distortion shunt maps to generate continuous enhancement conditional tensors.

[0009] As a preferred embodiment of the speech enhancement method based on energy spectrum deep modulation described in this invention, the steps of constructing and training the frequency-time joint enhancement network are as follows: A mapping layer is constructed using two-dimensional point convolution and activation functions, and a frequency-oriented state layer and a temporal state layer are built based on the Mamba architecture. At the same time, a local refinement branch is built through local convolution and recursive modeling. A bridge-constrained fusion layer is built based on gated fusion and conditional modulation, and a complex enhancement mask head and a depth filter head are built based on complex dual-channel projection branch and multi-tap coefficient projection branch, respectively. The mapping layer, frequency-oriented state layer and time-oriented state layer are connected sequentially according to the residuals, and the local refinement branches are connected in parallel to the bridge constraint fusion layer. At the same time, the complex enhancement mask head and the depth filter head are connected in parallel to construct the frequency-time joint enhancement network. Based on noisy single-channel speech samples and clean speech samples, the frequency-time joint augmentation network is trained under supervision to obtain the trained frequency-time joint augmentation network.

[0010] As a preferred embodiment of the speech enhancement method based on energy spectrum depth modulation described in this invention, the steps of inputting the continuous enhancement conditional tensor into the trained frequency-time joint enhancement network, and recursively combining it with local refinement branches to output a complex enhancement mask sequence and a depth filter coefficient sequence are as follows: Based on the trained frequency-time joint augmentation network, the complex short-time spectrum sequence, the continuous augmentation conditional tensor, and the predicted noise spectrum sequence are dimensionally aligned to form a pre-cleaned complex spectrum sequence. The complex short-time spectral sequence, the pre-cleaned complex spectral sequence, and the continuous enhancement conditional tensor are input into the mapping layer, and the full-frequency context sequence is output through frequency-directed state layer and time-directed state layer recursion. The complex short-time spectrum sequence, the continuous enhancement condition tensor, and the harmonic anchoring map are input into the local refinement branch. The harmonic anchor point is used as the center to recursively deduce the neighboring frequency points and neighboring frames, and the sub-band context sequence is output. The full-frequency context sequence, sub-band context sequence, continuous enhancement conditional tensor, predicted noise spectrum sequence, and sensitivity confidence map are input to the bridge constraint fusion layer and fused to output the path control fusion feature sequence. The path control fusion feature sequence is input into the complex enhancement mask head and the depth filter head respectively to generate the complex enhancement mask sequence and the depth filter coefficient sequence.

[0011] As a preferred embodiment of the speech enhancement method based on energy spectrum depth modulation described in this invention, the steps of performing complex weighting and multi-tap complex filtering on the complex short-time spectrum sequence based on the depth filter coefficient sequence and the complex enhancement mask sequence to generate a candidate enhanced spectrum sequence are as follows: Perform banding noise cancellation on the complex short-time spectral sequence to generate a pre-cancelled spectral sequence; Based on the complex enhancement mask sequence, the pre-cancellation spectrum sequence is complex-weighted, and multi-tap complex filtering is performed through the depth filter coefficient sequence to generate candidate enhancement spectrum sequences.

[0012] As a preferred embodiment of the speech enhancement method based on energy spectrum depth modulation described in this invention, the step of generating a consistency-enhanced spectral sequence through buffer locking and residual suppression is as follows: Recursive recovery is performed on low-confidence time-frequency points in the time-frequency plane corresponding to the candidate enhanced spectrum sequence, and buffer locking is performed on high-confidence time-frequency points to generate a recursive buffered enhanced spectrum sequence. Based on the harmonic anchoring diagram, harmonic continuous correction and modulation continuous correction are performed on the recursive buffered enhanced spectrum sequence, and residual suppression is performed on the recursive buffered enhanced spectrum sequence based on the predicted noise spectrum sequence to generate a uniform enhanced spectrum sequence.

[0013] As a preferred embodiment of the speech enhancement method based on energy spectrum depth modulation described in this invention, the steps of performing waveform reconstruction and steady-state constraints on the consistency enhancement spectrum sequence to generate a steady-state enhanced speech sequence are as follows: Perform inverse short-time Fourier transform and overlap addition on the consistency-enhanced spectral sequence to generate the initial enhanced speech sequence; Based on the initial enhanced speech sequence, frame-level gain constraints and silent segment noise floor shaping are performed, and a steady-state enhanced speech sequence is generated by maintaining harmonic amplitude and normalizing the overall amplitude.

[0014] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the speech enhancement method based on energy spectrum depth modulation as described in the first aspect of the present invention.

[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the speech enhancement method based on energy spectrum depth modulation as described in the first aspect of the present invention.

[0016] The beneficial effects of this invention are as follows: by extracting and encoding the sensitivity confidence map, harmonic anchoring map and distortion splitting map, the joint representation of sensitive information and harmonic fidelity constraints in speech recognition is achieved, thereby improving the discriminability; by constructing a frequency-time joint enhancement network, the collaborative recursion of the full-frequency context sequence and the sub-band context sequence is achieved, thereby achieving the effect of balancing enhanced continuity and local detail fidelity. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a speech enhancement method based on energy spectrum depth modulation.

[0019] Figure 2 This is a flowchart for preprocessing single-channel speech waveforms.

[0020] Figure 3 This is a flowchart of the frequency-time joint enhancement network generation process.

[0021] Figure 4 The flowchart for generating the bridge constraint fusion layer. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides a speech enhancement method based on energy spectrum depth modulation, comprising the following steps: S1. Preprocess the single-channel speech waveform to generate a complex short-time spectrum sequence and construct a predicted noise spectrum sequence.

[0026] The system acquires single-channel speech waveforms and performs pre-emphasis, framing, windowing, and short-time Fourier transform to generate complex short-time spectrum sequences.

[0027] Furthermore, after acquiring the single-channel speech waveform, the pre-emphasized speech waveform is obtained by calculating point by point according to the first-order difference coefficient (e.g., 0.97). The pre-emphasized speech waveform is continuously truncated (e.g., 20ms frame length, 10ms frame shift) to generate an overlapping frame sequence. The square root Hanning window is applied to each frame in the overlapping frame sequence to obtain a windowed frame sequence. The windowed frame sequence is then subjected to a short-time Fourier transform (e.g., 512 points) frame by frame, retaining the real and imaginary parts of each frame at each frequency point, and arranged in the order of time index and frequency index to generate a complex short-time spectrum sequence.

[0028] Based on the complex short-time spectral sequence, a noise candidate map is constructed, and the noise floor value of each frequency band is updated according to the noise candidate map. At the same time, interpolation extrapolation is performed on the time frequency point where speech is dominant to generate a predicted noise spectral sequence.

[0029] Furthermore, the amplitude values ​​at each time-frequency point are extracted from the complex short-time spectral sequence, and continuous (e.g., 15 frames) local minimum logarithmic amplitude values ​​are collected in the time direction to form a reference frequency band noise floor. The amplitude value at each time-frequency point is compared point by point with the reference frequency band noise floor of its respective frequency band. When the difference between the amplitude value at each time-frequency point and the reference frequency band noise floor is less than or equal to the amplitude approach threshold, and the inter-frame amplitude difference (the amplitude difference of the same frequency point between consecutive time frames) and the inter-frequency amplitude difference (the amplitude difference between adjacent frequency points in the same time frame) are both considered to be within the threshold, the following criteria are met: When all normalized values ​​are less than the stability threshold, the corresponding time-frequency points are marked as noise candidate points, generating a noise candidate map to locate suspected noise time-frequency points and constrain the update of noise floor values. The amplitude of noise candidate points in each frequency band (e.g., dividing the frequency band into four consecutive frequency points in each frame) is recursively smoothed according to the noise candidate map to update the noise floor values ​​of each frequency band. Using the noise floor values ​​of each frequency band as boundaries, bidirectional interpolation extrapolation is performed on the speech-dominant time-frequency points along the time and frequency directions to generate a predicted noise spectrum sequence.

[0030] It should be noted that the amplitude approach threshold is set based on the difference distribution between the amplitude at each time frequency point and the background noise of the reference frequency band in a pure noisy speech segment (a silent segment extracted from a noisy single-channel speech sample that contains only noise components and no speech components). The high quantile statistical value (e.g., 95%) of the difference distribution is taken as the fixed threshold, with an example value range of 0.5dB to 2dB. The noisy single-channel speech sample is a speech sample obtained from a single acquisition channel that contains both speech components and background noise components. The stability threshold is set based on the statistical distribution of the inter-frame amplitude difference and inter-frequency amplitude difference in the pure noisy speech segment. The two types of amplitude differences are normalized statistically and the quantile value (e.g., 95%) is taken as the fixed threshold, with an example value range of 0.08 to 0.15.

[0031] S2. Based on the predicted noise spectrum sequence, extract the sensitivity confidence map, harmonic anchoring map and distortion splitting map from the complex short-time spectrum sequence, and encode them to generate a continuous enhancement conditional tensor.

[0032] Based on the inter-frame energy changes, spectral stability, and transition segment response differences of each frequency band in the complex short-time spectral sequence, the consonant transition frequency band and the steady-state elementary frequency band are marked, and a sensitivity confidence map is generated.

[0033] Furthermore, the amplitude squares of each frequency point within each frequency band in the complex short-time spectral sequence are summed to obtain the frequency band energy sequence; based on the frequency band energy sequences of adjacent frames (e.g., 3 frames) within the same frequency band, the inter-frame energy change is obtained; the spectral stability is obtained based on the normalized amplitude distribution of the current frame and the frames before and after; using the current frame frequency band energy in the frequency band energy sequence as a benchmark, the mean energy of the front window frequency band and the mean energy of the back window frequency band are obtained respectively, and the transition segment response difference is obtained based on the difference between the mean energy of the front window frequency band and the mean energy of the back window frequency band; the frequency bands with inter-frame energy change not less than the inter-frame energy change threshold and the transition segment response difference not less than the transition segment response threshold are marked as consonant transition frequency bands, and the frequency bands with spectral stability not less than the spectral stability threshold and the inter-frame energy change less than the inter-frame energy change threshold are marked as steady-state atomic frequency bands, generating a sensitivity confidence map to indicate the enhancement priority and retention strength of sensitive frequency bands for speech recognition.

[0034] It should be noted that the inter-frame energy change threshold is set based on the statistical distribution of the frequency band energy difference in the noisy single-channel speech samples. For example, the 95th percentile is used as a fixed threshold, with an example value range of 0.12 to 0.25. The transition response threshold is set based on the statistical distribution of the difference between the mean energy of the front window frequency band and the mean energy of the back window frequency band in the noisy single-channel speech samples. For example, the 95th percentile is used as a fixed threshold, with an example value range of 0.18 to 0.35. The spectral stability threshold is set based on the statistical distribution of the consistency of the normalized amplitude distribution in the noisy single-channel speech samples. For example, the 90th percentile is used as a fixed threshold, with an example value range of 0.75 to 0.9. The consonant transition frequency band refers to the frequency band with large inter-frame energy changes and significant differences in response between the front and back time windows. It is used to mark the transient frequency bands that are more sensitive to recognition in the speech transition area. The steady-state vowel frequency band refers to the frequency band with stable spectral shape and small inter-frame energy changes. It is used to mark the stable harmonic and formant frequency bands that need to be maintained in the steady-state vowel region.

[0035] High-confidence time-frequency points are selected from the sensitivity confidence map, and noise pseudo-peaks are removed by combining the predicted noise spectrum sequence to obtain the harmonic anchoring map. At the same time, a distortion shunt map is generated based on the difference in the amount of disturbance decrease at each time-frequency point.

[0036] Furthermore, local maxima time-frequency points are extracted from the sensitivity confidence map as candidate peaks, and the candidate peaks are canceled out point by point with the predicted noise spectrum sequence. Joint sorting is performed, and the highest-ranked main peak is retained in the competitive peak group. Harmonic inheritance is performed at the second and third harmonic positions to eliminate noise pseudo-peaks and generate a harmonic anchoring map, which is used to lock the stable harmonic main peak and constrain subsequent fidelity enhancement. Based on the complex short-time spectrum sequence, amplitude-preserving write-back spectrum sequence and noise-suppressing write-back spectrum sequence are constructed. The perturbation reduction of the corresponding time-frequency points of the two write-back spectrum sequences relative to the sensitivity confidence map is obtained, and the path assignment is determined. A distortion split map is generated to distinguish between amplitude-preserving paths and noise-suppressing paths and reduce speech distortion.

[0037] Causal convolutional recursive encoding is performed on complex short-time spectral sequences, sensitivity confidence maps, harmonic anchoring maps, and distortion shunt maps to generate continuous enhancement conditional tensors.

[0038] Furthermore, the real and imaginary parts of the complex short-time spectral sequence (the real part represents the spectral components on the cosine basis, and the imaginary part represents the spectral components on the sine basis) are aligned with the sensitivity confidence map, harmonic anchoring map, and distortion splitting map at time-by-time frequency points, and then concatenated along the feature dimension to form a conditional input sequence. A causal convolution covering only the current frame and historical frames (e.g., two historical frames) is applied to the conditional input sequence frame by frame to obtain a local conditional sequence. The local conditional sequence is synthesized with the recursive state of the previous frame in chronological order, and the harmonic position is enhanced using the harmonic anchoring map. The amplitude preservation and noise suppression weights are adjusted using the distortion splitting map to obtain the conditional state vector of each frame. The conditional state vectors of each frame are stacked according to the time index and frequency index to generate a continuously enhanced conditional tensor.

[0039] S3. Construct and train a frequency-time joint augmentation network. Input the continuous augmentation conditional tensor into the trained frequency-time joint augmentation network, and perform recursion by combining the local refinement branch to output the complex augmentation mask sequence and the depth filter coefficient sequence.

[0040] A mapping layer is constructed using two-dimensional point convolution and activation functions, and a frequency-oriented state layer and a temporal state layer are built based on the Mamba architecture. At the same time, local refinement branches are built through local convolution and recursive modeling.

[0041] Furthermore, two-dimensional point convolution (with a kernel size of, for example, 1×1) and SiLU activation are applied to each time-frequency point of the complex short-time spectral sequence to form a mapping layer; selective state recursion is performed along the frequency order and time order according to the Mamba architecture to form a frequency-directed state layer and a time-directed state layer; local convolution and recursive modeling are performed on the neighboring frequency points and neighboring frames with the harmonic anchor map marker position as the center to build a local refinement branch.

[0042] It should be noted that the Mamba architecture is a linear complexity sequence modeling architecture based on a selective state space model. By recursively scanning the state updates related to the input, it can simultaneously model long-range dependent local dynamics without using attention. The Mamba architecture was chosen in this embodiment because speech enhancement scenarios require both modeling long-term contexts across time frames to suppress persistent noise and preserving local transients and harmonic neighborhood details. Compared to structures that only use local convolution, the Mamba architecture is more suitable for handling long sequence recursions while maintaining linear complexity.

[0043] A bridge-constrained fusion layer is built based on gated fusion and conditional modulation. Complex enhancement mask head and depth filter head are built based on complex dual-channel projection branch and multi-tap coefficient projection branch, respectively.

[0044] Furthermore, using the feature dimensions of the temporal state layer output and the local refinement branch output as a unified input dimension, gated convolution branches, fusion convolution branches, and conditional modulation mapping branches are set up respectively to construct a bridge-constrained fusion layer; using the output dimension of the bridge-constrained fusion layer as the input dimension, a complex dual-channel projection branch is concatenated, namely two layers (convolution kernel size, for example, 1×1) of convolution and one layer of dual-channel linear projection, to construct a complex enhancement mask head; a multi-tap coefficient projection branch is concatenated, namely two layers (convolution kernel size, for example, 1×1) of convolution and one layer (for example, the number of taps is fixed at 5) of linear projection, to output a depth filter coefficient sequence, and to construct a depth filter head.

[0045] It should be noted that gated fusion refers to the process of calculating full-frequency retention coefficients and performing weighted synthesis on the full-frequency context features and sub-band context features at the same time-frequency position, which is used to allocate the proportion of global enhancement and local refinement; conditional modulation refers to mapping the continuous enhancement conditional tensor into scaling modulation terms and offset modulation terms, and then performing scaling and offset processing on the normalized fusion features, which is used to inject sensitivity confidence, harmonic anchoring and distortion splitting constraints into subsequent enhancement paths; complex dual-channel projection branch is a projection branch that outputs real part masks and imaginary part masks at each time-frequency point, which is used to generate complex enhancement mask sequences; multi-tap coefficient projection branch is a projection branch that outputs depth filter coefficient sequences at each time-frequency point, which is used to support subsequent multi-tap complex filtering.

[0046] The mapping layer, frequency-oriented state layer, and temporal state layer are connected sequentially according to the residuals, and the local refinement branches are connected in parallel to the bridge constraint fusion layer. At the same time, the complex enhancement mask head and the depth filter head are connected in parallel to construct a frequency-temporal joint enhancement network.

[0047] It should be noted that the frequency-time joint enhancement network refers to a joint recursive network consisting of a mapping layer, a frequency-oriented state layer, a time-oriented state layer, a local refinement branch, a bridge constraint fusion layer, a complex enhancement mask head, and a depth filter head. It is used to generate the parameters required for enhancement by simultaneously utilizing cross-frequency correlation, cross-frame correlation, and harmonic neighborhood details.

[0048] Based on noisy single-channel speech samples and clean speech samples, the frequency-time joint augmentation network is trained under supervision to obtain the trained frequency-time joint augmentation network.

[0049] Furthermore, the noisy single-channel speech samples and the clean speech samples are synchronously segmented with the same time length to generate noisy complex spectrum sequences and clean complex spectrum sequences, respectively. The noisy complex spectrum sequences, the continuous enhancement conditional tensors generated based on the noisy complex spectrum sequences, and the predicted noise spectrum sequences generated based on the noisy complex spectrum sequences are fed into the frequency-time joint enhancement network in batches (e.g., 16). Through frequency-time recursion, bridge constraint fusion, and joint modulation, a consistent enhancement spectrum sequence is generated. The consistent enhancement spectrum sequence and the clean complex spectrum sequence are compared hourly at each frequency point. The network parameters are updated by backpropagating according to the complex difference, amplitude difference, and phase difference. The network is iterated (e.g., 80 rounds) at a learning rate (e.g., 0.001) until the loss converges, resulting in the trained frequency-time joint enhancement network.

[0050] It should be noted that a clean speech sample is a single-channel speech sample without background noise components, which is usually obtained by collecting single-person speech recordings in a quiet environment or by directly selecting from noise-free speech corpora.

[0051] Based on the trained frequency-time joint augmentation network, the complex short-time spectral sequence, the continuous augmentation conditional tensor, and the predicted noise spectral sequence are aligned in terms of execution dimension to form a pre-cleaned complex spectral sequence.

[0052] Furthermore, based on the input dimension of the mapping layer in the trained frequency-time joint augmentation network, the real and imaginary parts of the complex short-time spectrum sequence, the continuous augmentation conditional tensor, and the predicted noise spectrum sequence are rearranged and aligned according to the same time index, frequency index, and channel dimension; the predicted noise spectrum sequence is encoded as a noise reference input to form a pre-cleaned complex spectrum sequence.

[0053] The complex short-time spectral sequence, the pre-cleaned complex spectral sequence, and the continuously enhanced conditional tensor are input into the mapping layer, and the full-frequency context sequence is output through frequency-directed state layer and time-directed state layer recursion.

[0054] Furthermore, the real and imaginary parts of the complex short-time spectrum sequence and the real and imaginary parts of the pre-cleaned complex spectrum sequence are aligned point-by-point with the continuous enhancement conditional tensor, and then concatenated along the feature dimension before being fed into the mapping layer. The mapping layer applies two-dimensional point convolution (with a kernel size of, for example, 1×1) and SiLU activation to each time-frequency point to obtain the initial mapping features. The initial mapping features of the current frequency point are recursively synthesized with the state of the previous frequency point in frequency order to generate frequency-direction state features. The frequency-direction state features of the current time frame are recursively synthesized with the state of the previous time frame in time order to output the full-frequency context sequence, which is used to collect global enhancement basis across frequency points and across time frames.

[0055] The complex short-time spectrum sequence, the continuous enhancement conditional tensor, and the harmonic anchoring map are input into the local refinement branch. The harmonic anchor point is used as the center to recursively deduce the neighboring frequency points and neighboring frames, and the sub-band context sequence is output.

[0056] Furthermore, the real and imaginary parts of the complex short-time spectral sequence are aligned with the continuous enhancement conditional tensor and fed into the local refinement branch. Harmonic anchor points are located based on the harmonic anchoring map. The corresponding features of the neighboring frequency points (e.g., two before and two after) and the neighboring frames (e.g., one before and one after) are extracted with each harmonic anchor point as the center. Local convolution is applied along the frequency direction and recursively synthesized with the previous neighboring state to obtain the local state of the anchor point. The local state of the anchor point is recursively passed between adjacent frames to output the sub-band context sequence, which is used to enhance the preservation of harmonic neighborhood details.

[0057] The full-frequency context sequence, sub-band context sequence, continuous enhancement conditional tensor, predicted noise spectrum sequence, and sensitivity confidence map are input to the bridge constraint fusion layer and fused to output the path control fusion feature sequence.

[0058] Furthermore, the full-frequency context features, sub-band context features, sensitivity confidence values ​​in the sensitivity confidence map, harmonic anchoring values ​​in the harmonic anchoring map, distortion shunting values ​​in the distortion shunting map, and predicted noise spectrum values ​​in the predicted noise spectrum sequence at the same time-frequency location are aligned and input into the gating mapping weight matrix. A gating mapping bias term is then added to obtain the gating mapping value, which is compressed to between 0 and 1 using the Sigmoid activation function to generate the full-frequency retention coefficient. The sub-band retention coefficient is obtained by subtracting the full-frequency retention coefficient from 1. The full-frequency retention coefficient is applied point by point to the full-frequency context features, and the sub-band retention coefficient is applied point by point to the sub-band context features. The two parts are then weighted and fused and normalized at the execution layer to generate a normalized fused product. Features: The continuous enhancement condition tensor at the same time-frequency position is mapped into scaling modulation terms and offset modulation terms. The scaling modulation term is used to adjust the amplification and compression degree of the normalized fused features, and the offset modulation term is used to adjust the baseline position of the normalized fused features. Scaling processing controlled by the scaling modulation term and offset processing controlled by the offset modulation term are applied to the normalized fused features. The output path controls the fused feature sequence so that the sensitivity confidence value participates in the enhancement allocation of sensitive positions, the harmonic anchoring value participates in the detail preservation of stable harmonic positions, the distortion shunting value participates in the weight adjustment of the amplitude preservation path and the noise suppression path, and the predicted noise spectrum value participates in the residual noise intensity perception, thereby jointly constraining the path allocation of full-frequency enhancement and sub-band refinement.

[0059] It should be noted that the gated mapping weight matrix is ​​a learnable parameter in the bridge-constrained fusion layer. When constructing the bridge-constrained fusion layer, it is initialized according to the splicing dimension of full-frequency context features, sub-band context features, sensitivity confidence value, harmonic anchoring value, distortion splitting value and predicted noise spectrum value, and is obtained through loss backpropagation iterative update during the supervised training process of noisy single-channel speech samples and clean speech samples.

[0060] The expression for calculating the full-frequency retention factor is: ; in, These are the full-frequency retention coefficients, used to characterize the degree to which the current location retains the full-frequency context features; For time location; Frequency position; Use the Sigmoid activation function; This is the gated mapping weight matrix, used to map the concatenated multi-source features into gated discriminants; The full-frequency context features for the current location; The subband context features for the current position; For the first The time frame, the Sensitivity confidence value at each frequency point; For the first The time frame, the Harmonic anchorage values ​​at each frequency point; For the first The time frame, the Distortion shunt values ​​at each frequency point; For the first The time frame, the Predicted noise spectrum values ​​at each frequency point; This is the gating mapping bias term.

[0061] The calculation expression for the path control fusion feature sequence is as follows: ; in, For the first The time frame, the Path control fusion features at each frequency point; This is a layer normalization operation used to stabilize the gated and fused features to a numerical range suitable for subsequent modulation and prediction. This represents element-wise multiplication; Scaling modulation terms generated for continuously enhanced conditional tensors; For the first The time frame, the Continuous enhancement conditional tensor at each frequency point; The offset modulation term generated for the continuously enhanced conditional tensor.

[0062] The path control fusion feature sequence is input into the complex enhancement mask head and the depth filter head respectively to generate the complex enhancement mask sequence and the depth filter coefficient sequence.

[0063] Furthermore, the path control fusion feature sequence is fed into the complex enhancement mask head and the depth filter head respectively; the complex enhancement mask head passes through two layers (convolution kernel size, for example, 1×1) of convolution and one layer of dual-channel linear projection in sequence, and outputs real part mask and imaginary part mask at each time frequency point to form a complex enhancement mask sequence; the depth filter head passes through two layers (convolution kernel size, for example, 1×1) of convolution and one layer (for example, the number of taps is fixed at 5) of linear projection in sequence, and outputs a depth filter coefficient sequence at each time frequency point.

[0064] S4. Based on the depth filter coefficient sequence and the complex enhancement mask sequence, perform complex weighting and multi-tap complex filtering on the complex short-time spectrum sequence to generate candidate enhanced spectrum sequences, and generate a consistent enhanced spectrum sequence through buffer locking and residual suppression.

[0065] Banding noise cancellation is performed on the complex short-time spectral sequence to generate a pre-cancelled spectral sequence.

[0066] Furthermore, the mean noise amplitude and noise floor of each frequency band corresponding to the complex short-time spectrum sequence are extracted from the predicted noise spectrum sequence. The sum of the mean noise amplitude and the noise floor of each frequency band is used as the frequency band noise compensation amount. The noise compensation amount of each frequency band is allocated to each time frequency point in the corresponding frequency band of the complex short-time spectrum sequence. The compensation amount is subtracted point by point from the amplitude of each time frequency point while maintaining the phase relationship between the corresponding real and imaginary parts, thereby generating a pre-cancellation spectrum sequence, which is used to suppress the frequency band noise floor and broadband residual noise in advance.

[0067] Based on the complex enhancement mask sequence, the pre-cancellation spectrum sequence is complex-weighted, and multi-tap complex filtering is performed through the depth filter coefficient sequence to generate candidate enhancement spectrum sequences.

[0068] Furthermore, the real and imaginary masks corresponding to each time-frequency point in the complex enhancement mask sequence are multiplied point by point with the real and imaginary parts of the time-frequency point corresponding to the pre-cancellation spectrum sequence to obtain the mask-weighted spectrum sequence. Then, taking the current time-frequency point as the center, the complex spectrum values ​​of the mask-weighted spectrum sequence in the neighboring time frame and neighboring frequency point are obtained, and the values ​​are accumulated by complex weighting according to each tap coefficient in the depth filter coefficient sequence (for example, the number of taps is fixed at 5) to generate the candidate enhancement spectrum sequence.

[0069] Recursive recovery is performed on the low-confidence time-frequency points in the time-frequency plane corresponding to the candidate enhanced spectrum sequence, and buffer locking is performed on the high-confidence time-frequency points to generate a recursive buffered enhanced spectrum sequence.

[0070] Furthermore, based on the confidence values ​​of each time-frequency point in the sensitivity confidence map, time-frequency points with confidence values ​​not less than the high confidence threshold are marked as high-confidence time-frequency points, and time-frequency points with confidence values ​​not greater than the low confidence threshold are marked as low-confidence time-frequency points. For low-confidence time-frequency points, the real and imaginary parts of the candidate enhanced spectrum sequence of the same frequency point in the previous time frame are used as time references, and the average value of the candidate enhanced spectrum sequences of adjacent frequency points in the current time frame is used as frequency references. The time reference and frequency reference are weighted and synthesized according to the recursive coefficient to obtain the recovery value of the current time-frequency point. For high-confidence time-frequency points, the current candidate enhanced spectrum sequence value is written into a buffer queue (used to directly read the ordered sequence of stable spectrum values ​​during the buffer locking phase) and kept continuous (e.g., for 2 frames) without being updated, generating a recursive buffered enhanced spectrum sequence to suppress instantaneous fluctuations and maintain stable spectral peaks.

[0071] It should be noted that the high confidence threshold and low confidence threshold are set based on the confidence value distribution in the sensitivity confidence map generated from noisy single-channel speech samples. For example, the 90th percentile is taken as the high quantile and the 10th percentile as the low quantile. The example high confidence threshold ranges from 0.8 to 0.9, and the low confidence threshold ranges from 0.1 to 0.2. The recursive coefficient is set based on the statistical distribution of the correlation between the amplitudes of adjacent time frames before and after the same frequency point in the noisy single-channel speech samples. The example value ranges from 0.6 to 0.85.

[0072] Based on the harmonic anchoring diagram, harmonic continuous correction and modulation continuous correction are performed on the recursive buffered enhanced spectrum sequence, and residual suppression is performed on the recursive buffered enhanced spectrum sequence based on the predicted noise spectrum sequence to generate a uniform enhanced spectrum sequence.

[0073] Furthermore, the main harmonic anchor point is located according to the harmonic anchoring diagram, and the harmonic correction position is extended according to the second and third harmonic positions; the real and imaginary parts of the recursive buffer enhancement spectrum sequence corresponding to the main harmonic anchor point and the harmonic correction position are extracted, and the harmonic continuous correction sequence is obtained by correcting the amplitude ratio and phase difference of the same position in adjacent time frames point by point; the real and imaginary parts corresponding to the non-harmonic positions (referring to the time-frequency points other than the main harmonic anchor point and the second and third harmonic positions) are scaled point by point according to the ratio of the average amplitude of adjacent time frames in the same frequency band, and the modulation continuous correction sequence is obtained; based on the noise amplitude of the time-frequency point corresponding to the predicted noise spectrum sequence, the residual noise is subtracted from the amplitude of each time-frequency point of the modulation continuous correction sequence by a proportion (e.g., 20%), and a uniform enhancement spectrum sequence is generated.

[0074] S5. Perform waveform reconstruction and steady-state constraints on the consistency-enhanced spectral sequence to generate a steady-state enhanced speech sequence.

[0075] Perform inverse short-time Fourier transform and overlap addition on the consistency-enhanced spectral sequence to generate the initial enhanced speech sequence.

[0076] Furthermore, the consistency enhancement spectrum sequence is extracted frame by frame according to the time index, and the correspondence between the real and imaginary parts of each frame at each frequency point is maintained. An inverse short-time Fourier transform (e.g., 512 points) is performed to obtain the time-domain enhancement frame sequence. The time-domain enhancement frame sequence is played back sequentially according to the same frame shift (e.g., 10ms) at the time of framing, and the points are added point by point according to the overlapping positions of the window functions corresponding to the windowing stage to generate the initial enhanced speech sequence, which is used to restore the frequency domain enhancement amount to continuous time-domain speech.

[0077] Based on the initial enhanced speech sequence, frame-level gain constraints and silent segment noise floor shaping are performed, and a steady-state enhanced speech sequence is generated by maintaining harmonic amplitude and normalizing the overall amplitude.

[0078] Furthermore, the initial enhanced speech sequence is re-framed with the same frame length and frame shift as in the preprocessing stage. The sum of squared amplitudes of each frame sampling point is obtained to obtain the frame energy of each frame (referring to the frame-level energy value formed by the sum of squared amplitudes of all sampling points within a frame). The frame energy of the current frame is compared with the average frame energy of the adjacent frames to obtain the gain of the current frame, and the gain of the current frame is constrained by the average gain of the adjacent frames. Several frames (the first 5% to 10% of frames) are selected as silent segments according to the frame energy sorting, and the amplitude contour of low-amplitude noise in the silent segments is written back to the silent segment sampling points in chronological order. The amplitude ratio of the main harmonic and the octave position is maintained according to the harmonic anchoring diagram, and normalization is performed according to the peak amplitude of the whole segment to generate a steady-state enhanced speech sequence.

[0079] It should be noted that frame-level gain constraint refers to the process of smoothing the gain of the current frame by using the average gain of the preceding and following frames as the constraint benchmark; low-amplitude noise amplitude profile refers to the change pattern of low-amplitude noise amplitude of silent segment samples in time sequence, that is, the envelope trend formed by the amplitude of each silent segment sample over time; the amplitude ratio of the main harmonic to the harmonic position refers to the relative magnitude relationship between the amplitude corresponding to the main harmonic anchor point and the amplitude corresponding to the second and third harmonic positions, which is used to maintain the consistency of the amplitude hierarchy of the speech harmonic structure; silent segment noise floor shaping refers to the process of writing the low-amplitude noise amplitude profile in the silent segment back to the silent segment samples in time sequence, which is used to avoid the silent segment being excessively compressed and producing an unnatural sense of silence.

[0080] This embodiment also provides a computer device applicable to the speech enhancement method based on energy spectrum depth modulation, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the speech enhancement method based on energy spectrum depth modulation as proposed in the above embodiment.

[0081] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0082] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the speech enhancement method based on energy spectrum depth modulation as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0083] In summary, this invention achieves enhanced recognizability by extracting and encoding the sensitivity confidence map, harmonic anchoring map, and distortion splitting map, thereby jointly representing the sensitive information and harmonic fidelity constraints of speech recognition; and by constructing a frequency-time joint enhancement network, it achieves the collaborative recursion of the full-frequency context sequence and the sub-band context sequence, thus balancing the enhancement of continuity with the preservation of local details.

[0084] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A speech enhancement method based on energy spectrum depth modulation, characterized in that, include: The single-channel speech waveform is preprocessed to generate a complex short-time spectrum sequence, and a predicted noise spectrum sequence is constructed. Based on the predicted noise spectrum sequence, the sensitivity confidence map, harmonic anchoring map and distortion splitting map are extracted from the complex short-time spectrum sequence and encoded to generate a continuous enhancement conditional tensor; Construct and train a frequency-time joint augmentation network. Input the continuous augmentation condition tensor into the trained frequency-time joint augmentation network, and perform recursion by combining the local refinement branch to output the complex augmentation mask sequence and the depth filter coefficient sequence. Based on the depth filter coefficient sequence and the complex enhancement mask sequence, complex weighting and multi-tap complex filtering are performed on the complex short-time spectrum sequence to generate candidate enhanced spectrum sequences. Then, through buffer locking and residual suppression, a consistent enhanced spectrum sequence is generated. Waveform reconstruction and steady-state constraints are performed on the consistency-enhanced spectral sequence to generate a steady-state enhanced speech sequence.

2. The speech enhancement method based on energy spectrum depth modulation as described in claim 1, characterized in that, The steps for preprocessing the single-channel speech waveform to generate a complex short-time spectrum sequence and constructing a predicted noise spectrum sequence are as follows: Acquire single-channel speech waveforms and perform pre-emphasis, framing, windowing, and short-time Fourier transform to generate complex short-time spectrum sequences; Based on the complex short-time spectral sequence, a noise candidate map is constructed, and the noise floor value of each frequency band is updated according to the noise candidate map. At the same time, interpolation extrapolation is performed on the time frequency point where speech is dominant to generate a predicted noise spectral sequence.

3. The speech enhancement method based on energy spectrum depth modulation as described in claim 2, characterized in that, The steps for extracting the sensitivity confidence map, harmonic anchoring map, and distortion splitting map from the complex short-time spectrum sequence based on the predicted noise spectrum sequence, and encoding them to generate a continuous enhancement conditional tensor, are as follows: Based on the inter-frame energy changes, spectral stability, and transition segment response differences of each frequency band in the complex short-time spectral sequence, the consonant transition frequency band and the steady-state elementary frequency band are marked to generate a sensitivity confidence map; High-confidence time-frequency points are selected from the sensitivity confidence map, and noise pseudo-peaks are removed by combining the predicted noise spectrum sequence to obtain the harmonic anchoring map. At the same time, a distortion shunt map is generated based on the difference in the amount of disturbance decrease at each time-frequency point. Causal convolutional recursive encoding is performed on complex short-time spectral sequences, sensitivity confidence maps, harmonic anchoring maps, and distortion shunt maps to generate continuous enhancement conditional tensors.

4. The speech enhancement method based on energy spectrum depth modulation as described in claim 3, characterized in that, The steps for constructing and training the frequency-time joint enhancement network are as follows: A mapping layer is constructed using two-dimensional point convolution and activation functions, and a frequency-oriented state layer and a temporal state layer are built based on the Mamba architecture. At the same time, a local refinement branch is built through local convolution and recursive modeling. A bridge-constrained fusion layer is built based on gated fusion and conditional modulation, and a complex enhancement mask head and a depth filter head are built based on complex dual-channel projection branch and multi-tap coefficient projection branch, respectively. The mapping layer, frequency-oriented state layer and time-oriented state layer are connected sequentially according to the residuals, and the local refinement branches are connected in parallel to the bridge constraint fusion layer. At the same time, the complex enhancement mask head and the depth filter head are connected in parallel to construct the frequency-time joint enhancement network. Based on noisy single-channel speech samples and clean speech samples, the frequency-time joint augmentation network is trained under supervision to obtain the trained frequency-time joint augmentation network.

5. The speech enhancement method based on energy spectrum depth modulation as described in claim 1 or 4, characterized in that, The steps are as follows: Inputting the continuous enhancement conditional tensor into the trained frequency-time joint enhancement network, combining it with local refinement branches for recursion, and outputting a complex enhancement mask sequence and a depth filter coefficient sequence. Based on the trained frequency-time joint augmentation network, the complex short-time spectrum sequence, the continuous augmentation conditional tensor, and the predicted noise spectrum sequence are dimensionally aligned to form a pre-cleaned complex spectrum sequence. The complex short-time spectral sequence, the pre-cleaned complex spectral sequence, and the continuous enhancement conditional tensor are input into the mapping layer, and the full-frequency context sequence is output through frequency-directed state layer and time-directed state layer recursion. The complex short-time spectrum sequence, the continuous enhancement condition tensor, and the harmonic anchoring map are input into the local refinement branch. The harmonic anchor point is used as the center to recursively deduce the neighboring frequency points and neighboring frames, and the sub-band context sequence is output. The full-frequency context sequence, sub-band context sequence, continuous enhancement conditional tensor, predicted noise spectrum sequence, and sensitivity confidence map are input to the bridge constraint fusion layer and fused to output the path control fusion feature sequence. The path control fusion feature sequence is input into the complex enhancement mask head and the depth filter head respectively to generate the complex enhancement mask sequence and the depth filter coefficient sequence.

6. The speech enhancement method based on energy spectrum depth modulation as described in claim 5, characterized in that, The process of performing complex weighting and multi-tap complex filtering on the complex short-time spectrum sequence based on the depth filter coefficient sequence and the complex enhancement mask sequence to generate candidate enhanced spectrum sequences is as follows: Perform banding noise cancellation on the complex short-time spectral sequence to generate a pre-cancelled spectral sequence; Based on the complex enhancement mask sequence, the pre-cancellation spectrum sequence is complex-weighted, and multi-tap complex filtering is performed through the depth filter coefficient sequence to generate candidate enhancement spectrum sequences.

7. The speech enhancement method based on energy spectrum depth modulation as described in claim 1 or 6, characterized in that, The process of generating a consistency-enhanced spectral sequence through buffer locking and residual suppression is as follows: Recursive recovery is performed on low-confidence time-frequency points in the time-frequency plane corresponding to the candidate enhanced spectrum sequence, and buffer locking is performed on high-confidence time-frequency points to generate a recursive buffered enhanced spectrum sequence. Based on the harmonic anchoring diagram, harmonic continuous correction and modulation continuous correction are performed on the recursive buffered enhanced spectrum sequence, and residual suppression is performed on the recursive buffered enhanced spectrum sequence based on the predicted noise spectrum sequence to generate a uniform enhanced spectrum sequence.

8. The speech enhancement method based on energy spectrum depth modulation as described in claim 7, characterized in that, The steps for performing waveform reconstruction and steady-state constraints on the consistency-enhanced spectral sequence to generate a steady-state enhanced speech sequence are as follows: Perform inverse short-time Fourier transform and overlap addition on the consistency-enhanced spectral sequence to generate the initial enhanced speech sequence; Based on the initial enhanced speech sequence, frame-level gain constraints and silent segment noise floor shaping are performed, and a steady-state enhanced speech sequence is generated by maintaining harmonic amplitude and normalizing the overall amplitude.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the speech enhancement method based on energy spectrum depth modulation as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the speech enhancement method based on energy spectrum depth modulation as described in any one of claims 1 to 8.