Robust denoising processing method and system enhanced by acoustic signal self-supervised learning

By employing self-supervised learning and dynamic adaptive masking strategies, the problems of signal separation and feature preservation in complex noisy environments are solved, enabling accurate noise identification and suppression, and improving the robustness and signal fidelity of the system.

CN121415799BActive Publication Date: 2026-03-31BEIJING GUANYU INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing acoustic signal denoising techniques struggle to maintain stable performance in complex noise environments, and the cost of supervised training based on large amounts of labeled data is prohibitive, leading to signal filtering or distortion and ignoring the temporal continuity and semantic relevance of acoustic signals.

Method used

Employing a self-supervised learning framework and a dynamic adaptive masking strategy, this approach achieves adaptive separation of noise and signal and feature enhancement through time-frequency domain transformation, semantic manifold space separation, and probability propagation network. It also incorporates an iterative optimization mechanism to adjust processing parameters.

Benefits of technology

It significantly improves the ability to identify complex noise, preserves key semantic information, reduces signal distortion, and enhances the system's robustness and real-time application performance under low signal-to-noise ratio conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415799B_ABST
    Figure CN121415799B_ABST
Patent Text Reader

Abstract

The application provides a robust denoising processing method and system enhanced by acoustic wave signal self-supervised learning, relates to the technical field of signal processing, and comprises the following steps: performing feature enhancement on an initial time-frequency representation through a dynamic self-adaptive mask strategy; constructing a self-supervised reconstruction task based on the mask time-frequency representation; separating noise and signal subspaces in a semantic manifold space; combining time-domain continuity features to establish a probability propagation network to model local dependent relationships; and finally generating denoising weights and realizing semantic fidelity optimization. The application can effectively improve the denoising effect and semantic integrity of acoustic wave signals in a noise complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to signal processing technology, and more particularly to a robust noise reduction method and system for acoustic signal self-supervised learning enhancement. Background Technology

[0002] Acoustic signal processing, as a key technology in fields such as communication, speech recognition, and environmental monitoring, plays an important role in daily life and industrial applications. Especially in noisy environments, acoustic signals are often interfered with by various noises, such as background ambient noise, equipment noise, and random interference, which seriously affects the quality and usability of the acoustic signals. Therefore, noise reduction of acoustic signals has become an important research direction in the field of acoustic signal processing.

[0003] Traditional acoustic signal denoising techniques mainly include spectral subtraction, Wiener filtering, and Kalman filtering. With the development of deep learning technology, neural network-based denoising methods such as convolutional neural networks, recurrent neural networks, and attention mechanisms have been widely applied and have achieved certain results. These methods typically achieve noise suppression and signal recovery by establishing a mapping relationship between noise and clean signals.

[0004] Existing methods have limited ability to distinguish between noise and signal characteristics, especially in situations with low signal-to-noise ratios, which can easily lead to the filtering or distortion of useful signals and make it difficult to maintain stable performance in complex and changing noise environments.

[0005] Most methods rely on supervised training with large amounts of labeled data. Obtaining high-quality paired data (noisy signals and their corresponding clean signals) is costly and difficult to achieve, which limits the generalization ability and adaptability of the model in practical applications.

[0006] Existing technologies often neglect the inherent temporal continuity and semantic correlation of sound signals when processing time-frequency domain features, resulting in artifacts such as musical noise or speech distortion during processing, which reduces the naturalness and intelligibility of the processed signal. Summary of the Invention

[0007] The present invention provides a robust noise reduction method and system for acoustic signal self-supervised learning enhancement, which can solve the problems in the prior art.

[0008] A first aspect of the present invention provides a robust noise reduction processing method for acoustic signal self-supervised learning enhancement, comprising:

[0009] The original acoustic wave signal is acquired and transformed in the time-frequency domain to obtain an initial time-frequency representation. A dynamic adaptive masking strategy is used to enhance the features of the initial time-frequency representation to generate a masked time-frequency representation. A self-supervised reconstruction task is constructed based on the masked time-frequency representation. A semantic encoder is trained by minimizing the reconstruction loss to extract the semantic embedding vector of the initial time-frequency representation.

[0010] The semantic embedding vector is mapped to a semantic manifold space, and the semantic similarity matrix between time-frequency units is calculated in the semantic manifold space. Adaptive clustering is performed based on the semantic similarity matrix to obtain the separation boundary between the noise subspace and the signal subspace. A noise confidence distribution is generated based on the separation boundary.

[0011] A probability propagation network is established by combining the noise confidence distribution with the temporal continuity features in the semantic embedding vector. The probability propagation network is used to model local dependencies to obtain a corrected noise confidence distribution. A denoising weight is generated based on the corrected noise confidence distribution. The denoising weight is applied to the initial time-frequency representation to obtain a denoised time-frequency representation.

[0012] The output acoustic signal is obtained by performing an inverse time-frequency transformation on the denoised time-frequency representation. The semantic fidelity deviation of the output acoustic signal in the semantic manifold space is calculated. The mask parameters and coding parameters are iteratively optimized based on the semantic fidelity deviation.

[0013] A dynamic adaptive masking strategy is used to enhance the features of the initial time-frequency representation to generate a masked time-frequency representation. Based on this masked time-frequency representation, a self-supervised reconstruction task is constructed, including:

[0014] A dynamic adaptive masking strategy is used to perform local signal-to-noise ratio estimation on the initial time-frequency representation to obtain a signal-to-noise ratio distribution map at the time-frequency unit level, and a masking strength adjustment function is constructed based on the signal-to-noise ratio distribution map;

[0015] The signal-to-noise ratio (SNR) values ​​in the SNR distribution map are mapped to the mask probabilities of the corresponding time-frequency units using the mask strength modulation function. Differentiated masking operations are performed on different time-frequency units of the initial time-frequency representation based on the mask probabilities to generate masked time-frequency representations. A self-supervised reconstruction task is constructed based on the masked time-frequency representations.

[0016] A dynamic adaptive masking strategy is used to perform local signal-to-noise ratio (SNR) estimation on the initial time-frequency representation to obtain a time-frequency unit-level SNR distribution map. Based on this SNR distribution map, a masking strength adjustment function is constructed, including:

[0017] A dynamic adaptive coding strategy is used to perform multi-scale energy analysis on the initial time-frequency representation. A local analysis window is constructed for each time-frequency unit. An adaptive optimization algorithm is used to determine the optimal scale parameters of the local analysis window. The signal within the local analysis window is decoupled into the main signal component and the noise residual component. The energy ratio of the main signal component and the noise residual component is calculated to obtain the signal-to-noise ratio distribution map at the time-frequency unit level.

[0018] Wavelet coefficients are reconstructed on the signal-to-noise ratio (SNR) distribution map, and SNR variation features are extracted in the time axis and frequency axis directions respectively. The adaptive optimization algorithm is used to learn the SNR threshold parameter, and the separation boundary between signal and noise is constructed by combining the SNR threshold parameter to generate the distribution features of the SNR variation region.

[0019] Based on the separation boundary between the signal and noise and the distribution characteristics of the signal-to-noise ratio variation region, a mask strength mapping relationship is constructed. The adjustment parameters of the mask strength mapping relationship are determined through adaptive optimization iteration, and a mask strength adjustment function with nonlinear characteristics is generated.

[0020] In the semantic manifold space, the semantic similarity matrix between time-frequency units is calculated. Based on the semantic similarity matrix, adaptive clustering is performed to obtain the separation boundary between the noise subspace and the signal subspace. A noise confidence distribution is generated based on the separation boundary, including:

[0021] In the semantic manifold space, a deep association representation between time-frequency units is constructed. The semantic embedding vector of each pair of time-frequency units is mapped to the manifold surface. The evolution trajectory of the semantic embedding vector on the manifold surface is characterized by the Riemann geodesic optimization algorithm. Based on the deep association representation, adaptive weights are generated and filled into the corresponding positions in the semantic similarity matrix.

[0022] By integrating the eigenvalue sequence and eigenvector group of the semantic similarity matrix, structural analysis is performed on the deep association representation. The association strength of the eigenvalue sequence is used as the basis for division, and the subspace dimension division boundary is dynamically determined.

[0023] Based on the subspace dimension partition boundary, the feature vector group is reconstructed into a signal feature subset and a noise feature subset. The boundary modeling of the signal feature subset and the noise feature subset is performed by an adaptive mapping method to generate a subspace separation boundary with topological invariance.

[0024] For each time-frequency unit, the projection distance from its semantic embedding vector to the subspace separation boundary is calculated, and the projection distance is converted into a noise confidence distribution through the adaptive mapping method.

[0025] By using an adaptive mapping method to model the boundaries of the signal feature subset and the noise feature subset, a subspace separation boundary with topological invariance is generated, including:

[0026] For each feature vector in the signal feature subset, a tangent space basis is constructed in the semantic manifold space. Based on the tangent space basis, the projection relationship from the feature vector to the neighboring feature vector is calculated to obtain the tangent vector component. The tangent vector component is used to construct the manifold tangent space expansion coefficient. The manifold tangent space expansion coefficient is subjected to singular value decomposition to obtain the main direction vector group of the signal subset.

[0027] The method of constructing the tangent space basis is continued to process each feature vector in the noise feature subset. Based on the tangent space basis, the tangent vector component of the noise feature vector is calculated. The tangent vector component is mapped to the manifold tangent space to obtain the expansion coefficient. The expansion coefficient is subjected to singular value decomposition to obtain the main direction vector group of the noise subset.

[0028] The subspace angle between the main direction vector group of the signal subset and the main direction vector group of the noise subset is calculated, and the subspace angle is constructed as the boundary direction optimization objective. The gradient field of the boundary implicit function is constructed and solved in combination with the boundary direction optimization objective constraint. A manifold curvature regularization term is introduced into the boundary implicit function so that its zero level set maintains topological connectivity under local perturbation, thereby generating a subspace separation boundary with topological invariance.

[0029] A probability propagation network is established by combining the noise confidence distribution with the temporal continuity features in the semantic embedding vector. This probability propagation network is then used to model local dependencies, resulting in a corrected noise confidence distribution. Noise reduction weights are generated based on this corrected noise confidence distribution, including:

[0030] Temporal features are generated from the semantic embedding vectors. A tensor decomposition model is constructed for the semantic embedding vectors of each time-frequency unit. The correlation between time-frequency units is decoupled into principal components and coupling components. The semantic features of the unit are constructed through the principal components. The dynamic correlation between units is strengthened by the coupling components. The principal components and coupling components are nonlinearly reconstructed to obtain the deep features of the time-frequency units. The mapping relationship between adjacent time-frequency units is shaped based on the deep features.

[0031] A probability propagation network is constructed based on the mapping relationship, and the noise confidence distribution is mapped to the initial state of the network nodes. The propagation weights are obtained by optimizing the deep features through the tensor decomposition model.

[0032] In the probability propagation network, state information from neighboring nodes is aggregated for each node and dynamically reconstructed with the initial state. The state information is adaptively enhanced based on the propagation weights. When the network state reaches a steady state, the final state of the node is transformed to obtain the corrected noise confidence distribution. The corrected noise confidence distribution is then reconstructed into noise reduction weights.

[0033] The output acoustic signal is obtained by performing an inverse time-frequency transform on the denoised time-frequency representation. The semantic fidelity deviation of the output acoustic signal in the semantic manifold space is calculated by:

[0034] A semantic perception criterion is established for the noise reduction time-frequency representation. Based on the semantic perception criterion, the time-frequency domain signal is converted into an output acoustic signal. The output acoustic signal is mapped to the semantic manifold space. The signal features are reconstructed using the semantic perception criterion to generate signal fidelity features. The signal fidelity features in the semantic manifold space are measured and evaluated, and the signal fidelity features are converted into semantic fidelity deviation.

[0035] A second aspect of the present invention provides a robust noise reduction processing system for acoustic signal self-supervised learning enhancement, comprising:

[0036] The first unit is used to acquire the original acoustic wave signal and perform time-frequency domain transformation to obtain an initial time-frequency representation. A dynamic adaptive masking strategy is used to enhance the features of the initial time-frequency representation to generate a masked time-frequency representation. A self-supervised reconstruction task is constructed based on the masked time-frequency representation. A semantic encoder is trained by minimizing the reconstruction loss to extract the semantic embedding vector of the initial time-frequency representation.

[0037] The second unit is used to map the semantic embedding vector to a semantic manifold space, calculate the semantic similarity matrix between time-frequency units in the semantic manifold space, perform adaptive clustering based on the semantic similarity matrix to obtain the separation boundary between the noise subspace and the signal subspace, and generate a noise confidence distribution based on the separation boundary.

[0038] The third unit is used to establish a probability propagation network by combining the noise confidence distribution with the temporal continuity features in the semantic embedding vector, to model local dependencies using the probability propagation network, to obtain a corrected noise confidence distribution, to generate denoising weights based on the corrected noise confidence distribution, and to apply the denoising weights to the initial time-frequency representation to obtain a denoised time-frequency representation.

[0039] The fourth unit is used to perform inverse time-frequency transformation on the noise reduction time-frequency representation to obtain the output acoustic signal, calculate the semantic fidelity deviation of the output acoustic signal in the semantic manifold space, and iteratively optimize the mask parameters and coding parameters based on the semantic fidelity deviation.

[0040] A third aspect of the present invention provides an electronic device, comprising:

[0041] processor;

[0042] Memory used to store processor-executable instructions;

[0043] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0044] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0045] The beneficial effects of this application are as follows:

[0046] This invention significantly improves the ability to identify noise types in complex acoustic environments through a self-supervised learning framework and a dynamic adaptive masking strategy, effectively preserving key semantic information in the original sound wave signal and reducing semantic distortion problems in traditional noise reduction methods.

[0047] The present invention establishes a noise subspace and signal subspace separation mechanism in the semantic manifold space, and combines it with the modeling of time-domain dependencies by a probability propagation network, thereby achieving accurate identification and suppression of time-varying non-stationary noise and improving the robustness of the system under low signal-to-noise ratio conditions.

[0048] This invention is based on an iterative optimization mechanism for semantic fidelity deviation, which enables the system to adaptively adjust processing parameters, maintain stable noise reduction effect in different application scenarios, reduce signal distortion caused by excessive noise reduction, reduce computational complexity, and improve the system performance in real-time applications. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating the robust noise reduction method for acoustic signal self-supervised learning enhancement according to an embodiment of the present invention.

[0050] Figure 2 This is a flowchart illustrating the dynamic boundary construction process for semantic embedding vector adaptive clustering in an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0053] Figure 1 This is a flowchart illustrating the robust noise reduction method for acoustic signal self-supervised learning enhancement according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0054] The original acoustic wave signal is acquired and transformed in the time-frequency domain to obtain an initial time-frequency representation. A dynamic adaptive masking strategy is used to enhance the features of the initial time-frequency representation to generate a masked time-frequency representation. A self-supervised reconstruction task is constructed based on the masked time-frequency representation. A semantic encoder is trained by minimizing the reconstruction loss to extract the semantic embedding vector of the initial time-frequency representation.

[0055] The semantic embedding vector is mapped to a semantic manifold space, and the semantic similarity matrix between time-frequency units is calculated in the semantic manifold space. Adaptive clustering is performed based on the semantic similarity matrix to obtain the separation boundary between the noise subspace and the signal subspace. A noise confidence distribution is generated based on the separation boundary.

[0056] A probability propagation network is established by combining the noise confidence distribution with the temporal continuity features in the semantic embedding vector. The probability propagation network is used to model local dependencies to obtain a corrected noise confidence distribution. A denoising weight is generated based on the corrected noise confidence distribution. The denoising weight is applied to the initial time-frequency representation to obtain a denoised time-frequency representation.

[0057] The output acoustic signal is obtained by performing an inverse time-frequency transformation on the denoised time-frequency representation. The semantic fidelity deviation of the output acoustic signal in the semantic manifold space is calculated. The mask parameters and coding parameters are iteratively optimized based on the semantic fidelity deviation.

[0058] In one optional implementation, a dynamic adaptive masking strategy is used to enhance the features of the initial time-frequency representation to generate a masked time-frequency representation. The self-supervised reconstruction task based on the masked time-frequency representation includes:

[0059] A dynamic adaptive masking strategy is used to perform local signal-to-noise ratio estimation on the initial time-frequency representation to obtain a signal-to-noise ratio distribution map at the time-frequency unit level, and a masking strength adjustment function is constructed based on the signal-to-noise ratio distribution map;

[0060] The signal-to-noise ratio (SNR) values ​​in the SNR distribution map are mapped to the mask probabilities of the corresponding time-frequency units using the mask strength modulation function. Differentiated masking operations are performed on different time-frequency units of the initial time-frequency representation based on the mask probabilities to generate masked time-frequency representations. A self-supervised reconstruction task is constructed based on the masked time-frequency representations.

[0061] Local signal-to-noise ratio (SNR) estimation of the initial time-frequency representation is achieved by calculating the energy distribution of the local region surrounding each time-frequency unit. Specifically, a sliding window, with a size of 5×5 or 7×7, can be set to scan each time-frequency unit of the initial time-frequency representation. For each time-frequency unit (t, f), the average energy within its local window is calculated as a signal component estimate, while the background noise energy level in the entire time-frequency spectrum is calculated as a noise component estimate. The local SNR value of that time-frequency unit is obtained by the ratio of the signal component to the noise component. For example, for an audio signal with a sampling rate of 16 kHz, the initial time-frequency representation obtained after short-time Fourier transform has a time frame length of 25 ms, a frame shift of 10 ms, and 512 frequency points. In this case, the average energy value S(t, f) within its 5×5 window can be calculated for each time-frequency unit (t, f), while the average energy of the 10% of time-frequency units with the lowest energy in the entire time-frequency spectrum is estimated as the noise level N. Thus, the signal-to-noise ratio (SNR) SNR(t, f) for each time-frequency unit is S(t, f) / N. By performing this operation on all time-frequency units, a complete SNR distribution map can be obtained.

[0062] When constructing a mask strength adjustment function based on the signal-to-noise ratio (SNR) distribution map, a piecewise function can be used. The design principle is as follows: for regions with high SNR (representing regions rich in useful signal components), a lower mask probability should be assigned to retain more useful information; for regions with low SNR (representing regions mainly consisting of noise or less important signal components), a higher mask probability should be assigned to encourage the model to learn the characteristics of these regions. In specific implementation, a signal-to-noise ratio (SNR) threshold can be set. low 0dB and SNR high It is 20dB. When the signal-to-noise ratio (SNR) of the time-frequency cell is lower than the SNR low When the signal-to-noise ratio (SNR) is higher than the SNR, the mask strength modulation function outputs a high mask probability value, such as 0.8; when the SNR is higher than the SNR, the mask strength modulation function outputs a high mask probability value. high When the signal-to-noise ratio (SNR) is between two thresholds, a low mask probability value is output, such as 0.2. When the SNR is between two thresholds, the mask probability is calculated using linear interpolation. For example, for a time-frequency cell with an SNR of 10dB, its mask probability is calculated as 0.8 - (10-0) / (20-0)×(0.8-0.2) = 0.5. In this way, the mask strength adjustment function maps the SNR value to the mask probability of the corresponding time-frequency cell.

[0063] After mapping the signal-to-noise ratio (SNR) values ​​in the SNR distribution map to the mask probabilities of the corresponding time-frequency units using a mask strength adjustment function, a differential masking operation is performed. For each time-frequency unit X(t, f) in the initial time-frequency representation X, a random number r is generated based on its mask probability p(t, f), ranging from 0 to 1. If r is less than p(t, f), the time-frequency unit is set to a mask value (such as 0 or a specific mask marker); otherwise, the original value remains unchanged. For example, for a time-frequency unit with a mask probability of 0.6, there is approximately a 60% probability that it will be masked. In this way, regions with high SNR (important signal regions) are more likely to be preserved, while regions with low SNR (minor signal or noise regions) are more likely to be masked. This differential masking operation generates a masked time-frequency representation X. masked In this context, the masking level in different regions is adapted to the importance of their signals.

[0064] For a specific example, suppose there is an audio segment containing speech, lasting 3 seconds. After preprocessing, an initial time-frequency representation X is obtained, with a size of 300×257 (number of time frames × number of frequency points). Through the above local signal-to-noise ratio estimation process, the signal-to-noise ratio distribution map SNR is calculated. map It is also a 300×257 matrix. Using a mask strength adjustment function, the SNR is... map Convert to mask probability map prob map In the differential masking operation, for a time-frequency unit with time frame t=100 and frequency point f=120, if its signal-to-noise ratio (SNR)(100, 120)=15dB, then according to the mask strength adjustment function, its masking probability is 0.35. A random number r=0.42 is generated; since r>0.35, this time-frequency unit remains unchanged. However, for a time-frequency unit with time frame t=150 and frequency point f=80, if its signal-to-noise ratio (SNR)(150, 80)=-5dB, then its masking probability is 0.8. A random number r=0.3 is generated; since r<0.8, this time-frequency unit is masked. The final generated masked time-frequency representation X masked In the high signal-to-noise ratio region, most of the original information is preserved, while in the low signal-to-noise ratio region, most of the information is masked.

[0065] A self-supervised reconstruction task is constructed based on the generated mask time-frequency representation. Specifically, the mask time-frequency representation X can be reconstructed. masked The input is fed into a feature extraction network, which can be a multi-layer convolutional neural network or a transformer structure. The network's task is to attempt to recover the masked time-frequency units, predict their original values, and form a reconstructed time-frequency representation X. recon By minimizing the difference between the original time-frequency representation X and the reconstructed time-frequency representation X... reconThe network is trained using differences in data (such as mean squared error). This self-supervised reconstruction task enables the network to learn the intrinsic structure and features of audio signals, resulting in better performance in downstream tasks. During training, each batch of data generates a new dynamic mask, ensuring the network learns diverse feature representations.

[0066] In one optional implementation, a dynamic adaptive masking strategy is used to perform local signal-to-noise ratio (SNR) estimation on the initial time-frequency representation to obtain a time-frequency unit-level SNR distribution map. The masking strength adjustment function is constructed based on the SNR distribution map, including:

[0067] A dynamic adaptive coding strategy is used to perform multi-scale energy analysis on the initial time-frequency representation. A local analysis window is constructed for each time-frequency unit. An adaptive optimization algorithm is used to determine the optimal scale parameters of the local analysis window. The signal within the local analysis window is decoupled into the main signal component and the noise residual component. The energy ratio of the main signal component and the noise residual component is calculated to obtain the signal-to-noise ratio distribution map at the time-frequency unit level.

[0068] Wavelet coefficients are reconstructed on the signal-to-noise ratio (SNR) distribution map, and SNR variation features are extracted in the time axis and frequency axis directions respectively. The adaptive optimization algorithm is used to learn the SNR threshold parameter, and the separation boundary between signal and noise is constructed by combining the SNR threshold parameter to generate the distribution features of the SNR variation region.

[0069] Based on the separation boundary between the signal and noise and the distribution characteristics of the signal-to-noise ratio variation region, a mask strength mapping relationship is constructed. The adjustment parameters of the mask strength mapping relationship are determined through adaptive optimization iteration, and a mask strength adjustment function with nonlinear characteristics is generated.

[0070] The dynamic adaptive masking strategy is implemented based on a multi-scale energy analysis framework, which performs point-by-point scanning processing on the initial time-frequency representation. The initial time-frequency representation is obtained through a short-time Fourier transform, with a transform window length of 1024 sampling points and a frame shift step of 256 sampling points to ensure that the time-frequency resolution meets engineering application standards. The multi-scale energy analysis employs a sliding window mechanism, expanding the window scale from a minimum of 3×3 time-frequency units to a maximum of 31×31 time-frequency units, with a scale step of 2 units, forming 15 different analysis scales. The local analysis window constructed for each time-frequency unit adopts a rectangular shape, with the window center aligned with the coordinate position of the current time-frequency unit.

[0071] The adaptive optimization algorithm determines the optimal scale parameter for the local analysis window using the energy variance minimization criterion. The algorithm calculates the standard deviation of the time-frequency coefficients within the window at each candidate scale and selects the scale with the smallest standard deviation as the optimal parameter. The standard deviation calculation covers the amplitude values ​​of all time-frequency units within the window. When the standard deviation is less than 0.1, the signal characteristics within the window are considered relatively uniform; when the standard deviation is greater than 0.8, the window is considered to have crossed the signal boundary and the scale needs to be reduced. The optimal scale parameter is limited to a range of 5×5 to 21×21 time-frequency units to avoid statistical instability due to an excessively small window or boundary ambiguity due to an excessively large window.

[0072] The signal decoupling process separates the time-frequency coefficients within a local analysis window into the main signal component and the noise residual component. The main signal component is obtained through amplitude thresholding, which is set to the 75th percentile of the amplitude values ​​within the window. Time-frequency units above the threshold are assigned to the main signal component, while those below the threshold are assigned to the noise residual component. The energy ratio is calculated using the power spectral density summation method; the energy of the main signal component is equal to the sum of the squares of the amplitudes of its time-frequency units, and the energy of the noise residual component is also equal to the sum of the squares of the amplitudes of its time-frequency units. The signal-to-noise ratio at the time-frequency unit level is defined as the ratio of the energy of the main signal component to the energy of the noise residual component. A minimum value is set when the energy of the noise residual component is less than 0.001 to avoid division-by-zero anomalies.

[0073] After the signal-to-noise ratio (SNR) distribution map is constructed, wavelet coefficient reconstruction is performed. The reconstruction uses a bioorthogonal wavelet transform with a decomposition level of four to ensure the capture of SNR variation characteristics across different frequency ranges. The SNR variation characteristics along the time axis are obtained by wavelet decomposition of the SNR sequence for each frequency channel, extracting the energy distribution of the second and third layer wavelet coefficients as time-varying features. Similarly, the SNR variation characteristics along the frequency axis are obtained by wavelet decomposition of the SNR sequence for each time frame, extracting the energy distribution of the second and third layer wavelet coefficients as frequency-varying features.

[0074] The adaptive optimization algorithm learns the signal-to-noise ratio (SNR) threshold parameter using a gradient descent strategy. The initial threshold is set to 1.2 times the median of the SNR distribution, the learning rate is set to 0.01, and the number of iterations is limited to 100. The objective function is defined as minimizing the classification error rate, i.e., minimizing the sum of the proportion of signal components misclassified as noise and the proportion of noise components misclassified as signals. Gradient calculation uses the numerical differencing method, with a step size set to 0.001 times the current threshold value. The iteration is terminated early when the change in the objective function is less than 0.0001 over 10 consecutive iterations.

[0075] The signal-to-noise separation boundary is constructed based on the learned signal-to-noise ratio (SNR) threshold parameter. The separation boundary is represented by contour lines, with the contour line values ​​equal to the optimized SNR threshold. Boundary generation employs a linear interpolation method, inserting intermediate points between adjacent time-frequency units, with the interpolation precision set to one-tenth of the original sampling interval. The distribution characteristics of SNR variation regions are obtained through boundary gradient calculation. Regions with gradients greater than 0.5 are marked as abrupt change regions, regions with gradients less than 0.1 are marked as smooth regions, and the remaining regions are marked as transition regions.

[0076] The mask strength mapping relationship is constructed using a piecewise linear function. The mapping relationship maps the signal-to-noise ratio (SNR) value to a mask strength value between 0 and 1. Time-frequency units with an SNR higher than 1.5 times the threshold correspond to a mask strength of 1.0, while those with an SNR lower than 0.5 times the threshold correspond to a mask strength of 0.0. Linear interpolation is used to determine the mask strength in the intermediate regions. The adaptive optimization iterative adjustment of the mapping relationship parameters includes the slope of the linear segment and the position of the inflection point. The initial value of the slope parameter is set to 2.0, with a range limited to 0.5 to 5.0. The inflection point position parameter is expressed as a multiple of the threshold, with the initial value of the upper inflection point being 1.5 times the threshold and the initial value of the lower inflection point being 0.5 times the threshold.

[0077] The parameter optimization was implemented using a particle swarm optimization algorithm, with a particle swarm size of 20 particles, 50 iterations, an inertia weight of 0.9, and an acceleration constant of 2.0. The fitness function was defined as the weighted sum of the improvement in signal-to-noise ratio and the degree of signal distortion after masking, with a weight ratio of 7:3. A velocity constraint strategy was used for particle position updates during each iteration, with the maximum velocity set to 10% of the parameter value range.

[0078] The mask strength adjustment function with nonlinear characteristics is implemented by introducing a sigmoid activation function. The final mask strength adjustment function combines piecewise linear mapping and sigmoid nonlinear transformation. The steepness parameter of the sigmoid function is set to 5.0, and the center point position corresponds to the signal-to-noise ratio threshold. In practical applications, 16-bit fixed-point numbers are used to represent the mask strength value, with a quantization precision of 1 / 65536, to ensure the stability of numerical calculation.

[0079] In a specific data example, the initial input time-frequency representation is a complex matrix of 256 time frames multiplied by 513 frequency channels, with a sampling frequency of 16000Hz. The signal is speech signal superimposed with Gaussian white noise, and the initial signal-to-noise ratio (SNR) is 5dB. After processing with a dynamic adaptive masking strategy, the time-frequency unit-level SNR distribution map shows that the SNR in the speech fundamental frequency region reaches above 15dB, while the SNR in the high-frequency noise region remains below -5dB. The learned SNR threshold parameter is 8.2dB, and the separation boundary clearly delineates the speech and noise regions. The final generated mask strength adjustment function outputs a mask strength of 0.8 to 1.0 in the active speech region and a mask strength of 0.0 to 0.2 in the pure noise region. After processing, the overall signal SNR is improved to 12.8dB, and the speech intelligibility score is improved by 30%.

[0080] In one optional implementation, the semantic similarity matrix between time-frequency units is calculated in the semantic manifold space, adaptive clustering is performed based on the semantic similarity matrix to obtain the separation boundary between the noise subspace and the signal subspace, and a noise confidence distribution is generated based on the separation boundary, including:

[0081] In the semantic manifold space, a deep association representation between time-frequency units is constructed. The semantic embedding vector of each pair of time-frequency units is mapped to the manifold surface. The evolution trajectory of the semantic embedding vector on the manifold surface is characterized by the Riemann geodesic optimization algorithm. Based on the deep association representation, adaptive weights are generated and filled into the corresponding positions in the semantic similarity matrix.

[0082] By integrating the eigenvalue sequence and eigenvector group of the semantic similarity matrix, structural analysis is performed on the deep association representation. The association strength of the eigenvalue sequence is used as the basis for division, and the subspace dimension division boundary is dynamically determined.

[0083] Based on the subspace dimension partition boundary, the feature vector group is reconstructed into a signal feature subset and a noise feature subset. The boundary modeling of the signal feature subset and the noise feature subset is performed by an adaptive mapping method to generate a subspace separation boundary with topological invariance.

[0084] For each time-frequency unit, the projection distance from its semantic embedding vector to the subspace separation boundary is calculated, and the projection distance is converted into a noise confidence distribution through the adaptive mapping method.

[0085] like Figure 2 As shown, the method includes:

[0086] The deep association representation between time-frequency units in the semantic manifold space is constructed using a high-dimensional embedding vector mapping mechanism. Each time-frequency unit corresponds to a 128-dimensional semantic embedding vector, with vector elements ranging from -1 to +1. The manifold surface is constructed using a local linear embedding method, projecting the high-dimensional semantic embedding vectors onto a 64-dimensional manifold space while preserving local neighborhood relationships. The manifold surface is parameterized using spherical coordinates, with the radius parameter fixed at 1.0, and the angle parameter obtained through principal component analysis for dimensionality reduction. Each pair of time-frequency unit semantic embedding vectors corresponds to two coordinate points on the manifold surface, with the Euclidean distance between these points serving as the initial association metric.

[0087] The Riemann geodesic optimization algorithm characterizes the semantic embedding vector evolution trajectory based on gradient descent. The geodesic path is approximated using piecewise linear approximation, with 50 segments of equal length. The objective function for path optimization is defined as minimizing the total path length, with the constraint that the path's start and end points are fixed. The gradient is calculated using the finite difference method with a step size of 0.001 and an iteration limit of 200. The convergence criterion is that the path length change is less than 0.0001 over 10 consecutive iterations. The evolution trajectory is represented by the path integral length, calculated using a trapezoidal rule with a precision of 6 decimal places. Shorter path lengths indicate higher semantic similarity and thus larger weight values.

[0088] The adaptive weight generation for deep association representation employs a path length inverse mapping strategy. Weights are calculated by multiplying the path length by a normalization factor, which ensures that the sum of all weights equals the total number of time-frequency units. Weight values ​​are then filled into the corresponding positions in the semantic similarity matrix, with the matrix dimension being the square of the total number of time-frequency units. The matrix uses a symmetric structure, with the main diagonal elements set to 1.0 to represent self-similarity. Weight values ​​are maintained at 32-bit floating-point precision to avoid the accumulation of numerical calculation errors. The matrix is ​​stored in a sparse format, only storing non-zero elements and their index positions, reducing memory usage to 15% of the original matrix.

[0089] The semantic similarity matrix eigenvalue decomposition is implemented using the Jacobi iterative method, with an iteration convergence threshold set to 1e-8 and a maximum number of iterations set to 1000. The eigenvalue sequences are arranged in descending order, and the number of eigenvalues ​​equals the matrix dimension. The eigenvector groups form an orthogonal basis, with the magnitude of each eigenvector normalized to 1.0. The correlation strength of the eigenvalue sequences is calculated using the cumulative contribution rate; the number of eigenvalues ​​corresponding to a cumulative contribution rate of 85% is considered the effective dimension. The correlation strength threshold is set to 5% of the total variance; eigenvalues ​​below the threshold correspond to noise components, while those above the threshold correspond to signal components.

[0090] Deep correlation characterization and structural analysis integrate eigenvalue sequences and eigenvector sets. The eigenvalue sequences reflect the importance of each principal component, while the eigenvector sets describe the spatial orientation of the principal components. Structural analysis identifies significant component boundaries by calculating the jump magnitude between eigenvalues; the jump magnitude is defined as the ratio of the difference between adjacent eigenvalues ​​to the larger eigenvalue. A jump magnitude exceeding 0.3 is considered indicative of significant structural changes, corresponding to candidate boundaries for subspace partitioning. Structural stability is verified through eigenvector correlation; an inner product of adjacent eigenvectors greater than 0.8 indicates structural continuity, while a product less than 0.3 indicates structural discontinuity.

[0091] The dynamic determination of subspace dimension boundaries employs the gap statistic method, where the gap statistic is defined as the ratio of the difference between adjacent eigenvalues ​​to the larger eigenvalue. Local maxima in the gap statistic sequence are used as potential partitioning points, and these local maxima must be greater than 1.5 times the adjacent gap statistics. The partitioning boundary is selected such that the signal subspace dimension is no less than 30% and no more than 70% of the total dimension. After the boundary is determined, the signal subspace corresponds to the first N largest eigenvalues, and the noise subspace corresponds to the subsequent eigenvalues. The value of N is determined by the location of the maximum gap statistic. Boundary positions are stored as integer indices, ranging from 30% to 70% of the total dimension.

[0092] The eigenvector group is reconstructed into a signal feature subset and a noise feature subset, with boundaries defined by the subspace dimension. The signal feature subset contains the feature vectors corresponding to the first N eigenvalues, and the noise feature subset contains the feature vectors corresponding to the subsequent eigenvalues. The eigenvector reconstruction employs a Gram-Schmidt orthogonalization process to ensure the reconstructed eigenvector group maintains orthogonality. The reconstruction error is calculated using the inner product of the eigenvectors; the degree to which the inner product deviates from 0 or 1 indicates the reconstruction accuracy, which is controlled within 0.001. The reconstruction process uses an iterative refinement strategy, correcting the accumulated error in each iteration, with no more than 10 iterations.

[0093] The adaptive mapping method employs a support vector machine (SVM) approach for boundary modeling of signal and noise feature subsets. Boundary modeling treats the two feature subsets as a binary classification problem, seeking the optimal separating hyperplane. The hyperplane parameters are solved using the Lagrange multiplier method, with the constraint of maximizing the classification margin. A radial basis function is used as the kernel function, with the kernel parameter set to the square root of the reciprocal of the feature space dimension. The separating hyperplane normal vector and intercept parameter constitute the boundary model, and the model parameters are stored as a 64-dimensional vector. Boundary optimization uses a sequential minimum optimization algorithm, with a convergence tolerance of 0.001 and a maximum of 500 iterations.

[0094] The topologically invariant subspace separation boundary is achieved through continuous deformation preservation. Topological invariance is verified by homotopy equivalence; the boundary maintains its separation characteristics under continuous deformation. The boundary is represented using an implicit function: a function value greater than 0 indicates a signal region, a function value less than 0 indicates a noise region, and a function value equal to 0 indicates a separation boundary. Boundary smoothness is guaranteed by the continuity of the second derivative, with a smoothness parameter set to 0.1 to control the magnitude of boundary curvature variation. Topological verification is achieved by calculating the boundary winding number; a winding number of 0 indicates topological consistency.

[0095] The projection distance from the semantic embedding vector of a time-frequency unit to the subspace separation boundary is calculated using the point-to-hyperplane distance formula. The projection distance equals the inner product of the semantic embedding vector and the hyperplane normal vector, plus the intercept parameter, divided by the normal vector magnitude. A positive distance value indicates that the time-frequency unit belongs to the signal region, while a negative distance value indicates that it belongs to the noise region. The absolute value of the distance represents the confidence level. The projection calculation uses vectorized operations, and a single calculation can process all units of the entire time-frequency representation. The computational complexity is linear time, and the processing speed reaches 1 million time-frequency units per second.

[0096] The adaptive mapping method converts the projected distance into a noise confidence distribution using a sigmoid function transformation. The sigmoid function steepness parameter is set to 2.0, and the center point parameter is set to 0.0. The projected distance serves as the input to the sigmoid function, and the output value ranges from 0 to 1, representing the noise confidence. Smaller distance values ​​correspond to higher noise confidence, and larger distance values ​​correspond to lower noise confidence. The confidence distribution precision is maintained at 16-bit fixed-point, with a quantization step size of 1 / 65536. The mapping function is implemented using a lookup table, pre-calculating function values ​​for 4096 sampling points, with intermediate values ​​obtained through linear interpolation.

[0097] In a specific data example, the input semantic embedding vector matrix has a dimension of 131072×128, corresponding to a time-frequency representation of 256 time frames multiplied by 512 frequency channels. The semantic similarity matrix calculation involves 850 million distance operations, which are reduced to 2.3 seconds using GPU parallel acceleration. Eigenvalue decomposition yields 131072 eigenvalues, with the first 15000 eigenvalues ​​contributing 85% of the total. The gap statistic reaches its maximum value of 0.42 at the 8500th eigenvalue. The subspace dimension division boundary determines the signal subspace dimension to be 8500 and the noise subspace dimension to be 122572. The normal vector of the separating boundary hyperplane has a magnitude of 1.0 and an intercept parameter of 0.15. The projection distance calculation results show that the average distance in the active speech region is 0.8, and the average distance in the noise region is -0.6. After sigmoid transformation, the noise confidence distribution shows that the noise confidence in the speech region is below 0.2, the noise confidence in the pure noise region is above 0.8, and the confidence in the boundary transition region smoothly varies between 0.3 and 0.7. After processing, the overall signal-to-noise ratio of the signal is improved from the original 5dB to 13.2dB, and the subjective speech quality score is improved by 40%.

[0098] In one optional implementation, boundary modeling of the signal feature subset and the noise feature subset using an adaptive mapping method to generate a subspace separation boundary with topological invariance includes:

[0099] For each feature vector in the signal feature subset, a tangent space basis is constructed in the semantic manifold space. Based on the tangent space basis, the projection relationship from the feature vector to the neighboring feature vector is calculated to obtain the tangent vector component. The tangent vector component is used to construct the manifold tangent space expansion coefficient. The manifold tangent space expansion coefficient is subjected to singular value decomposition to obtain the main direction vector group of the signal subset.

[0100] The method of constructing the tangent space basis is continued to process each feature vector in the noise feature subset. Based on the tangent space basis, the tangent vector component of the noise feature vector is calculated. The tangent vector component is mapped to the manifold tangent space to obtain the expansion coefficient. The expansion coefficient is subjected to singular value decomposition to obtain the main direction vector group of the noise subset.

[0101] The subspace angle between the main direction vector group of the signal subset and the main direction vector group of the noise subset is calculated, and the subspace angle is constructed as the boundary direction optimization objective. The gradient field of the boundary implicit function is constructed and solved in combination with the boundary direction optimization objective constraint. A manifold curvature regularization term is introduced into the boundary implicit function so that its zero level set maintains topological connectivity under local perturbation, thereby generating a subspace separation boundary with topological invariance.

[0102] The adaptive mapping method models the boundaries of the signal and noise feature subsets by constructing a tangent space basis in the semantic manifold space. For each feature vector in the signal feature subset, the tangent space basis is constructed using a local coordinate system method, establishing an orthogonal coordinate system at the location of the feature vector. The tangent space basis dimension is set to 64, consistent with the semantic manifold space dimension. The basis vectors are generated through a Gram-Schmidt orthogonalization process, ensuring that the basis vectors are mutually orthogonal and have a magnitude of 1.0. The construction radius of the tangent space basis is set to the average distance of the feature vector's neighborhood, and the neighborhood range is determined using the k-nearest neighbor method, with k set to 10.

[0103] The projection relationship from a feature vector to its neighboring feature vectors is calculated based on the vector dot product operation. The 10 nearest neighboring feature vectors are selected, and Euclidean distance is used as the distance metric. The projection relationship is obtained by calculating the dot product between the current feature vector and its neighboring feature vectors; the dot product value represents the projection strength. The projection relationship matrix has a dimension of 64×10, and the matrix elements range from -1 to +1. The projection calculation uses vectorized operations, allowing for the simultaneous calculation of the projection relationships of all neighboring vectors in a single process, with a computational complexity of linear time.

[0104] The tangent vector components are obtained through linear combinations of projection relations on the tangent space basis. The tangent vector component is calculated as the product of the projection relation matrix and the tangent space basis matrix, resulting in a 64-dimensional vector. The magnitude of the tangent vector components is controlled between 0.1 and 2.0 through normalization to avoid computational instability caused by excessively large or small values. The direction of the tangent vector components reflects the local variation trend of the eigenvectors on the manifold; directional consistency is verified by the angle between adjacent tangent vectors being less than 30 degrees.

[0105] The expansion coefficients of the manifold tangent space are constructed using a weighted combination strategy of tangent vector components. The expansion coefficients are obtained by calculating the inner product of the tangent vector components and predefined weight vectors, where the weight vectors reflect the importance of different directions. The weight vectors are initialized with a uniform distribution, ranging from 0.5 to 1.5. The dimension of the expansion coefficient matrix is ​​equal to the number of eigenvectors in the signal feature subset multiplied by 64, and the matrix uses a sparse storage format to reduce memory usage. The calculation precision of the expansion coefficients is maintained at 32-bit floating-point numbers to avoid accumulated errors affecting subsequent decomposition processes.

[0106] The singular value decomposition (SVD) of the manifold tangent space expansion coefficients is implemented using the SVD algorithm. SVD decomposes the expansion coefficient matrix into the product of three matrices: a left singular vector matrix, a singular value diagonal matrix, and a right singular vector matrix. The singular values ​​are arranged in descending order, and the left singular vectors corresponding to the first 20 largest singular values ​​constitute the principal direction vector group of the signal subset. The principal direction vector group has a dimension of 64×20, and the magnitude of each principal direction vector is normalized to 1.0. The convergence threshold of SVD decomposition is set to 1e-8, and the maximum number of iterations is set to 500.

[0107] The method for constructing the tangent space basis of the noise feature subset continues the approach used for the signal feature subset. Each feature vector in the noise feature subset also employs a local coordinate system method to establish a tangent space basis, maintaining a basis dimension of 64. The tangent space basis construction parameters are consistent with those of the signal feature subset, including a neighborhood range k of 10 and a construction radius equal to the average neighborhood distance. The calculation process for the tangent vector components of the noise feature vectors is the same as that for the signal feature vectors, with a projection relation matrix dimension of 64×10 and a tangent vector component dimension of 64.

[0108] The expansion coefficients of the noise eigenvector tangent components mapped to the manifold tangent space are calculated using the same weighted combination strategy. The weight vectors are consistent with those used in the signal feature subset processing, ensuring the comparability of the processing methods for the two subsets. The dimension of the expansion coefficient matrix is ​​equal to the number of eigenvectors in the noise feature subset multiplied by 64, and the matrix storage format and precision requirements are the same as those for the signal feature subset. The singular value decomposition of the expansion coefficients also uses the SVD algorithm, and the left singular vectors corresponding to the first 20 largest singular values ​​constitute the principal direction vector group of the noise subset.

[0109] The subspace angle between the principal direction vector groups of the signal subset and the principal direction vector group of the noise subset is calculated based on the principal component angle method. The subspace angle is defined as the minimum angle between the subspaces spanned by the two vector groups, obtained by calculating the maximum canonical correlation coefficient between the two vector groups. The canonical correlation coefficient is calculated through the singular value decomposition of the cross-covariance matrix of the two vector groups; the maximum singular value corresponds to the maximum canonical correlation coefficient. The subspace angle is equal to the inverse cosine of the maximum canonical correlation coefficient, and the angle ranges from 0 to 90 degrees.

[0110] The boundary direction optimization objective is to maximize the angle between the subspaces as the objective function. The objective function is defined as the cosine of the angle between the subspaces, aiming to make the two subspaces orthogonal for optimal separation. Optimization constraints include orthogonality and magnitude constraints on the principal direction vectors. The orthogonality constraint requires that all vectors within a vector group be pairwise orthogonal, and the magnitude constraint requires that the magnitude of each vector equals 1.0. The optimization algorithm uses the Lagrange multiplier method, with the initial value of the Lagrange multiplier set to 0.1 and the iteration step size set to 0.01.

[0111] The gradient field of the implicit boundary function is constructed based on the objective constraint of the boundary direction. The gradient field is defined as the vector of partial derivatives of the implicit boundary function with respect to spatial coordinates, and the gradient direction points in the direction of the fastest increase in function value. The gradient field is calculated using the finite difference method, with the difference step size set to one-thousandth of the spatial resolution. The gradient field is solved using the numerical solution of the Poisson equation, with the boundary condition that the normal component of the gradient field on the boundary is equal to the gradient of the objective in the boundary direction.

[0112] The boundary implicit function is solved by integrating the gradient field. The integration path is a straight line from the reference point to the target point, and the integration step size is set to one-hundredth of the path length. The integration calculation uses the fourth-order Runge-Kutta method to ensure the accuracy and stability of the numerical integration. The zero level set of the boundary implicit function corresponds to the separation boundary between the signal subspace and the noise subspace, and the zero level set is obtained through a contour extraction algorithm.

[0113] The manifold curvature regularization term is introduced by adding a second-order differential term to the implicit boundary function. The regularization term is defined as the Laplace operator of the implicit boundary function, reflecting the local curvature characteristics of the function. The regularization strength parameter is set to 0.01 to control the degree of influence of curvature regularization on boundary smoothness. The regularization term is calculated using a five-point difference scheme to ensure second-order precision numerical differentiation.

[0114] Topological connectivity preservation is verified through the stability of the zero-level set under local perturbations. The local perturbation employs a Gaussian noise model, with the noise standard deviation set to 1% of the function's dynamic range. Topological connectivity is verified by calculating the number of connected components in the zero-level set; the unchanged number of connected components before and after the perturbation indicates topological invariance. Connectivity analysis uses a depth-first search algorithm, with the search threshold set to twice the accuracy of the zero-level set extraction.

[0115] In the specific data example, the signal feature subset contains 8500 eigenvectors, each with a dimension of 64. The tangent space basis construction generates 8500 orthogonal matrices of 64×64, and the calculation of the tangent vector components involves 850,000 vector projection operations. The manifold tangent space expansion coefficient matrix has a dimension of 8500×64, and singular value decomposition yields 544,000 singular values, with the top 20 singular values ​​accounting for 92% of the cumulative energy. The noise feature subset contains 122,572 eigenvectors, with a corresponding expansion coefficient matrix of 122,572×64. The calculated subspace angle between the principal direction vector groups of the two subsets is 73.2 degrees, indicating good separability between the two subspaces. Solving the gradient field of the boundary implicit function involves a 4096×4096 grid calculation, with an average gradient field magnitude of 0.85. After manifold curvature regularization, the average curvature of the zero-level set is 0.23, and local perturbation tests show that the number of connected components remains 1, verifying topological invariance. The generated subspace separation boundary improves the separation accuracy of the speech signal region from the noise region to 94.6%, with a boundary smoothness score of 0.92.

[0116] In one optional implementation, a probability propagation network is established by combining the noise confidence distribution with the temporal continuity features in the semantic embedding vector. The probability propagation network is then used to model local dependencies to obtain a corrected noise confidence distribution. Noise reduction weights are generated based on the corrected noise confidence distribution, including:

[0117] Temporal features are generated from the semantic embedding vectors. A tensor decomposition model is constructed for the semantic embedding vectors of each time-frequency unit. The correlation between time-frequency units is decoupled into principal components and coupling components. The semantic features of the unit are constructed through the principal components. The dynamic correlation between units is strengthened by the coupling components. The principal components and coupling components are nonlinearly reconstructed to obtain the deep features of the time-frequency units. The mapping relationship between adjacent time-frequency units is shaped based on the deep features.

[0118] A probability propagation network is constructed based on the mapping relationship, and the noise confidence distribution is mapped to the initial state of the network nodes. The propagation weights are obtained by optimizing the deep features through the tensor decomposition model.

[0119] In the probability propagation network, state information from neighboring nodes is aggregated for each node and dynamically reconstructed with the initial state. The state information is adaptively enhanced based on the propagation weights. When the network state reaches a steady state, the final state of the node is transformed to obtain the corrected noise confidence distribution. The corrected noise confidence distribution is then reconstructed into noise reduction weights.

[0120] Temporal features in the semantic embedding vectors are generated using a sliding window extraction mechanism. Each time-frequency unit corresponds to a 128-dimensional semantic embedding vector, and the temporal features are calculated by the rate of change of the semantic embedding vectors within adjacent time frames. The sliding window length is set to 5 time frames, and the window step size is set to 1 time frame to ensure the continuity and smoothness of the temporal features. The temporal feature dimension is set to 32 dimensions, obtained by principal component analysis (PCA) of the semantic embedding vectors within the window. The PCA process retains principal components with a cumulative variance contribution rate of 95% to ensure that the temporal features contain sufficient information.

[0121] The tensor decomposition model constructs a semantic embedding vector for each time-frequency unit using a third-order tensor representation. The tensor dimensions are set to 128×5×5, corresponding to the semantic dimension, time window, and frequency window, respectively. Tensor decomposition employs the CANDECOMP-PARAFAC decomposition method, decomposing the third-order tensor into a linear combination of multiple rank tensors. The decomposition rank is set to 16 to ensure sufficient expressive power while maintaining computational efficiency. The decomposition convergence threshold is set to 1e-6, and the maximum number of iterations is set to 200.

[0122] The decoupling of relationships between time-frequency units is achieved through factor matrices of tensor decomposition, separating principal components from coupled components. Principal components correspond to the top 8 rank-one tensors with the largest weights in the tensor decomposition, reflecting the main semantic features of the time-frequency units. Coupled components correspond to the bottom 8 rank-one tensors with smaller weights, characterizing the complex interactions between time-frequency units. A weight threshold is set at 80% of the total weight; components above the threshold are classified as principal components, and components below the threshold are classified as coupled components.

[0123] Semantic feature construction is based on principal component factor matrix reconstruction. The principal component factor matrix is ​​reconstructed into a 64-dimensional semantic feature vector through weighted summation, with the weight coefficients equal to the decomposition weights of the corresponding rank tensor. The semantic feature vector is L2 normalized to ensure that the vector magnitude is 1.0. The semantic features reflect the core semantic content of the time-frequency unit and are used for subsequent similarity calculation and cluster analysis.

[0124] Dynamic correlation enhancement is achieved through nonlinear transformation of the coupled components. The factor matrix of the coupled components is processed by a tanh activation function to enhance the nonlinear interaction characteristics between components. The steepness parameter of the activation function is set to 2.0 to ensure appropriate nonlinear strength. The enhanced coupled components are used to construct dynamic correlation weights between time-frequency units, with weight values ​​ranging from 0 to 1.

[0125] The nonlinear reconstruction of principal components and coupled components is achieved using a multilayer perceptron structure. The reconstruction network contains two hidden layers with 128 and 64 neurons respectively. The ReLU activation function is used to avoid the vanishing gradient problem. The network is trained using the Adam optimizer with a learning rate of 0.001 and a batch size of 128. The training objective is to minimize the reconstruction error, which is calculated using the mean squared error loss function.

[0126] Deep features are generated by reconstructing the network's output layer. The deep feature dimension is set to 32, containing fused information from principal components and coupling components. Batch normalization is applied to ensure the stability of the feature distribution. The feature value range is limited to between -2 and +2 through clipping operations to avoid the impact of numerical anomalies on subsequent processing.

[0127] The mapping relationship between adjacent time-frequency units is established based on the similarity calculation of deep features. The similarity calculation adopts the cosine similarity method to calculate the deep feature similarity between the current time-frequency unit and all its 8-connected neighbor units. The mapping relationship strength is equal to the similarity value, which ranges from -1 to +1. The mapping relationship is stored in a sparse matrix format, and only the connections with an absolute similarity value greater than 0.1 are saved.

[0128] The probability propagation network is constructed using a graph structure based on the mapping relationship between time-frequency units. Network nodes correspond to time-frequency units, and edge weights correspond to the strength of the mapping relationship. The network adopts an undirected graph structure, and the edge weights are symmetrically processed to ensure the symmetry of the network. Network connectivity is verified through maximum connected component analysis, and the number of nodes in a connected component should be no less than 95% of the total number of nodes.

[0129] The noise confidence distribution is mapped to the initial state of network nodes through direct assignment. The initial state value of each node is equal to the noise confidence value of the corresponding time-frequency unit, and the state value ranges from 0 to 1. The dimension of the initial state vector is equal to the total number of network nodes, and the vector is stored using 32-bit floating-point numbers to ensure numerical precision.

[0130] The propagation weights are optimized through further processing of deep features using a tensor decomposition model. The optimization process employs a gradient descent algorithm, and the objective function is defined as minimizing information loss during propagation. Automatic differentiation is used for gradient calculation to avoid the complexity and errors of manual differentiation. The optimization learning rate is set to 0.01, the momentum parameter to 0.9, and the weight decay coefficient to 1e⁻⁴.

[0131] In the probabilistic propagation network, node state information aggregation employs a weighted averaging strategy. Each node collects state information from its neighboring nodes, and the aggregation weight is equal to the propagation weight of the corresponding edge. The neighboring state information and the node's initial state are dynamically reconstructed through a linear combination, with the combination coefficients obtained through adaptive learning. The initial value of the combination coefficients is set to 0.5, and the learning rate is set to 0.001.

[0132] Adaptive augmentation processing performs a nonlinear transformation on the state information based on propagation weights. The augmentation function uses the sigmoid activation function, and the steepness parameter is dynamically adjusted according to the propagation weights. Larger propagation weights correspond to a larger steepness parameter, resulting in a more significant augmentation effect. The steepness parameter ranges from 1.0 to 5.0 and is calculated from the propagation weights through a linear mapping.

[0133] Network steady-state detection is achieved by judging the magnitude of state changes between consecutive iterations. The magnitude of state changes is defined as the L2 norm of the difference in state values ​​of all nodes in adjacent iterations. The steady-state criterion is that the magnitude of state changes is less than 1e-5 for 10 consecutive iterations. The maximum number of iterations is set to 1000 to avoid infinite loops. The iteration process adopts a synchronous update strategy, where the states of all nodes are updated simultaneously.

[0134] The corrected noise confidence distribution is obtained through the final state transformation of nodes in the steady state of the network. The final state value of the node is directly used as the corrected noise confidence value, maintaining the value range between 0 and 1. The correction process uses the sigmoid function for boundary constraints to ensure the validity of the output value. The correction effect is evaluated by the degree of difference from the initial noise confidence distribution, and the degree of difference is calculated using Kullback-Leibler divergence.

[0135] The noise reduction weight reconstruction is based on the inverse transform of the corrected noise confidence distribution. The noise reduction weight is defined as 1 minus the corrected noise confidence value, ensuring that regions with high noise confidence correspond to smaller noise reduction weights. The weight values ​​are non-linearly adjusted through a power function transformation, with the power exponent set to 0.8 to enhance the contrast of the weight distribution. Weight normalization ensures that the sum of all weight values ​​equals the total number of time-frequency units.

[0136] In the specific data example, the input semantic embedding vector matrix has dimensions of 256×512×128, corresponding to 256 time frames, 512 frequency channels, and 128-dimensional semantic features. Temporal feature generation produces a 256×512×32 feature matrix, and tensor decomposition involves optimization calculations of 167 million parameters. Principal component weights account for 82.3% of the total weights, and coupled component weights account for 17.7%. The deep feature reconstruction network converges after 50 epochs of training, with the training loss reduced to 0.003. The probabilistic propagation network contains 131,072 nodes and 524,288 edges, with a network connectivity of 98.7%. Propagation weight optimization converges after 100 iterations, with a weight distribution standard deviation of 0.15. Network state propagation reaches steady state after 245 iterations, with the state change amplitude reduced to 8.3e-6. The KL divergence of the noise confidence distribution before and after correction is 0.23, indicating a significant correction effect. The generated noise reduction weights averaged 0.85 in the active speech region and 0.12 in the noisy region. After processing, the signal-to-noise ratio increased from the initial 5.2dB to 14.7dB, and the objective speech quality score improved by 45%.

[0137] In one optional implementation, performing an inverse time-frequency transform on the noise-reduced time-frequency representation to obtain an output acoustic signal, and calculating the semantic fidelity deviation of the output acoustic signal in the semantic manifold space includes:

[0138] A semantic perception criterion is established for the noise reduction time-frequency representation. Based on the semantic perception criterion, the time-frequency domain signal is converted into an output acoustic signal. The output acoustic signal is mapped to the semantic manifold space. The signal features are reconstructed using the semantic perception criterion to generate signal fidelity features. The signal fidelity features in the semantic manifold space are measured and evaluated, and the signal fidelity features are converted into semantic fidelity deviation.

[0139] For noise reduction time-frequency representation, a semantic perception criterion is established. This criterion is designed based on the perceptual characteristics of the human auditory system in understanding the semantic content of sound, including the assessment of the degree of preservation of spectral contours, transient characteristics, and harmonic structure. Specifically, the semantic perception criterion assigns importance to different regions in the time-frequency domain through a weighting function. For example, in speech signals, the mid-frequency band from 300Hz to 3000Hz is given higher weight because it contains most of the semantic information; while for music signals, harmonic structure and low-frequency rhythm are given higher weight. This criterion can be represented as a set of weight vectors describing the relationship between signal features and semantic content, such as a weight of 0.7 for spectral details, 0.8 for transient features, and 0.6 for harmonic structure.

[0140] Based on the established semantic awareness criteria, the denoised time-frequency representation is converted into the output acoustic signal. The conversion process employs inverse short-time Fourier transform (ISFT) technology, considering the importance assessment of different time-frequency regions by the semantic awareness criteria during the conversion. For example, when processing speech signals, the system prioritizes the reconstruction quality of frequency bands containing key semantic information (such as consonant regions). During the conversion, an overlap-addition method is used to reduce discontinuities caused by frame-by-frame processing, and a Hanning window function is employed for smoothing. The window length is set to 512 sampling points, with an overlap rate of 75%, to ensure the temporal continuity of the acoustic signal.

[0141] After obtaining the output acoustic signal, it is mapped to a semantic manifold space. This mapping process is implemented using a deep neural network, which consists of four layers: the first layer is a time-frequency analysis layer, using 64 filter banks to extract time-frequency features; the second layer is a feature extraction layer, using 128 convolutional kernels to extract local patterns; the third layer is a context integration layer, using two layers of 256-unit bidirectional gated recurrent units to handle temporal dependencies; and the fourth layer is a mapping layer, mapping the extracted features to a 64-dimensional semantic manifold space. The network learns semantic feature representations through pre-training. The training dataset contains approximately 10,000 hours of multi-scene speech and music data to capture the semantic features of different types of sounds. For example, a speech signal "The weather is nice today" is represented as a 64-dimensional feature vector in the semantic manifold space after mapping, where the first 20 dimensions mainly represent the identity features of the speech, the middle 20 dimensions represent the semantic content, and the last 24 dimensions represent the acoustic environment features.

[0142] In the semantic manifold space, signal features are reconstructed using semantic perception criteria to generate signal fidelity features. The reconstruction process is implemented through a feature transformation network, which receives feature vectors from the semantic manifold space and reconstructs an ideal semantic feature representation based on the weight allocation of the semantic perception criteria. The reconstruction network employs a three-layer fully connected structure, with each layer containing 128, 96, and 64 neurons respectively, and ReLU activation function. For example, when processing denoised speech segments, the system uses the representation of the original signal in the semantic manifold space as a reference to calculate the feature differences between the reconstructed signal and the reference signal; for instance, the cosine similarity in the semantic content dimension reaches 0.92, and the similarity in the timbre feature dimension is 0.85.

[0143] When evaluating signal fidelity features in the semantic manifold space, a multi-dimensional evaluation index is used. First, a semantic consistency index is calculated to measure the similarity between the output signal and the original signal in the semantic content dimension. Second, a perceptual quality index is calculated to evaluate the clarity and naturalness of the output signal. Finally, a comprehensive semantic fidelity score is calculated. During the evaluation process, different weights are assigned to different types of distortion; for example, semantic content distortion has a weight of 0.6, timbre distortion has a weight of 0.25, and residual background noise has a weight of 0.15. In a specific example, a processed speech signal obtains a semantic consistency score of 0.88, a perceptual quality score of 0.92, and a comprehensive semantic fidelity score of 0.89.

[0144] The signal fidelity features are converted into semantic fidelity deviation. This conversion process is based on a predefined reference standard, calculating the deviation value as the difference between the overall score and the ideal reference value (usually set to 1.0). The system also considers the different needs of various application scenarios, performing normalization through an adjustable threshold function. For example, in a teleconferencing system, a semantic fidelity score of 0.89 is converted to a deviation value of 0.11; after scenario adaptation adjustment, considering the inherent limitations of telephone audio quality, the adjusted deviation value is 0.08, indicating that the sound quality has reached a good level in this application scenario. This deviation value can be used to guide system parameter optimization. When the deviation value exceeds a preset threshold (e.g., 0.15), the system triggers a re-noise reduction or parameter adjustment process to improve the semantic fidelity of the output signal.

[0145] A second aspect of the present invention provides a robust noise reduction processing system for acoustic signal self-supervised learning enhancement, comprising:

[0146] The first unit is used to acquire the original acoustic wave signal and perform time-frequency domain transformation to obtain an initial time-frequency representation. A dynamic adaptive masking strategy is used to enhance the features of the initial time-frequency representation to generate a masked time-frequency representation. A self-supervised reconstruction task is constructed based on the masked time-frequency representation. A semantic encoder is trained by minimizing the reconstruction loss to extract the semantic embedding vector of the initial time-frequency representation.

[0147] The second unit is used to map the semantic embedding vector to a semantic manifold space, calculate the semantic similarity matrix between time-frequency units in the semantic manifold space, perform adaptive clustering based on the semantic similarity matrix to obtain the separation boundary between the noise subspace and the signal subspace, and generate a noise confidence distribution based on the separation boundary.

[0148] The third unit is used to establish a probability propagation network by combining the noise confidence distribution with the temporal continuity features in the semantic embedding vector, to model local dependencies using the probability propagation network, to obtain a corrected noise confidence distribution, to generate denoising weights based on the corrected noise confidence distribution, and to apply the denoising weights to the initial time-frequency representation to obtain a denoised time-frequency representation.

[0149] The fourth unit is used to perform inverse time-frequency transformation on the noise reduction time-frequency representation to obtain the output acoustic signal, calculate the semantic fidelity deviation of the output acoustic signal in the semantic manifold space, and iteratively optimize the mask parameters and coding parameters based on the semantic fidelity deviation.

[0150] A third aspect of the present invention provides an electronic device, comprising:

[0151] processor;

[0152] Memory used to store processor-executable instructions;

[0153] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0154] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0155] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for robust denoising processing enhanced by self-supervised learning of acoustic signals, characterized in that, The method comprises the following steps: obtaining an original sound wave signal and performing time-frequency domain transformation to obtain an initial time-frequency representation, using a dynamic adaptive mask strategy to enhance the features of the initial time-frequency representation to generate a mask time-frequency representation, constructing a self-supervised reconstruction task based on the mask time-frequency representation, training a semantic encoder by minimizing the reconstruction loss, and extracting a semantic embedding vector of the initial time-frequency representation; mapping the semantic embedding vector to a semantic manifold space, calculating a semantic similarity matrix between time-frequency units in the semantic manifold space, performing adaptive clustering based on the semantic similarity matrix to obtain a separation boundary of a noise subspace and a signal subspace, and generating a noise confidence distribution according to the separation boundary, comprising: constructing a deep association representation between time-frequency units in the semantic manifold space, mapping the semantic embedding vectors of each pair of time-frequency units to a manifold surface, describing the evolution trajectory of the semantic embedding vectors on the manifold surface by using a Riemannian geodesic optimization algorithm, and generating an adaptive weight based on the deep association representation and filling it into the corresponding position in the semantic similarity matrix; fusing the eigenvalue sequence and eigenvector group of the semantic similarity matrix, performing structural analysis on the deep association representation, taking the association strength of the eigenvalue sequence as the division basis, and dynamically determining the subspace dimension division boundary; based on the subspace dimension division boundary, reconstructing the eigenvector group into a signal feature subset and a noise feature subset, and modeling the boundary of the signal feature subset and the noise feature subset by using an adaptive mapping method to generate a subspace separation boundary with topological invariance; for each time-frequency unit, calculate the projection distance of its semantic embedding vector to the subspace separation boundary, and convert the projection distance to a noise confidence distribution by using the adaptive mapping method; establishing a probability propagation network for the noise confidence distribution combined with the time-domain continuity feature in the semantic embedding vector, modeling the local dependency relationship by using the probability propagation network to obtain a corrected noise confidence distribution, generating a denoising weight based on the corrected noise confidence distribution, and applying the denoising weight to the initial time-frequency representation to obtain a denoised time-frequency representation; performing inverse time-frequency transformation on the denoised time-frequency representation to obtain an output sound wave signal, calculating the semantic fidelity deviation of the output sound wave signal in the semantic manifold space, and iteratively optimizing the mask parameters and the encoding parameters according to the semantic fidelity deviation.

2. The method of claim 1, wherein, The method comprises the following steps: using a dynamic adaptive mask strategy to enhance the features of the initial time-frequency representation to generate a mask time-frequency representation, and constructing a self-supervised reconstruction task based on the mask time-frequency representation comprises: using a dynamic adaptive mask strategy to estimate the signal-to-noise ratio of the initial time-frequency representation to obtain a signal-to-noise ratio distribution map at the time-frequency unit level, and constructing a mask intensity control function based on the signal-to-noise ratio distribution map; mapping the signal-to-noise ratio value in the signal-to-noise ratio distribution map to the mask probability of the corresponding time-frequency unit by using the mask intensity control function, performing differential mask operation on different time-frequency units of the initial time-frequency representation based on the mask probability to generate a mask time-frequency representation, and constructing a self-supervised reconstruction task based on the mask time-frequency representation.

3. The method of claim 2, wherein, The initial time-frequency representation is locally estimated by a dynamic adaptive masking strategy to obtain a time-frequency unit level signal-to-noise ratio distribution map, and a masking intensity control function is constructed based on the signal-to-noise ratio distribution map, including: The initial time-frequency representation is analyzed by a multi-scale energy analysis using a dynamic adaptive coding strategy, a local analysis window is constructed for each time-frequency unit, the optimal scale parameter of the local analysis window is determined using an adaptive optimization algorithm, the signal in the local analysis window is decoupled into a signal main component and a noise residual component, the energy ratio of the signal main component and the noise residual component is calculated to obtain a time-frequency unit level signal-to-noise ratio distribution map; The signal-to-noise ratio distribution map is reconstructed by wavelet coefficients, the signal-to-noise ratio change characteristics are extracted in the time axis direction and the frequency axis direction respectively, the signal-to-noise ratio threshold parameter is learned using the adaptive optimization algorithm, the signal and noise separation boundary is constructed combined with the signal-to-noise ratio threshold parameter, and the distribution characteristics of the signal-to-noise ratio change area are generated; According to the signal and noise separation boundary and the distribution characteristics of the signal-to-noise ratio change area, a masking intensity mapping relationship is constructed, the control parameters of the masking intensity mapping relationship are determined by adaptive optimization iteration, and a masking intensity control function with nonlinear characteristics is generated.

4. The method of claim 1, wherein, The boundary modeling of the signal feature subset and the noise feature subset is performed by an adaptive mapping method, and a subspace separation boundary with topological invariance is generated, including: For each feature vector in the signal feature subset, a tangent space base is constructed in the semantic manifold space, the projection relationship of the feature vector to the neighborhood feature vector is calculated based on the tangent space base to obtain a tangent vector component, the manifold tangent space expansion coefficient is constructed using the tangent vector component, and the signal subset main direction vector group is obtained by singular value decomposition of the manifold tangent space expansion coefficient; The construction method of the tangent space base is continued to process each feature vector in the noise feature subset, the tangent vector component of the noise feature vector is calculated based on the tangent space base, the tangent vector component is mapped to the manifold tangent space to obtain an expansion coefficient, and the noise subset main direction vector group is obtained by singular value decomposition of the expansion coefficient; Based on the signal subset main direction vector group and the noise subset main direction vector group, the subspace included angle thereof is calculated, the subspace included angle is constructed as a boundary direction optimization target, a boundary implicit function gradient field is constructed and solved by combining the boundary direction optimization target constraint, a manifold curvature regularization term is introduced in the boundary implicit function to make its zero level set maintain topological connectivity under local disturbance, and a subspace separation boundary with topological invariance is generated.

5. The method of claim 1, wherein, The noise confidence distribution is combined with the time domain continuity feature in the semantic embedding vector to establish a probability propagation network, the local dependency relationship is modeled using the probability propagation network to obtain a corrected noise confidence distribution, and a denoising weight is generated based on the corrected noise confidence distribution, including: generate a temporal feature from the semantic embedding vector, construct a tensor decomposition model for the semantic embedding vector of each time-frequency unit, decouple the correlation between time-frequency units into a principal component and a coupled component, construct a semantic feature of a unit through the principal component, and strengthen the dynamic correlation between units by using the coupled component, nonlinearly reconstruct the principal component and the coupled component to obtain a deep feature of the time-frequency unit, and shape a mapping relationship between adjacent time-frequency units based on the deep feature; construct a probability propagation network based on the mapping relationship, map the noise confidence distribution to the initial state of the network node, and optimize the deep feature through the tensor decomposition model to obtain a propagation weight; each node in the probability propagation network converges state information from the neighborhood nodes and dynamically reconstructs the initial state based on the propagation weight, and when the network state reaches a steady state, the final state of the conversion node is obtained to obtain a corrected noise confidence distribution, and the corrected noise confidence distribution is reconstructed into a denoising weight.

6. The method of claim 1, wherein, perform inverse time-frequency transformation on the denoised time-frequency representation to obtain an output sound wave signal, and calculate the semantic fidelity deviation of the output sound wave signal in the semantic manifold space, including: establish a semantic perception criterion for the denoised time-frequency representation, convert the time-frequency domain signal to obtain an output sound wave signal based on the semantic perception criterion, map the output sound wave signal to the semantic manifold space, reconstruct the signal feature using the semantic perception criterion, and generate a signal fidelity feature; and measure and evaluate the signal fidelity feature in the semantic manifold space, and convert the signal fidelity feature into a semantic fidelity deviation.

7. Robust denoising processing system with acoustic signal self-supervised learning enhancement, for implementing the method of any one of the preceding claims 1-6, characterized in that, including: a first unit configured to obtain an original sound wave signal and perform time-frequency domain transformation to obtain an initial time-frequency representation, enhance the features of the initial time-frequency representation using a dynamic adaptive mask strategy to generate a mask time-frequency representation, construct a self-supervised reconstruction task based on the mask time-frequency representation, train a semantic encoder by minimizing the reconstruction loss, and extract a semantic embedding vector of the initial time-frequency representation; a second unit configured to map the semantic embedding vector to a semantic manifold space, calculate a semantic similarity matrix between time-frequency units in the semantic manifold space, perform adaptive clustering based on the semantic similarity matrix to obtain a separation boundary of a noise subspace and a signal subspace, and generate a noise confidence distribution according to the separation boundary; a third unit configured to establish a probability propagation network for the noise confidence distribution in combination with the temporal continuity feature in the semantic embedding vector, model local dependency relationships using the probability propagation network, obtain a corrected noise confidence distribution, generate a denoising weight based on the corrected noise confidence distribution, and apply the denoising weight to the initial time-frequency representation to obtain a denoised time-frequency representation; a fourth unit configured to perform inverse time-frequency transformation on the denoised time-frequency representation to obtain an output sound wave signal, calculate the semantic fidelity deviation of the output sound wave signal in the semantic manifold space, and iteratively optimize the mask parameters and the encoding parameters according to the semantic fidelity deviation.

8. An electronic device, comprising: including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored by the memory to perform the method of any one of claims 1 to 6.

9. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by a processor, implement the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice noise reduction method and device based on artificial intelligence, equipment and storage medium

    CN114694674A

  • Man-machine interaction voice perception method and system based on gradient intelligent dispatch subnet pool

    CN121148370A