A speech enhancement method and device based on flow matching, equipment and medium

By using a stream matching-based speech enhancement method, the problems of poor fidelity and high latency in existing technologies are solved, achieving efficient speech enhancement in the financial and medical fields, improving the fidelity of speech details and reducing latency.

CN120913576BActive Publication Date: 2026-05-12平安科技(上海)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
平安科技(上海)有限公司
Filing Date
2025-08-26
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing speech enhancement technologies suffer from poor fidelity and high inference latency, making them particularly difficult to meet the requirements of real-time scenarios, especially in the financial and medical fields.

Method used

A stream-matching-based speech enhancement method is adopted, which enhances noisy speech data through time-frequency transformation, normalization processing, stream-matching input construction, feature encoding of the stream-matching model, and numerical solution of ordinary differential equations.

Benefits of technology

It improves the fidelity of voice details and reduces latency, making it suitable for real-time voice processing needs in the medical and financial fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913576B_ABST
    Figure CN120913576B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a voice enhancement method and device based on flow matching, equipment and a medium, which can be applied to the fields of finance and medicine. The method comprises the following steps: acquiring noise voice data, performing time-frequency conversion processing and standardization processing on the noise voice data, and obtaining standardized time-frequency representation data; performing flow matching input construction processing on the basis of the standardized time-frequency representation data, inputting the standardized time-frequency representation data into a flow matching model for processing, and obtaining enhanced time-frequency representation data; performing anti-standardization processing on the enhanced time-frequency representation data, and obtaining anti-standardized time-frequency representation data; and performing time-domain reconstruction processing on the anti-standardized time-frequency representation data, and obtaining enhanced voice waveform data. In the application, the problem that the existing voice enhancement method is generally poor in fidelity and long in inference time delay caused by discrete quantization can be solved by introducing flow matching modeling and time-frequency-time-domain consistent processing chain to obtain enhanced voice waveform data, so that the voice detail fidelity is improved and the time delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device and medium for speech enhancement based on stream matching. Background Technology

[0002] Existing speech enhancement technologies mainly fall into three categories: First, discrete modeling methods based on language models (LMs) first map speech to discrete tokens before modeling. However, the discrete quantization process inevitably introduces quantization loss, causing artifacts and a decrease in sound quality, which in turn affects speaker similarity and intelligibility. Second, methods based on diffusion models (such as CDiffuSE, SGMSE, and StoRM) typically rely on a large number of stepwise inversion iterations to complete the reconstruction from noise to data distribution. The numerous inference steps and long processing time make it difficult to meet the requirements of real-time scenarios. Third, generative methods based on VAEs or GANs, while improving perceptual quality, suffer from complex model structures, poor training stability, and are prone to problems such as pattern collapse or overly smoothed synthesis results. In summary, existing technologies in speech representation suffer from poor fidelity and high inference latency, which are problematic for applications in the financial and medical fields. Summary of the Invention

[0003] This invention provides a speech enhancement method, apparatus, device, and medium based on stream matching to solve the technical problems of poor fidelity and large inference delay caused by discrete quantization, which are common in existing speech enhancement methods.

[0004] Firstly, a speech enhancement method based on stream matching is provided, including:

[0005] Noisy speech data is acquired and time-frequency transform is performed to obtain noise time-frequency representation data;

[0006] The noise time-frequency representation data is standardized to obtain standardized time-frequency representation data.

[0007] Based on the standardized time-frequency representation data, stream matching input construction processing is performed to obtain stream matching input representation data;

[0008] The flow matching input representation data is input into the feature encoding submodule of the flow matching model to obtain flow matching feature representation data; context modeling processing is performed on the flow matching feature representation data to obtain context-enhanced feature representation data; the context-enhanced feature representation data is input into the velocity field estimation submodule to output velocity field estimation data; the time-frequency state increment is calculated based on the velocity field estimation data and a preset numerical step size parameter to obtain time-frequency state increment data; the time-frequency state increment data is used to numerically solve the ordinary differential equation for the intermediate state time-frequency tensor in the flow matching input representation data to obtain the updated intermediate state time-frequency tensor; the updated intermediate state time-frequency tensor is used as input to the flow matching model to update until the target path parameter is reached, thus obtaining enhanced time-frequency representation data;

[0009] The enhanced time-frequency representation data is denormalized to obtain denormalized time-frequency representation data;

[0010] The denormalized time-frequency representation data is subjected to time-domain reconstruction processing to obtain enhanced speech waveform data.

[0011] Secondly, a stream-matching-based speech enhancement device is provided, comprising:

[0012] The noise acquisition module is used to acquire noisy speech data and perform time-frequency transformation processing to obtain noise time-frequency representation data;

[0013] The noise processing module is used to standardize the noise time-frequency representation data to obtain standardized time-frequency representation data.

[0014] The data processing module is used to perform stream matching input construction processing based on the standardized time-frequency representation data to obtain stream matching input representation data;

[0015] The model optimization module is used to input the flow matching input representation data into the feature encoding submodule of the flow matching model to obtain flow matching feature representation data; perform context modeling processing on the flow matching feature representation data to obtain context-enhanced feature representation data; input the context-enhanced feature representation data into the velocity field estimation submodule to output velocity field estimation data; calculate the time-frequency state increment based on the velocity field estimation data and a preset numerical step size parameter to obtain time-frequency state increment data; use the time-frequency state increment data to numerically solve the ordinary differential equation for the intermediate state time-frequency tensor in the flow matching input representation data to obtain the updated intermediate state time-frequency tensor; and use the updated intermediate state time-frequency tensor as input to update the flow matching model until the target path parameter is reached to obtain enhanced time-frequency representation data.

[0016] The data back-inference module is used to perform denormalization processing on the enhanced time-frequency representation data to obtain denormalized time-frequency representation data;

[0017] The reconstruction output module is used to perform time-domain reconstruction processing on the denormalized time-frequency representation data to obtain enhanced speech waveform data.

[0018] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described stream-matching-based speech enhancement method.

[0019] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described stream-matching-based speech enhancement method.

[0020] In the aforementioned scheme implemented by a stream matching-based speech enhancement method, apparatus, device, and medium, noisy speech data can be acquired and subjected to time-frequency transformation processing to obtain noisy time-frequency representation data; the noisy time-frequency representation data can be standardized to obtain standardized time-frequency representation data; stream matching input construction processing can be performed based on the standardized time-frequency representation data to obtain stream matching input representation data; the stream matching input representation data can be input into a stream matching model for velocity field estimation and numerical solution of ordinary differential equations to obtain enhanced time-frequency representation data; the enhanced time-frequency representation data can be de-standardized to obtain de-standardized time-frequency representation data; and the de-standardized time-frequency representation data can be time-domain reconstructed to obtain enhanced speech waveform data. In this invention, addressing the common problems of poor fidelity and large inference delay caused by discrete quantization in existing speech enhancement methods, enhanced speech waveform data can be obtained by introducing stream matching modeling and a time-frequency-time domain consistent processing chain, thereby improving speech detail fidelity and reducing latency. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating a speech enhancement method based on stream matching in one embodiment of the present invention;

[0023] Figure 2 yes Figure 1 A detailed implementation flow diagram of step S10 Figure 1 ;

[0024] Figure 3 yes Figure 1 A detailed implementation flow diagram of step S20 Figure 2 ;

[0025] Figure 4 yes Figure 1 A detailed implementation flow diagram of step S30 Figure 3 ;

[0026] Figure 5 yes Figure 1 A detailed implementation flow diagram of step S40 Figure 4 ;

[0027] Figure 6 yes Figure 1 A schematic flowchart of a specific implementation method for step S50 Figure 5 ;

[0028] Figure 7 yes Figure 1 A detailed implementation flow diagram of step S60 Figure 6 ;

[0029] Figure 8 This is a schematic diagram of a speech enhancement device based on stream matching in one embodiment of the present invention;

[0030] Figure 9 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0031] Figure 10 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Please see Figure 1 As shown, Figure 1 A flowchart illustrating a stream-matching-based speech enhancement method provided in this embodiment of the invention includes the following steps:

[0034] S10: Acquire noisy speech data and perform time-frequency transformation processing to obtain noise time-frequency representation data.

[0035] Audio input streams can be received at medical or financial acquisition terminals. The file header and encoding / decoding parameters are parsed to obtain noisy speech data. The sampling rate of the noisy speech data is unified and the channels are mapped (for example, medical speech is unified to 16kHz mono, and financial counter recordings can be unified to 8 / 16kHz mono). After time-frequency transformation processing, the noise time-frequency representation data can be obtained, with an attached business domain identifier (such as medical / financial) for subsequent selection of appropriate engineering parameters.

[0036] Combination Figure 2 As shown, step S10 specifically includes the following steps:

[0037] S101: Receives the audio input stream and parses the file header information to obtain the raw noisy speech data.

[0038] The system receives a continuous audio input stream from the acquisition end, parses the container and file header information to obtain the raw noisy speech data. The parsed content includes: sampling rate, bit depth, number of channels, encoding format (e.g., PCM / WAV, FLAC, AAC, Opus), duration, timestamp, and optional service domain identifier and session identifier (e.g., department / equipment identifier in medical scenarios or agent / work order number in financial scenarios).

[0039] S102: Perform sampling rate unification and channel conversion processing on the original noise speech data to obtain formatted noise speech data.

[0040] Sampling rate unification can be achieved using band-limited interpolation or multiphase filtering resampling strategies. For example, in medical scenarios, the unification can be 16kHz mono, while in financial scenarios, it can be 8kHz or 16kHz mono. Channel conversion supports selecting the main channel from stereo or merging it into mono according to weights. After unification, the samples are normalized to a unified numerical domain (such as floating-point [-1,1]) to obtain formatted noisy speech data, ensuring numerical consistency in subsequent time-frequency processing.

[0041] S103: Perform DC bias removal and pre-emphasis processing on the formatted noisy speech data to obtain preprocessed noisy speech data.

[0042] DC bias removal is accomplished by calculating and subtracting the mean of the whole or segment. Pre-emphasis can be achieved by using a first-order high-pass pre-emphasis filter, whose differential form is: the current sample minus the weight of the previous sample. The weight coefficient α can be between 0.95 and 0.98 (e.g., α = 0.97), and finally the preprocessed noisy speech data is obtained.

[0043] S104: The preprocessed noisy speech data is divided into frames based on a preset window length and frame shift to obtain a noisy speech frame sequence.

[0044] The window length and frame shift are determined based on the service domain and recognition link configuration. For example, the window length is 20 to 32 ms (25 ms is commonly used), and the frame shift is 8 to 12 ms (10 ms is commonly used). A zero-padding strategy is used for samples that are less than one frame at the end to ensure that the length of each frame is consistent, forming a sequence of noisy speech frames arranged in chronological order.

[0045] S105: Apply a preset window function to the noisy speech frame sequence frame by frame to obtain a windowed noisy speech frame sequence.

[0046] The window function can be a Hamming window, a Hanning window, or other equivalent smoothing windows, with the window length consistent with the frame length in step S104 above. To avoid energy mismatch, a window energy normalization strategy can also be configured in the implementation to maintain a consistent engineering scale of energy before and after windowing, ultimately obtaining a windowed noisy speech frame sequence.

[0047] S106: Perform a short-time Fourier transform on the windowed noisy speech frame sequence to obtain a complex spectral matrix.

[0048] Perform an N-point Discrete Fourier Transform on each frame of the windowed noisy speech frame sequence (N is not less than the frame length, preferably a power of 2, such as 512 / 1024). If necessary, use zero padding in the time or frequency domain to meet frequency resolution requirements. Stack the complex spectra of each frame in the order of "time × frequency" to form a complex spectrum matrix, whose elements include real and imaginary part information.

[0049] S107: Calculate the amplitude spectrum and phase spectrum of the complex spectrum matrix and combine them according to a preset time-frequency characteristic format to obtain noise time-frequency representation data.

[0050] The amplitude spectrum can be represented by a complex modulus or its logarithmic form, while the phase spectrum is calculated from complex arguments; both are arranged from low to high frequency and from earliest to latest time. The specific encoding format for the feature combination can include representation as a dual-channel "amplitude-phase" tensor, or a split structure of "amplitude tensor + phase tensor." In medical and financial scenarios, the number of frequency bands, frequency truncation range, and quantization accuracy can be configured at the engineering level without changing the mathematical definitions. The final output noise time-frequency representation data serves as the direct input for subsequent standardization processing and stream matching input construction.

[0051] S20: Standardize the noise time-frequency representation data to obtain standardized time-frequency representation data.

[0052] By extracting amplitude-related features from noisy time-frequency representation data and performing standardization, standardized time-frequency representation data can be obtained. For medical and financial scenarios, independent statistical parameter caches can be maintained to improve cross-domain adaptability without changing the standardized calculation process.

[0053] Combination Figure 3As shown, step S20 specifically includes the following steps:

[0054] S201: Extract the amplitude spectrum matrix and phase spectrum matrix from the noise time-frequency representation data respectively to obtain the amplitude spectrum matrix and phase spectrum matrix to be standardized.

[0055] The amplitude spectrum matrix to be standardized can be taken as the modulus of the complex spectrum or the square root of the power spectrum; the phase spectrum matrix is ​​the argument of the complex spectrum. For ease of subsequent statistics, it is stored as a two-dimensional tensor of "time frame × frequency point (or Mel band)," and the service domain identifier (medical / financial) and acquisition channel information are retained for selecting subsequent statistical parameter sources.

[0056] S202: Perform numerical stabilization on the amplitude spectrum matrix to be standardized to obtain a stabilized amplitude spectrum matrix.

[0057] A preferred approach is a combination of lower bound truncation and the addition of a small constant: the amplitude value is truncated to a lower bound to avoid zero values, and then a small constant is added before calculating the logarithm or ratio to suppress numerical divergence, ultimately yielding a stable amplitude spectrum matrix. In medical scenarios, considering that low-energy components such as equipment noise floor and breathing sounds are more common, the lower bound and small constant can be set to more conservative values ​​than in financial scenarios; in financial scenarios (such as 8kHz call audio), the lower bound can be lowered accordingly to avoid excessively raising the noise floor.

[0058] S203: Perform dynamic range compression processing based on the stabilized amplitude spectrum matrix to obtain a compressed amplitude spectrum matrix.

[0059] Compression strategies can be logarithmic amplitude mapping or power-law compression, resulting in a compressed amplitude spectrum matrix. Logarithmic amplitude mapping performs a logarithmic operation on the stabilized amplitude to compress high-energy differences and improve the discriminability of low-energy speech details. Power-law compression performs a non-linear mapping on the stabilized amplitude using an exponent less than 1 to achieve dynamic range compression approximating logarithmic compression. In medical scenarios, logarithmic amplitude mapping is preferred to preserve low-energy details such as heart sounds / breathing sounds. In financial scenarios, power-law compression or logarithmic mapping can be used to adapt to the speech-dominated spectral distribution, while limiting the maximum dynamic range.

[0060] S204: Calculate the mean and standard deviation parameters of the compression amplitude spectrum matrix according to the preset dimensions to obtain standardized statistical parameters.

[0061] Standardized statistical parameters are obtained by statistically analyzing the mean and standard deviation of the compressed amplitude spectrum matrix over time according to preset dimensions (e.g., independent by frequency point or independent by frequency band). Two strategies can be employed for statistical analysis: sliding window or whole-segment statistics. First, in the medical scenario, long-term statistics are preferably maintained at the department / equipment level (for stable equipment characteristics), supplemented by sliding window statistics for the current session to adapt to scenario changes. Second, in the financial scenario, long-term statistics are preferably maintained at the agent / call line level, with short-window updates based on call rounds to adapt to changes in rate and speaker. When the current session duration is insufficient to support stable statistics, a fallback to global prior statistics (pre-trained) for the corresponding business domain is implemented to ensure the availability of standardized parameters.

[0062] S205: The compressed amplitude spectrum matrix is ​​normalized by using the standardized statistical parameters to obtain a normalized amplitude spectrum matrix.

[0063] To avoid the amplification effect caused by extremely small variances, a lower limit threshold can be set for the standard deviation. The normalized matrix is ​​consistent with the original matrix in the "time frame × frequency point (or Mel band)" dimension and retains the same time index as the phase spectrum matrix, thus obtaining a normalized amplitude spectrum matrix to ensure consistency in subsequent coding combinations. If necessary, statistical parameters (mean, standard deviation) and service domain identifiers can be cached together for subsequent de-normalization and cross-segment consistency maintenance.

[0064] S206: Based on a preset time-frequency feature format, the normalized amplitude spectrum matrix and the phase spectrum matrix are encoded and combined to obtain standardized time-frequency representation data.

[0065] Based on a preset time-frequency characteristic format, the normalized amplitude spectrum matrix and phase spectrum matrix are encoded and combined to obtain standardized time-frequency representation data. Preferred encoding formats include:

[0066] Dual-channel overlay format: Overlaying in the channel dimension to form a tensor of size "time frame × frequency point × 2", where channel 0 is the normalized amplitude spectrum and channel 1 is the phase spectrum;

[0067] Split format: Maintain the "amplitude tensor" and "phase tensor" separately, and record their alignment index in the metadata;

[0068] Phase encoding extension: While keeping the phase information unchanged, the phase spectrum can be replaced with its sine / cosine pairs according to a preset format to improve numerical continuity.

[0069] In medical scenarios, a dual-channel overlay format can be prioritized to simplify edge-side inference deployment; in financial scenarios, if integration with existing recognition front-ends is required, a split format can be used, with splicing completed at the interface layer. The final output standardized time-frequency representation data serves as input for subsequent stream matching input construction.

[0070] S30: Perform stream matching input construction processing based on the standardized time-frequency representation data to obtain stream matching input representation data.

[0071] By extracting a portion of data from the standardized time-frequency representation data and performing stream matching input construction processing, we can obtain stream matching input representation data.

[0072] Combination Figure 4 As shown, step S30 specifically includes the following steps:

[0073] S301: Extract the feature tensor corresponding to the model input size from the standardized time-frequency representation data to obtain the standardized time-frequency tensor.

[0074] From the standardized time-frequency representation data, feature tensors with the same input size as the stream matching model are extracted to obtain the standardized time-frequency tensor. The input size is represented by a three-dimensional structure of "number of time frames × number of frequency points × number of channels", where the channel dimension includes normalized amplitude and phase correlation encoding. For samples that do not meet the target size, zero-padding or center pruning in the time dimension and alignment strategies in the frequency dimension are used to reshape them. Longer time windows can be configured for medical scenarios (e.g., 16kHz sampling), while tighter time windows can be configured for financial scenarios (e.g., 8 / 16kHz call audio), but both remain consistent with the static input size of the model.

[0075] S302: Generate a random noise tensor with zero mean and unit variance based on the standardized time-frequency tensor to obtain the initial time-frequency noise tensor.

[0076] The random number generator sets a deterministic seed at the session level to ensure repeatability; the tensor elements follow an independent and identically distributed Gaussian distribution, and the tensor dimension is strictly consistent with the normalized time-frequency tensor without any scaling or offset, resulting in a noisy time-frequency initial tensor after processing.

[0077] S303: Sample path parameters within a preset interval, and perform linear interpolation between the initial noise time-frequency tensor and the standardized time-frequency tensor to obtain the intermediate state time-frequency tensor.

[0078] The path parameter t is sampled within a preset interval (preferably [0,1]); the path parameter can be sampled according to a uniform distribution or a preset density function (e.g., Beta family) to form different interpolation coverages. Linear interpolation is performed using the initial noisy time-frequency tensor Z and the normalized time-frequency tensor X to obtain the intermediate-state time-frequency tensor Xt, where Xt = (1-t)·Z + t·X. To ensure numerical stability, the interpolation is performed in the floating-point domain and uses the same data type and memory layout as the input tensor.

[0079] Step S303 specifically includes the following steps:

[0080] S3031: Analyze the start and end boundaries of the preset interval and configure the sampling strategy to obtain the path parameters.

[0081] The processing unit reads the start and end values ​​of the path parameter space from the configuration storage, verifies their data types and value ranges, and confirms that they satisfy the sequential relationship and are within the allowed range. Then, it determines the sampling strategy based on the task configuration. The sampling strategy can be equal-interval sampling, segmented encrypted sampling, or adaptive sampling based on historical distribution. In scenarios requiring repeatability, a random number seed is set to ensure the reproducibility of the sampling. After completing the strategy initialization, the processing unit performs sampling within the preset range formed by the start and end values ​​to obtain at least one target path parameter. When only a single interpolation is needed, a single path parameter is directly output; when batch processing is required, a sequence of path parameters arranged in the generation order is output, along with an index for subsequent steps.

[0082] S3032: Perform relative position conversion and range constraints on the target path parameters respectively to obtain interpolation weights.

[0083] The processing unit receives the path parameters (or path parameter sequence) output in step S3031 and calculates their corresponding scaling coefficients based on the relative positions of the path parameters between the start and end boundaries. These scaling coefficients characterize whether the parameter is closer to the start or end value. To avoid exceeding limits, the processing unit applies upper and lower bound constraints to the scaling coefficients, limiting them to a closed interval between zero and one. When the scaling coefficient is less than the lower bound, the lower bound is used; when it is greater than the upper bound, the upper bound is used. In batch processing scenarios, each item in the sequence is converted and constrained one by one, and a one-to-one mapping relationship is established with the original path parameters to form interpolation weights, ensuring consistent tensor processing and synthesis according to the weights in subsequent steps.

[0084] S3033: Based on the interpolation weights, the noise time-frequency initial tensor and the normalized time-frequency tensor are aligned in terms of size, channels, and step size to obtain an aligned tensor set.

[0085] The system reads two types of tensors to be processed: the initial noisy time-frequency tensor and the normalized time-frequency tensor, and verifies whether their dimensional descriptions (batch size, number of channels, number of time steps, and number of frequency points) and memory layout are consistent with the model input specifications. If inconsistencies exist, alignment is performed in the following order: First, size alignment: when the time or frequency dimensions are inconsistent, a time and frequency grid consistent with the model is used first. The smaller side is padded with no information or boundary repetition, and the larger side is truncated according to the center or first / last pruning strategy. Second, channel alignment: when the number of channels is not synchronized, channel selection, channel duplication, or linear mapping is performed according to the model channel definition to ensure that the two tensors correspond one-to-one in the channel dimension. Third, step size alignment: when the time step size or frequency resolution is inconsistent, resampling is performed according to the reference step size of the model input to align the two tensors on the same time and frequency grid. During the above alignment process, interpolation weights are used to determine the priority order and bias selection during boundary pruning and padding to reduce the impact on the consistency of subsequent synthesis. After the above processing, the output is a set of aligned tensors consisting of two sets of aligned tensors.

[0086] S3034: Perform linear weighted synthesis on the aligned tensor set to obtain the intermediate state time-frequency tensor.

[0087] Based on the mapping relationship with the interpolation weights, a weighted sum is performed on the corresponding elements of the aligned two types of tensors: the complementary part of the weights is used to measure the share of the initial noisy time-frequency tensor, and the weights themselves are used to measure the share of the normalized time-frequency tensor. In batch scenarios, the weight list is traversed item by item to generate a corresponding set of intermediate state results. To ensure numerical stability, the processing unit truncates outliers during the synthesis process, replaces missing data with zero or nearest neighbor values, and performs a consistency check and data type normalization after synthesis. Finally, the intermediate state time-frequency tensor with the same size as the model input is output. When there are multiple path parameters, an intermediate state time-frequency tensor is formed in order of path parameter index, which can be directly called by the subsequent time coding module and velocity field estimation submodule.

[0088] S304: Perform frame block division and position index encoding on the intermediate state time-frequency tensor to obtain an intermediate state time-frequency tensor with position encoding.

[0089] The intermediate-state time-frequency tensor is partitioned along the time and frequency dimensions to obtain a regular grid of blocks. The block size and step size are fixed in the model configuration (e.g., time dimension Bt × frequency dimension Bf, step size St, Sf), and zero-padding is used to fill in insufficient boundary regions. Subsequently, a two-dimensional position index is generated based on the start and center indices of the block in the global time-frequency plane. The above index is mapped to a position encoding vector (two-dimensional sine and cosine position encoding or equivalent learnable encoding can be used) through a position encoding function, and aligned with the tensor fragment of the corresponding block to obtain the intermediate-state time-frequency tensor with position encoding. The medical and financial domains can use the same encoding function, with only the block size and step size configured differently to adapt to different speech time-frequency resolutions.

[0090] S305: Input the intermediate state time-frequency tensor with position encoding and the path parameters into the time encoding module, and output the condition vector.

[0091] The intermediate state time-frequency tensor with position encoding and the path parameter t are input into the time encoding module, which outputs a condition vector. The time encoding module maps the path parameter t to a fixed-dimensional time embedding (e.g., time embedding implemented through a sine / cosine time-frequency basis or a multilayer perceptron), and aligns it with position-encoded statistics (such as normalized statistics of the frame block index) when needed, ensuring that the condition vector is compatible with subsequent network inputs in the channel dimension. In both the medical and financial domains, the time embedding dimension and frequency coverage can be adjusted to meet different real-time requirements, without changing the functional relationship of "t → time embedding → condition vector".

[0092] S306: The condition vector and the intermediate state time-frequency tensor with position encoding are concatenated by features and processed by channel mapping to obtain the stream matching input representation data.

[0093] The conditional vector and the intermediate state time-frequency tensor with positional encoding are concatenated along the channel dimension to form an extended feature tensor. Then, channel mapping is used to map this extended feature tensor to the desired number of input channels for the flow matching model, resulting in the flow matching input representation data. Channel mapping can employ pointwise convolution (1×1 convolution) or an equivalent linear mapping layer, keeping the time and frequency dimensions unchanged and adjusting only the channel dimension. To ensure cross-domain consistency, the weights and normalization strategies of the channel mapping layer are uniformly fitted during training and executed according to the same graph during inference, without altering their mathematical form due to changes in the business domain.

[0094] S40: Input the flow matching input representation data into the flow matching model to perform velocity field estimation and numerical solution of ordinary differential equations to obtain enhanced time-frequency representation data.

[0095] The input representation data of the stream matching model is fed into the feature encoding and context modeling submodule (e.g., using a time-frequency Transformer backbone) for processing to obtain enhanced time-frequency representation data. To meet the real-time processing needs of medical consultations and financial call centers, the corresponding numerical step size and iteration rounds can be configured according to the edge computing power.

[0096] Combination Figure 5 As shown, step S40 specifically includes the following steps:

[0097] S401: Input the stream matching input representation data into the feature encoding submodule of the stream matching model to obtain stream matching feature representation data.

[0098] The generated stream is matched to the input representation data (which includes the intermediate state time-frequency tensor M). t The location, time encoding, and path parameter t) are input to the feature encoding submodule of the flow matching model. A two-dimensional time-frequency encoding backbone based on DiT (DiffusionTransformer) is preferred.

[0099] First, the intermediate state time-frequency tensor is linearly projected (or 1×1 convolution) into frame blocks to obtain the initial embedding;

[0100] Second, the time encoding (generated by the path parameter t) and the two-dimensional position encoding (from the start and end index of the frame block) are added or concatenated and fused with the initial embedding;

[0101] Third, the flow matching feature representation data is obtained through several Transformer blocks (multi-head self-attention + feedforward network + layer normalization + residual connection).

[0102] In medical scenarios, a larger time attention window can be configured to cover long statements; in financial scenarios, the window can be tightened to reduce latency on the client side, without changing the coding process.

[0103] S402: Perform context modeling processing on the stream matching feature representation data to obtain context-enhanced feature representation data.

[0104] Context modeling preferably considers the correlation between the time and frequency axes simultaneously: time-axis self-attention aggregates the dependencies of cross-frame speech units (syllables / words); frequency-axis self-attention aggregates harmonic and formant structures; optional local convolutional feedforward (Conv-FFN) enhances neighborhood robustness. To reduce the differences between different recording environments, domain embeddings (learnable vectors of medical / financial business domain identifiers) can be introduced at this stage and lightly adjusted through gating (such as FiLM / gated residuals) without changing the form of the self-attention operator, ultimately obtaining context-enhanced feature representation data.

[0105] S403: Input the context-enhanced feature representation data into the velocity field estimation submodule and output the velocity field estimation data.

[0106] Input the context-enhanced feature representation data into the velocity field estimation submodule, and output the velocity field estimation data v. θ (M t ,t,C). Where: v θ The head network consists of several layers of 1×1 convolutional MLPs, and its output is related to M... t A time-frequency velocity field of the same shape; t is the path parameter; C is an optional condition input.

[0107] When the scenario provides textual conditions, a text encoder (such as a ConvNeXtV2 text branch or an equivalent semantic encoder) can be used to generate a semantic vector T(C). This vector is then fused with the feature backbone through cross-attention or channel modulation to enhance semantic consistency under extremely low signal-to-noise ratios. Technical terminology readings in medical scenarios and key information verification statements in financial scenarios can both serve as sources of C. If no textual conditions are provided, the branch is left empty without affecting the solution chain.

[0108] S404: Calculate the time-frequency state increment based on the velocity field estimation data and the preset numerical step size parameter to obtain the time-frequency state increment data.

[0109] The time-frequency state increment ΔM is calculated based on the velocity field estimation data and the preset numerical step size parameter h. t Finally, the time-frequency state increment data is obtained:

[0110] ΔM t =h·v θ (M t ,t,C)

[0111] h can adopt a fixed step size (such as 1 / K) or a segmented step size strategy (larger in the early stage and smaller in the later stage). To balance real-time performance and stability, K≈10–16 can be used in medical scenarios and K≈6–12 can be used in financial scenarios. When the input energy is detected to be too low or the speech rate changes abruptly, h can be pruned to control numerical oscillation.

[0112] S405: Using the time-frequency state increment data, numerically solve the ordinary differential equation for the intermediate state time-frequency tensor in the stream-matched input representation data to obtain the updated intermediate state time-frequency tensor.

[0113] The intermediate state time-frequency tensor is updated by performing ODE numerical solution using the time-frequency state increment, resulting in the updated intermediate state time-frequency tensor, as follows: Explicit Euler: M t+h =M t +ΔM tAlternatively, an improved Euler-Heyn approach could be used: predict first, then correct to reduce truncation error. On servers with sufficient computing power, RK4 can also be used, but Euler or Heyn is preferred for mobile / edge devices. During the solution process, the time and frequency dimensions of the tensor remain unchanged; only the channel values ​​are updated; and compared with the generated M... t The memory layout is consistent.

[0114] S406: The updated intermediate state time-frequency tensor is used as input to the flow matching model to update it until the target path parameters are reached, thereby obtaining enhanced time-frequency representation data.

[0115] Using the updated intermediate state time-frequency tensor as the new input, repeat steps S401–S405 until the path parameters reach the target endpoint (t→1), outputting the enhanced time-frequency representation data. During training, linear interpolation is used:

[0116] M t = (1-t)·M y +t·M x

[0117] Sampling intermediate states and minimizing the velocity field for supervision:

[0118]

[0119] During inference, the target end is reached through K forward numerical integration steps from the noise end (or the initial value end), typically less than 20 steps. The only differences between medical and financial scenarios are the engineering values ​​of K and h, and whether or not the text condition C is enabled, without changing the unified processing flow of "feature encoding - context modeling - velocity field estimation - numerical integration - iterative convergence".

[0120] S50: Perform denormalization processing on the enhanced time-frequency representation data to obtain denormalized time-frequency representation data.

[0121] By performing denormalization on the enhanced time-frequency representation data, we can obtain denormalized time-frequency representation data.

[0122] Combination Figure 6 As shown, step S50 specifically includes the following steps:

[0123] S501: Extract the enhanced normalized amplitude spectrum matrix and the enhanced phase spectrum matrix from the enhanced time-frequency representation data, respectively.

[0124] The enhanced normalized amplitude spectrum matrix is ​​the normalized amplitude channel, and the enhanced phase spectrum matrix is ​​the phase channel corresponding one-to-one with the time frame and frequency point of the enhanced normalized amplitude spectrum matrix. To ensure the consistency of subsequent denormalization, the normalized metadata bound to the current session is read and verified (including the dimension used to calculate the mean and standard deviation, the sliding window length, and the business domain identifier, etc.); if the session-level metadata is missing, it falls back to the global prior metadata of the corresponding business domain (medical / financial).

[0125] S502: Perform an inverse transformation on the enhanced normalized amplitude spectrum matrix to obtain the inverse normalized amplitude spectrum matrix.

[0126] Optimal execution is performed per frequency band (or per frequency point) to ensure consistency with the statistical dimensions in the standardization phase; a lower bound truncation is used for extremely small standard deviations to avoid numerical amplification, resulting in an inversely normalized amplitude spectrum matrix. In medical scenarios, long-term statistical parameters at the department / equipment level can be prioritized to improve consistency across time periods; in financial scenarios, seat / line-level statistical parameters can be used to adapt to call characteristics.

[0127] S503: Perform inverse compression processing on the inverse normalized amplitude spectrum matrix according to the preset dynamic range compression configuration parameters to obtain the inverse compressed amplitude spectrum matrix.

[0128] If logarithmic amplitude compression is used, exponential mapping and bias backoff are performed; if power-law compression is used, the inverse transform of the corresponding exponent is performed, resulting in the inverse compressed amplitude spectrum matrix. To avoid overflow of high-energy points after the inverse transform, an upper limit is set for output pruning, and the pruning ratio is recorded for quality monitoring. In medical scenarios, a wider low-energy resolution range is typically retained to cover details such as breathing / heart sounds; in financial scenarios, under voice-dominated conditions, the maximum dynamic range can be limited to suppress line noise playback.

[0129] S504: Perform numerical restoration processing on the inverse compression amplitude spectrum matrix to obtain the restored amplitude spectrum matrix.

[0130] Numerical reconstruction of the inverse compression amplitude spectrum matrix yields the restored amplitude spectrum matrix. This process includes: removing small stability constants introduced during the normalization phase; applying zero-threshold correction to negative values; performing unit uniformity according to the system-defined numerical domain (e.g., linear amplitude or power square root amplitude); and performing band-limited smoothing when necessary to suppress isolated spikes introduced by the inverse transform. In medical settings, mild band regularization can be enabled to smooth the equipment noise floor, while in financial settings, out-of-band suppression can be performed according to telecommunications standards.

[0131] S505: Encode and combine the restored amplitude spectrum matrix and the enhanced phase spectrum matrix to obtain candidate inverse normalized time-frequency representation data.

[0132] The restored amplitude spectrum matrix and the enhanced phase spectrum matrix are encoded and combined to obtain candidate inversely normalized time-frequency representation data. The encoding format is consistent with the time-frequency feature format in step S20: that is, combined into amplitude and phase by channel dimension; the amplitude tensor and phase tensor are output separately, and the alignment index is retained in the metadata. In addition, without changing the phase value, it can be stored in the form of [sinφ, cosφ] to improve the numerical continuity (whether to enable it depends on the business domain strategy). The alignment relationship between the time frame and the frequency point is strictly maintained during combination to ensure that the subsequent inverse short-time Fourier transform can be directly consumed.

[0133] S506: Perform dimensional rearrangement on the candidate denormalized time-frequency representation data to obtain denormalized time-frequency representation data.

[0134] The goal of dimensional rearrangement is to align with the input conventions of the time-domain reconstruction module (e.g., rearranging the data into a continuous memory layout of "frequency × time × channel" or "time × frequency × channel"), while simultaneously filling in the spectral symmetry and boundary frames required for reconstruction. If real-valued FFT is used in steps S10 and S20 and only positive frequencies are retained, then the conjugate frequency band is reconstructed here according to the convention. The zero-padding portion of the last frame is explicitly labeled so that window function energy compensation can be performed during the reconstruction stage. In medical scenarios, the output typically outputs a data layout compatible with a 16kHz / 16bit reconstruction link; in financial scenarios, it can be aligned with the ISTFT implementation of an 8 / 16kHz call link. The final denormalized time-frequency representation data is used for complex synthesis and inverse short-time Fourier transform.

[0135] S60: Perform time-domain reconstruction processing on the denormalized time-frequency representation data to obtain enhanced speech waveform data.

[0136] The denormalized time-frequency representation data is reconstructed in the time domain to generate enhanced speech waveform data for archiving, quality inspection, or downstream recognition.

[0137] Combination Figure 7 As shown, step S60 specifically includes the following steps:

[0138] S601: Extract the inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix from the inverse normalized time-frequency representation data, respectively.

[0139] To ensure consistency with the preceding link, the metadata carried by the denormalized time-frequency representation data (such as FFT points, window function type, frame shift, and whether only positive frequencies are saved) is extracted and verified. In the medical domain, the same Mel / line spectrum configuration as the front end is used for 16kHz mono data; in the financial domain, an 8 / 16kHz line spectrum configuration can be used, yielding the denormalized amplitude spectrum matrix and denormalized phase spectrum matrix.

[0140] S602: Perform complex number synthesis processing on the inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix to obtain the reconstructed complex spectrum matrix.

[0141] The inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix are combined using complex numbers to obtain the reconstructed complex spectrum matrix. Specifically, the amplitude and phase can be mapped to a complex plane representation by frame or by frequency point, and the negative frequency band is supplemented according to the conjugate symmetry principle when only positive frequencies are preserved; DC and Nyquist frequencies are treated as real numbers.

[0142] S603: Perform inverse short-time Fourier transform on the reconstructed complex spectrum matrix to obtain a windowed time-domain frame sequence.

[0143] The inverse short-time Fourier transform can employ FFT points and window functions (such as Hamming / Hanning windows), and the implementation should enable platform-consistent real-valued IFFT optimization. For boundary frequencies caused by frequency truncation or zero-padding, consistent complex symmetry should be maintained to avoid time-domain leakage.

[0144] S604: Perform windowing compensation and overlap addition processing on the windowed time-domain frame sequence to obtain a continuous time-domain speech sequence.

[0145] Windowing compensation and overlap addition are performed on the windowed time-domain frame sequence to obtain a continuous time-domain speech sequence. The WOLA (Weighted Overlap-Add) strategy is preferred.

[0146] First, windowing compensation: Amplitude correction is performed on each frame by window power weighting to ensure that the reconstructed energy is consistent with the frequency domain;

[0147] Second, overlapping and addition: adjacent frames are superimposed according to the same frame shift as the front end, and insufficient frames at the end are handled according to the zero-tail strategy.

[0148] Third, boundary processing: smooth the transition area between the beginning and end to suppress amplitude abrupt changes at the splicing point.

[0149] The medical domain can be configured with longer windows (such as 25ms, 10ms frame shift) to maintain smooth detail; the financial domain can use a more compact window shift combination to reduce edge latency, but without changing the WOLA process.

[0150] S605: Perform DC bias correction on the continuous time-domain speech sequence to obtain corrected time-domain speech data.

[0151] The correction method can be to calculate the moving average and subtract it, supplemented by a first-order high-pass micro-filter if necessary to suppress very low frequency drift. To prevent amplitude clipping, a safe normalization of the amplitude domain is then performed (without changing the relative dynamic range), and a peak margin matching the subsequent coding bit depth is maintained to obtain the corrected time-domain speech data.

[0152] S606: Perform sampling rate and bit depth output encoding processing on the corrected time-domain speech data to obtain enhanced speech waveform data.

[0153] The medical domain prioritizes outputting 16kHz / 16bit PCM (WAV container) for archiving and subsequent ASR; the financial domain can output 8kHz / 16bit PCM or 16kHz / 16bit PCM according to the quality inspection platform specifications, and write metadata such as call identifiers and timestamps when necessary. To avoid quantization noise concentration, low-amplitude unbiased jitter can be used before shaping to the integer bit depth to obtain enhanced speech waveform data.

[0154] It should be noted that because step S40 uses numerical integration of ordinary differential equations with a small step size to obtain the enhanced time-frequency representation during the inference stage, step S60 can complete the reconstruction output under low latency conditions. While keeping the ISTFT reconstruction link unchanged, if the upstream uses Mel spectrum as the enhancement target and enables a high-fidelity vocoder according to system specifications, steps S603-S605 can be replaced with a vocoder reconstruction sub-process in the engineering implementation.

[0155] As can be seen, this invention organically integrates generative modeling with flow matching modeling based on ordinary differential equations, constructing a consistent enhancement link of "standardization—velocity field estimation and numerical integration—destandardization—temporal reconstruction": replacing high-step diffusion inversion with velocity field estimation combined with numerical integration significantly reduces the number of inference steps and the real-time factor (RTF), achieving rapid inference; it does not rely on adversarial training and variational constraints, avoiding GAN training instability and VAE over-smoothing, thus improving the stability of training and inference; with the synergy of contextual modeling and optional semantic conditions, it can better maintain time-frequency details and phase consistency, improve speech intelligibility, naturalness, and speaker similarity in complex noise scenarios, and achieve better performance on objective evaluation metrics such as DNSMOS. Furthermore, the method is decoupled from the high-fidelity vocoder, facilitating engineering integration and stable deployment in edge and online systems.

[0156] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0157] In one embodiment, a stream-matching-based speech enhancement device is provided, which corresponds one-to-one with the stream-matching-based speech enhancement methods described in the above embodiments. For example... Figure 8 As shown, the stream matching-based speech enhancement device includes: a noise acquisition module 100, a noise processing module 200, a data processing module 300, a model optimization module 400, a data inversion module 500, and a reconstruction output module 600. Detailed descriptions of each functional module are as follows:

[0158] The noise acquisition module 100 is used to acquire noise speech data and perform time-frequency transformation processing to obtain noise time-frequency representation data.

[0159] The noise processing module 200 is used to standardize the noise time-frequency representation data to obtain standardized time-frequency representation data;

[0160] Data processing module 300 is used to perform stream matching input construction processing based on the standardized time-frequency representation data to obtain stream matching input representation data;

[0161] The model optimization module 400 is used to input the flow matching input representation data into the feature encoding submodule of the flow matching model to obtain flow matching feature representation data; perform context modeling processing on the flow matching feature representation data to obtain context-enhanced feature representation data; input the context-enhanced feature representation data into the velocity field estimation submodule to output velocity field estimation data; calculate the time-frequency state increment based on the velocity field estimation data and a preset numerical step size parameter to obtain time-frequency state increment data; use the time-frequency state increment data to numerically solve the ordinary differential equation for the intermediate state time-frequency tensor in the flow matching input representation data to obtain the updated intermediate state time-frequency tensor; and use the updated intermediate state time-frequency tensor as input to update the flow matching model until the target path parameter is reached to obtain enhanced time-frequency representation data.

[0162] The data back-inference module 500 is used to perform denormalization processing on the enhanced time-frequency representation data to obtain denormalized time-frequency representation data;

[0163] The reconstruction output module 600 is used to perform time-domain reconstruction processing on the denormalized time-frequency representation data to obtain enhanced speech waveform data.

[0164] In one embodiment, the noise acquisition module 100 is specifically used for:

[0165] Receive the audio input stream and parse the file header information to obtain the raw noisy speech data;

[0166] The original noisy speech data is subjected to sampling rate unification and channel conversion processing to obtain formatted noisy speech data;

[0167] The formatted noisy speech data is subjected to DC bias removal and pre-emphasis processing respectively to obtain preprocessed noisy speech data;

[0168] The preprocessed noisy speech data is segmented into frames based on a preset window length and frame shift to obtain a noisy speech frame sequence.

[0169] A preset window function is applied to each frame of the noisy speech frame sequence to obtain a windowed noisy speech frame sequence;

[0170] The windowed noisy speech frame sequence is subjected to a short-time Fourier transform to obtain a complex spectral matrix;

[0171] The amplitude spectrum and phase spectrum of the complex spectrum matrix are calculated and combined according to a preset time-frequency characteristic format to obtain noise time-frequency representation data.

[0172] In one embodiment, the noise processing module 200 is specifically used for:

[0173] The amplitude spectrum matrix and phase spectrum matrix are extracted from the noise time-frequency representation data respectively, and the amplitude spectrum matrix and phase spectrum matrix to be normalized are obtained accordingly.

[0174] The amplitude spectrum matrix to be standardized is numerically stabilized to obtain a stabilized amplitude spectrum matrix;

[0175] Based on the stabilized amplitude spectrum matrix, dynamic range compression processing is performed to obtain the compressed amplitude spectrum matrix;

[0176] Standardized statistical parameters are obtained by statistically analyzing the mean and standard deviation parameters of the compression amplitude spectrum matrix according to a preset dimension.

[0177] The compressed amplitude spectrum matrix is ​​normalized by using the standardized statistical parameters to obtain a normalized amplitude spectrum matrix.

[0178] The normalized amplitude spectrum matrix and the phase spectrum matrix are encoded and combined based on a preset time-frequency feature format to obtain standardized time-frequency representation data.

[0179] In one embodiment, the data processing module 300 is specifically used for:

[0180] The feature tensor corresponding to the model input size is extracted from the standardized time-frequency representation data to obtain the standardized time-frequency tensor;

[0181] Based on the standardized time-frequency tensor, a random noise tensor with zero mean and unit variance is generated to obtain the initial time-frequency noise tensor;

[0182] The path parameters are sampled within a preset interval, and the noise time-frequency initial tensor and the standardized time-frequency tensor are used to perform linear interpolation to obtain the intermediate state time-frequency tensor.

[0183] The intermediate state time-frequency tensor is subjected to frame block division and position index encoding to obtain an intermediate state time-frequency tensor with position encoding.

[0184] The intermediate state time-frequency tensor with position encoding and the path parameters are input together into the time encoding module, and a condition vector is output.

[0185] The conditional vector and the intermediate state time-frequency tensor with position encoding are concatenated by features and then processed by channel mapping to obtain the stream matching input representation data.

[0186] In one embodiment, the data processing module 300 is further specifically used for:

[0187] The start and end boundaries of the preset interval are parsed and a sampling strategy is configured to obtain the path parameters;

[0188] The relative position conversion and range constraint are applied to the target path parameters respectively to obtain the interpolation weights;

[0189] Based on the interpolation weights, the noise time-frequency initial tensor and the normalized time-frequency tensor are aligned in terms of size, channels, and step size to obtain an aligned tensor set.

[0190] The aligned tensor set is linearly weighted and synthesized to obtain the intermediate state time-frequency tensor.

[0191] In one embodiment, the data back-reasoning module 500 is specifically used for:

[0192] The enhanced normalized amplitude spectrum matrix and the enhanced phase spectrum matrix are extracted from the enhanced time-frequency representation data, respectively.

[0193] Perform an inverse transformation on the enhanced normalized amplitude spectrum matrix to obtain the inverse normalized amplitude spectrum matrix;

[0194] The inverse normalized amplitude spectrum matrix is ​​inversely compressed according to the preset dynamic range compression configuration parameters to obtain the inverse compressed amplitude spectrum matrix.

[0195] The inverse compression amplitude spectrum matrix is ​​numerically restored to obtain the restored amplitude spectrum matrix;

[0196] The restored amplitude spectrum matrix and the enhanced phase spectrum matrix are encoded and combined to obtain candidate inverse normalized time-frequency representation data;

[0197] The candidate denormalized time-frequency representation data is rearranged in dimensions to obtain the denormalized time-frequency representation data.

[0198] In one embodiment, the reconstruction output module 600 is specifically used for:

[0199] Extract the inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix from the inverse normalized time-frequency representation data, respectively;

[0200] The inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix are combined using complex number synthesis to obtain the reconstructed complex spectrum matrix.

[0201] The reconstructed complex spectrum matrix is ​​subjected to inverse short-time Fourier transform to obtain a windowed time-domain frame sequence;

[0202] The windowed time-domain frame sequence is subjected to dewindowing compensation and overlap addition processing respectively to obtain a continuous time-domain speech sequence;

[0203] DC bias correction is applied to the continuous time-domain speech sequence to obtain corrected time-domain speech data;

[0204] The corrected time-domain speech data is subjected to sampling rate and bit depth output encoding processing to obtain enhanced speech waveform data.

[0205] For specific limitations regarding the stream-matching-based speech enhancement device, please refer to the limitations of the stream-matching-based speech enhancement method above, which will not be repeated here. Each module in the aforementioned stream-matching-based speech enhancement device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0206] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a stream-matching-based speech enhancement method on the server side.

[0207] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 10As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a stream-matching-based speech enhancement method.

[0208] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed, can perform the steps provided in the above embodiments.

[0209] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0210] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0211] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0212] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A speech enhancement method based on stream matching, characterized in that, include: Noisy speech data is acquired and time-frequency transform is performed to obtain noise time-frequency representation data; The noise time-frequency representation data is standardized to obtain standardized time-frequency representation data. Based on the standardized time-frequency representation data, stream matching input construction processing is performed to obtain stream matching input representation data; The flow matching input representation data is input into the feature encoding submodule of the flow matching model to obtain flow matching feature representation data; the flow matching feature representation data is subjected to context modeling processing to obtain context-enhanced feature representation data; the context-enhanced feature representation data is input into the velocity field estimation submodule to output velocity field estimation data. The time-frequency state increment is calculated based on the velocity field estimation data and the preset numerical step size parameter to obtain the time-frequency state increment data; the time-frequency state increment data is used to numerically solve the ordinary differential equation of the intermediate state time-frequency tensor in the flow matching input representation data to obtain the updated intermediate state time-frequency tensor; the updated intermediate state time-frequency tensor is used as input to the flow matching model to update until the target path parameter is reached, thus obtaining the enhanced time-frequency representation data; The enhanced time-frequency representation data is denormalized to obtain denormalized time-frequency representation data; The denormalized time-frequency representation data is subjected to time-domain reconstruction processing to obtain enhanced speech waveform data; The standardization process for the noise time-frequency representation data to obtain standardized time-frequency representation data includes: extracting the amplitude spectrum matrix and phase spectrum matrix from the noise time-frequency representation data to obtain the amplitude spectrum matrix and phase spectrum matrix to be standardized; performing numerical stabilization on the amplitude spectrum matrix to be standardized to obtain a stabilized amplitude spectrum matrix; performing dynamic range compression on the stabilized amplitude spectrum matrix to obtain a compressed amplitude spectrum matrix; statistically analyzing the mean and standard deviation parameters of the compressed amplitude spectrum matrix according to a preset dimension to obtain standardized statistical parameters; using the standardized statistical parameters to normalize the mean and variance of the compressed amplitude spectrum matrix to obtain a normalized amplitude spectrum matrix; and encoding and combining the normalized amplitude spectrum matrix and the phase spectrum matrix according to a preset time-frequency feature format to obtain standardized time-frequency representation data. The process of constructing stream matching input based on the standardized time-frequency representation data to obtain stream matching input representation data includes: extracting feature tensors corresponding to the model input size from the standardized time-frequency representation data to obtain a standardized time-frequency tensor; generating a zero-mean, unit-variance random noise tensor based on the standardized time-frequency tensor to obtain an initial noise time-frequency tensor; sampling path parameters within a preset interval and performing linear interpolation between the initial noise time-frequency tensor and the standardized time-frequency tensor to obtain an intermediate state time-frequency tensor; performing frame block partitioning and position index encoding on the intermediate state time-frequency tensor to obtain an intermediate state time-frequency tensor with position encoding; inputting the intermediate state time-frequency tensor with position encoding and the path parameters into a time encoding module to output a condition vector; and concatenating the condition vector with the intermediate state time-frequency tensor with position encoding and performing channel mapping to obtain stream matching input representation data.

2. The speech enhancement method based on stream matching according to claim 1, characterized in that, The process of acquiring noisy speech data and performing time-frequency transformation to obtain noise time-frequency representation data includes: Receive the audio input stream and parse the file header information to obtain the raw noisy speech data; The original noisy speech data is subjected to sampling rate unification and channel conversion processing to obtain formatted noisy speech data; The formatted noisy speech data is subjected to DC bias removal and pre-emphasis processing respectively to obtain preprocessed noisy speech data; The preprocessed noisy speech data is segmented into frames based on a preset window length and frame shift to obtain a noisy speech frame sequence. A preset window function is applied to each frame of the noisy speech frame sequence to obtain a windowed noisy speech frame sequence; The windowed noisy speech frame sequence is subjected to a short-time Fourier transform to obtain a complex spectral matrix; The amplitude spectrum and phase spectrum of the complex spectrum matrix are calculated and combined according to a preset time-frequency characteristic format to obtain noise time-frequency representation data.

3. The speech enhancement method based on stream matching according to claim 1, characterized in that, The step of sampling path parameters within a preset interval and performing linear interpolation between the initial noise time-frequency tensor and the standardized time-frequency tensor to obtain an intermediate state time-frequency tensor includes: The start and end boundaries of the preset interval are parsed and a sampling strategy is configured to obtain the path parameters; The relative position conversion and range constraint are applied to the target path parameters respectively to obtain the interpolation weights; Based on the interpolation weights, the noise time-frequency initial tensor and the normalized time-frequency tensor are aligned in terms of size, channels, and step size to obtain an aligned tensor set. The aligned tensor set is linearly weighted and synthesized to obtain the intermediate state time-frequency tensor.

4. The speech enhancement method based on stream matching according to claim 1, characterized in that, The step of performing denormalization processing on the enhanced time-frequency representation data to obtain denormalized time-frequency representation data includes: The enhanced normalized amplitude spectrum matrix and the enhanced phase spectrum matrix are extracted from the enhanced time-frequency representation data, respectively. Perform an inverse transformation on the enhanced normalized amplitude spectrum matrix to obtain the inverse normalized amplitude spectrum matrix; The inverse normalized amplitude spectrum matrix is ​​inversely compressed according to the preset dynamic range compression configuration parameters to obtain the inverse compressed amplitude spectrum matrix. The inverse compression amplitude spectrum matrix is ​​numerically restored to obtain the restored amplitude spectrum matrix; The restored amplitude spectrum matrix and the enhanced phase spectrum matrix are encoded and combined to obtain candidate inverse normalized time-frequency representation data; The candidate denormalized time-frequency representation data is rearranged in dimensions to obtain the denormalized time-frequency representation data.

5. The speech enhancement method based on stream matching according to claim 1, characterized in that, The step of performing time-domain reconstruction processing on the denormalized time-frequency representation data to obtain enhanced speech waveform data includes: Extract the inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix from the inverse normalized time-frequency representation data, respectively; The inverse normalized amplitude spectrum matrix and the inverse normalized phase spectrum matrix are combined using complex number synthesis to obtain the reconstructed complex spectrum matrix. The reconstructed complex spectrum matrix is ​​subjected to inverse short-time Fourier transform to obtain a windowed time-domain frame sequence; The windowed time-domain frame sequence is subjected to dewindowing compensation and overlap addition processing respectively to obtain a continuous time-domain speech sequence; DC bias correction is applied to the continuous time-domain speech sequence to obtain corrected time-domain speech data; The corrected time-domain speech data is subjected to sampling rate and bit depth output encoding processing to obtain enhanced speech waveform data.

6. A speech enhancement device based on stream matching, characterized in that, include: The noise acquisition module is used to acquire noisy speech data and perform time-frequency transformation processing to obtain noise time-frequency representation data; The noise processing module is used to standardize the noise time-frequency representation data to obtain standardized time-frequency representation data. The data processing module is used to perform stream matching input construction processing based on the standardized time-frequency representation data to obtain stream matching input representation data. The model optimization module is used to input the flow matching input representation data into the feature encoding submodule of the flow matching model to obtain flow matching feature representation data; perform context modeling processing on the flow matching feature representation data to obtain context-enhanced feature representation data; and input the context-enhanced feature representation data into the velocity field estimation submodule to output velocity field estimation data. The time-frequency state increment is calculated based on the velocity field estimation data and the preset numerical step size parameter to obtain the time-frequency state increment data; the time-frequency state increment data is used to numerically solve the ordinary differential equation of the intermediate state time-frequency tensor in the flow matching input representation data to obtain the updated intermediate state time-frequency tensor; the updated intermediate state time-frequency tensor is used as input to the flow matching model to update until the target path parameter is reached, thus obtaining the enhanced time-frequency representation data; The data back-inference module is used to perform denormalization processing on the enhanced time-frequency representation data to obtain denormalized time-frequency representation data; The reconstruction output module is used to perform time-domain reconstruction processing on the denormalized time-frequency representation data to obtain enhanced speech waveform data. The noise processing module is specifically used to extract the amplitude spectrum matrix and the phase spectrum matrix from the noise time-frequency representation data, respectively, to obtain the amplitude spectrum matrix and the phase spectrum matrix to be standardized. The amplitude spectrum matrix to be standardized is numerically stabilized to obtain a stabilized amplitude spectrum matrix; Based on the stabilized amplitude spectrum matrix, dynamic range compression processing is performed to obtain the compressed amplitude spectrum matrix; The mean and standard deviation parameters of the compressed amplitude spectrum matrix are statistically analyzed according to a preset dimension to obtain standardized statistical parameters; the mean and variance of the compressed amplitude spectrum matrix are normalized using the standardized statistical parameters to obtain a normalized amplitude spectrum matrix; the normalized amplitude spectrum matrix and the phase spectrum matrix are encoded and combined based on a preset time-frequency feature format to obtain standardized time-frequency representation data. The data processing module is specifically used to extract feature tensors corresponding to the model input size from the standardized time-frequency representation data to obtain a standardized time-frequency tensor; generate a zero-mean, unit-variance random noise tensor based on the standardized time-frequency tensor to obtain an initial noise time-frequency tensor; sample path parameters within a preset interval and perform linear interpolation between the initial noise time-frequency tensor and the standardized time-frequency tensor to obtain an intermediate state time-frequency tensor; perform frame block partitioning and position index encoding on the intermediate state time-frequency tensor to obtain an intermediate state time-frequency tensor with position encoding; input the intermediate state time-frequency tensor with position encoding and the path parameters into the time encoding module to output a condition vector; and perform feature concatenation between the condition vector and the intermediate state time-frequency tensor with position encoding and perform channel mapping to obtain stream matching input representation data.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the stream-matching-based speech enhancement method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the stream-matching-based speech enhancement method as described in any one of claims 1 to 5.