Speech processing method, apparatus, device, storage medium, and program product

By constructing a phase mask to directly correct the phase angle in the audio domain features and using a multi-layer neural network model to process the initial phase spectrum, the problem of phase mismatch in frequency domain speech enhancement is solved, thus improving speech quality.

CN122135726APending Publication Date: 2026-06-02UNISOC CHONGQING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNISOC CHONGQING TECH CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing frequency domain speech enhancement techniques neglect the restoration of phase information, resulting in phase mismatch during time domain reconstruction of the enhanced speech, which affects speech quality.

Method used

The phase angle in the audio domain features is directly corrected by constructing a phase mask. The initial phase spectrum is processed by a multi-layer neural network model to generate the target phase spectrum. The target time-domain speech signal is then reconstructed by combining the initial amplitude spectrum.

Benefits of technology

It significantly improves the accuracy of phase perception in frequency domain speech processing, reduces phase mismatch during time domain reconstruction, and optimizes speech quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135726A_ABST
    Figure CN122135726A_ABST
Patent Text Reader

Abstract

This application provides a speech processing method, apparatus, device, storage medium, and program product. The method includes: acquiring an initial time-domain speech signal; performing a Fourier transform on the initial time-domain speech signal to obtain an initial amplitude spectrum and an initial phase spectrum; determining a phase mask based on the initial phase spectrum; processing the initial phase spectrum based on the phase mask to obtain a target phase spectrum; and determining a target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum. This method can improve speech quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a speech processing method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the popularization of technologies such as intelligent voice assistants, intelligent customer service, voice recognition systems (such as voice input methods and voice control of home appliances), and remote conferencing systems, users have placed higher demands on the clarity, accuracy, and naturalness of voice interaction.

[0003] In practical applications, speech signals are often contaminated by various factors, including environmental noise (such as traffic noise and background voices), reverberation (such as room echoes), and echo interference (such as feedback between speakers and microphones), leading to a significant decrease in intelligibility and recognition accuracy. Furthermore, in voice communication scenarios, network packet loss or signal attenuation can also cause intermittent or distorted speech. Therefore, using frequency domain speech enhancement techniques to reduce noise, suppress reverberation, cancel echoes, or repair signals from the original speech signal becomes a crucial step in improving speech quality.

[0004] Frequency domain processing has become a mainstream speech enhancement method due to its ability to transform complex time-domain signals into more easily analyzable spectral features. For example, in speech recognition systems, frequency domain enhancement can directly optimize the spectral structure of speech signals, thereby improving the input quality of recognition models; in voice communication, frequency domain processing can separate speech and noise components, achieving clearer call quality. However, related technologies generally neglect the direct restoration of phase information in frequency domain speech enhancement, leading to phase mismatch in the enhanced speech during time-domain reconstruction, which in turn causes speech distortion or auditory discomfort. Summary of the Invention

[0005] This application provides a speech processing method, apparatus, device, storage medium, and program product, which avoids phase mismatch during time-domain reconstruction of enhanced speech, thereby optimizing speech quality.

[0006] In a first aspect, embodiments of this application provide a voice processing method, including:

[0007] Acquire the initial time-domain speech signal;

[0008] Perform a Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum;

[0009] Determine the phase mask based on the initial phase spectrum;

[0010] The initial phase spectrum is processed based on the phase mask to obtain the target phase spectrum;

[0011] The target time-domain speech signal is determined based on the target phase spectrum and the initial amplitude spectrum.

[0012] In one possible implementation, determining the phase mask based on the initial phase spectrum includes:

[0013] The initial phase spectrum is input into the first neural network model to obtain the sinusoidal intermediate features;

[0014] A sinusoidal phase mask is obtained by performing a sinusoidal transform on the sinusoidal intermediate features.

[0015] The initial phase spectrum is input into the second neural network model to obtain the cosine intermediate features;

[0016] A cosine transform is performed on the cosine intermediate features to obtain a cosine phase mask.

[0017] In one possible implementation, the first neural network model and the second neural network model are multi-layer neural network models, including convolutional neural network models or fully connected neural network models.

[0018] In one possible implementation, the phase mask includes the sine phase mask and the cosine phase mask;

[0019] The step of processing the initial phase spectrum based on the phase mask to obtain the target phase spectrum includes:

[0020] The initial phase spectrum is subjected to a sine transform to obtain an initial sine phase component; and the initial phase spectrum is subjected to a cosine transform to obtain an initial cosine phase component.

[0021] The initial sinusoidal phase component and the initial cosine phase component are processed based on the sinusoidal phase mask and the cosine phase mask to obtain the target sinusoidal phase component and the target cosine phase component.

[0022] The target phase spectrum is obtained by performing an arctangent operation on the target sinusoidal phase component and the target cosine phase component.

[0023] In one possible implementation, determining the target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum includes:

[0024] Determine the amplitude mask based on the initial amplitude spectrum;

[0025] The initial amplitude spectrum is processed based on the amplitude mask to obtain the target amplitude spectrum;

[0026] The target phase spectrum and the target amplitude spectrum are subjected to inverse Fourier transform to obtain the target time-domain speech signal.

[0027] In one possible implementation, determining the amplitude mask based on the initial amplitude spectrum includes:

[0028] The initial amplitude spectrum is input into the third neural network model to obtain intermediate amplitude features;

[0029] A nonlinear transformation is performed on the intermediate amplitude feature to obtain the amplitude mask.

[0030] Secondly, embodiments of this application provide a voice processing apparatus, including:

[0031] The acquisition module is used to acquire the initial time-domain speech signal;

[0032] The Fourier transform module is used to perform Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum.

[0033] The first determining module is used to determine the phase mask based on the initial phase spectrum;

[0034] The processing module is used to process the initial phase spectrum based on the phase mask to obtain the target phase spectrum;

[0035] The second determining module is used to determine the target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum.

[0036] In one possible implementation, the first determining module is specifically used for:

[0037] The initial phase spectrum is input into the first neural network model to obtain the sinusoidal intermediate features;

[0038] A sinusoidal phase mask is obtained by performing a sinusoidal transform on the sinusoidal intermediate features.

[0039] The initial phase spectrum is input into the second neural network model to obtain the cosine intermediate features;

[0040] A cosine transform is performed on the cosine intermediate features to obtain a cosine phase mask.

[0041] In one possible implementation, the first neural network model and the second neural network model are multi-layer neural network models, including convolutional neural network models or fully connected neural network models.

[0042] In one possible implementation, the phase mask includes the sine phase mask and the cosine phase mask; the processing module is specifically used for:

[0043] The initial phase spectrum is subjected to a sine transform to obtain an initial sine phase component; and the initial phase spectrum is subjected to a cosine transform to obtain an initial cosine phase component.

[0044] The initial sinusoidal phase component and the initial cosine phase component are processed based on the sinusoidal phase mask and the cosine phase mask to obtain the target sinusoidal phase component and the target cosine phase component.

[0045] The target phase spectrum is obtained by performing an arctangent operation on the target sinusoidal phase component and the target cosine phase component.

[0046] In one possible implementation, the second determining module is specifically used for:

[0047] Determine the amplitude mask based on the initial amplitude spectrum;

[0048] The initial amplitude spectrum is processed based on the amplitude mask to obtain the target amplitude spectrum;

[0049] The target phase spectrum and the target amplitude spectrum are subjected to inverse Fourier transform to obtain the target time-domain speech signal.

[0050] In one possible implementation, the second determining module is specifically used for:

[0051] The initial amplitude spectrum is input into the third neural network model to obtain intermediate amplitude features;

[0052] A nonlinear transformation is performed on the intermediate amplitude feature to obtain the amplitude mask.

[0053] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0054] The memory stores computer-executed instructions;

[0055] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0056] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0057] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0058] The speech processing methods, apparatus, devices, storage media, and program products provided in this application embodiment directly correct the phase angle in the audio domain features by constructing a phase mask, significantly improving the accuracy of phase perception in frequency domain speech processing, thereby reducing phase mismatch and optimizing speech quality during time domain reconstruction. Attached Figure Description

[0059] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0060] Figure 1 Flowchart of the speech processing method provided in the embodiments of this application Figure 1 ;

[0061] Figure 2 Flowchart of the speech processing method provided in the embodiments of this application Figure 2 ;

[0062] Figure 3 Flowchart of the speech processing method provided in the embodiments of this application Figure 3 ;

[0063] Figure 4 This is a schematic diagram of the structure of the voice processing device provided in the embodiments of this application;

[0064] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0065] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0067] The terms "first," "second," etc., used in the embodiments of this application are for illustrative purposes and to distinguish the objects being described. They do not indicate any order or limit on the number of objects in the embodiments of this application, and cannot constitute any limitation on the embodiments of this application. For example, the use of terms such as "first neural network model" and "second neural network model" is only to distinguish different neural network models, and does not indicate any difference in the size, priority, or importance of the two neural network models.

[0068] It should be further understood that the terms "comprising" or "including" indicate the presence of the aforementioned features, steps, operations, elements, components, types, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, types, and / or groups.

[0069] In this application, terms such as "exemplary," "in some embodiments," and "in other embodiments" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the term "exemplary" is used to present the concept in a specific manner.

[0070] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include at least one sub-step or at least one stage. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0071] Speech enhancement refers to the process in digital signal processing where algorithms are used to process raw speech signals to improve their quality, clarity, or intelligibility. This technology primarily reduces background noise, echo, reverberation, packet loss, spectral spread, and other interference factors, resulting in clearer, more natural, and easier-to-understand speech. Speech enhancement technology is widely used in various scenarios, including but not limited to: telephone communication, speech recognition, voice assistants, audio recording, and broadcasting.

[0072] In related technologies, frequency domain speech signals can be enhanced in the following two ways.

[0073] I. The time-domain speech signal is converted into an amplitude spectrum using a Fast Fourier Transform (FFT), and the amplitude spectrum is repaired (e.g., noise suppression, gain adjustment). Then, the repaired amplitude spectrum is combined with the original phase using an Inverse Fast Fourier Transform (IFFT) to reconstruct the time-domain signal. Since the original phase may contain noise interference or reverberation components, directly retaining it will lead to phase distortion in the enhanced speech, especially in high-noise environments, where the improvement in speech quality is limited.

[0074] Second, the time-domain speech signal is converted into a complex spectrum using Fourier transform, the complex spectrum is repaired, and then the repaired complex spectrum is converted back into a time-domain signal. Although repairing the complex spectrum can indirectly enhance the amplitude and phase of the audio, it lacks direct perception of amplitude and phase. In addition, the complexity of complex number operations increases computational overhead and has low deployment efficiency on low-resource devices (such as embedded terminals).

[0075] Because the structured features of audio phase are weak, the phase exhibits a degree of randomness during audio propagation, making it difficult to directly generate or fit the phase. Therefore, this application provides a speech processing method that directly corrects the phase angle in the audio domain features by constructing a phase mask, significantly improving the accuracy of phase perception in frequency domain speech processing. This reduces phase mismatch during time domain reconstruction and optimizes speech quality.

[0076] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0077] Figure 1 Flowchart of the speech processing method provided in the embodiments of this application Figure 1 ,like Figure 1 As shown, the method includes the following steps:

[0078] S101. Obtain the initial time-domain speech signal.

[0079] The execution subject of this application embodiment can be an electronic device, chip, or chip module with voice processing function, or a voice processing device set in the above-mentioned electronic device, chip, or chip module. The voice processing device can be implemented by software or by a combination of software and hardware.

[0080] For example, the electronic device can be headphones, mobile phone, tablet, etc., and this application does not limit it.

[0081] For example, the initial time-domain speech signal can be a noisy speech signal.

[0082] The initial time-domain speech signal can be a continuous time-domain signal or a discrete sampled signal.

[0083] The initial time-domain speech signal can be obtained from the speech processing device itself or received from other devices; this application does not impose any restrictions on this.

[0084] S102. Perform Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum.

[0085] The Fourier transform may include the FFT, the Short-Time Fourier Transform (STFT), or the Discrete Fourier Transform (DFT), and this application does not limit it to any of these.

[0086] The initial amplitude spectrum and initial phase spectrum can be referred to as the frequency domain characteristics of a speech signal.

[0087] The initial time-domain speech signal, initial amplitude spectrum, and initial phase spectrum satisfy the following formula:

[0088] amp,pha = FFT(frame)

[0089] Where frame represents a frame of time-domain audio after the initial time-domain speech signal is framed, amp represents the initial amplitude spectrum, and pha represents the initial phase spectrum.

[0090] Specifically, after framing and windowing the initial time-domain speech signal, an FFT is performed on each frame of time-domain audio to obtain a complex spectrum. Each frequency point corresponds to a complex number, which includes a real part and an imaginary part. For each frequency point, a modulo operation is performed on the real and imaginary parts of each frequency point, i.e., The initial amplitude spectrum is obtained; the arctangent operation is performed on the real and imaginary parts corresponding to each frequency point, that is, the phase angle = atan2(imaginary part, real part), to obtain the initial phase spectrum.

[0091] The phase angles in the initial phase spectrum are within the range of [-π, π], and are bounded periodic phase angles.

[0092] S103. Determine the phase mask based on the initial phase spectrum.

[0093] A phase mask can refer to a phase correction value used to map the phase spectrum of noisy speech to the phase spectrum of clean speech.

[0094] The number of phase masks can be one or more.

[0095] In one possible implementation, the phase mask can be determined based on the initial phase spectrum in the following manner:

[0096] The initial phase spectrum is input into the first neural network model to obtain the sinusoidal intermediate feature; the sinusoidal intermediate feature is subjected to a sine transform to obtain the sinusoidal phase mask; the initial phase spectrum is input into the second neural network model to obtain the cosine intermediate feature; the cosine intermediate feature is subjected to a cosine transform to obtain the cosine phase mask.

[0097] The sine transform maps the intermediate features of the sine wave to the interval [-π, π], while the cosine transform maps the intermediate features of the cosine wave to the interval [-π, π].

[0098] The initial phase spectrum and the phase mask satisfy the following formula:

[0099] sin_mask = math.sin(model_sin(pha))

[0100] cos_mask = math.cos(model_cos(pha))

[0101] Where model_sin() represents the first neural network model, model_cos() represents the second neural network model, math.sin() represents the sine operation, math.cos() represents the cosine operation, sin_mask represents the sine phase mask, cos_mask represents the cosine phase mask, and pha represents the initial phase spectrum.

[0102] In one possible implementation, the first neural network model and the second neural network model are multi-layer neural network models, which include convolutional neural network models or fully connected neural network models.

[0103] The following details how to train the first and second neural network models.

[0104] 1. Training data preparation

[0105] 1.1 Constructing the Pairing Dataset

[0106] A large amount of paired speech data is identified, with each pair containing two versions of the same speech: one clean speech (the original recording without noise) and the other noisy speech (various noises added to the clean speech). These two speech pairs must be perfectly aligned in time, meaning that they correspond to the same speech content at the same moment.

[0107] 1.2 Extraction of Phase Spectrum

[0108] For each pair of speech data, a short-time Fourier transform is performed to convert the time-domain signal into a frequency-domain representation. The resulting complex spectrum is used to extract the phase spectrum. That is, each pair of speech data corresponds to two phase spectra: one is the phase spectrum of the noisy speech (used as input to the subsequent network), and the other is the phase spectrum of the corresponding clean speech (used to calculate the training target).

[0109] 1.3 Calculate the training objective

[0110] The phase spectrum of the clean speech is decomposed using sine and cosine decomposition, which involves calculating the sine and cosine values ​​of the phase value at each frequency point in each time frame. This yields two target matrices: a sine target matrix and a cosine target matrix. These two matrices will serve as the "standard answer" for training the neural network.

[0111] 1.4 Dataset Partitioning

[0112] All paired speech data is divided into three subsets: a training set, a validation set, and a test set, with no overlap between the three subsets. The training set is used for model parameter learning, the validation set is used for hyperparameter tuning and determining early stopping timing, and the test set is used for final model performance evaluation. For example, the training set comprises 80% of the total data, the validation set comprises 10%, and the test set comprises 10%.

[0113] 2. Model Input / Output Definition

[0114] 2.1 Input Data Format

[0115] The input to the first and second neural network models is the phase spectrum of the noisy speech. This phase spectrum is a three-dimensional data structure: the first dimension is the time frame index, the second dimension is the frequency point index, and the third dimension is the number of channels (here, 1). The value of each element represents the phase angle at that frequency point in that time frame, in radians, and ranges from -π to π.

[0116] 2.2 Output Data Definition

[0117] The outputs of both the first and second neural network models are three-dimensional data structures with the same dimensions as the inputs. However, the outputs of both models are not the final phase masks, but rather intermediate features. The intermediate features output by the first neural network model are transformed using a sine function to obtain a sine phase mask; the intermediate features output by the second neural network model are transformed using a cosine function to obtain a cosine phase mask.

[0118] 2.3 Definition of Training Objective

[0119] The training objective of the first neural network model is to make the output after sine function transformation (i.e., the sine phase mask) as close as possible to the sine phase component of the pure speech; the training objective of the second neural network model is to make the output after cosine function transformation (i.e., the cosine phase mask) as close as possible to the cosine phase component of the pure speech.

[0120] 3. Loss Function

[0121] Sine component loss: This term calculates the mean square error between the predicted sinusoidal phase mask obtained by sinusoidally transforming the intermediate features output by the first neural network model and the sinusoidal phase component of the clean speech. This loss term measures the accuracy of the network's prediction of the sinusoidal phase component.

[0122] Cosine component loss: This term calculates the mean square error between the predicted cosine phase mask (obtained by cosine transforming the intermediate features output by the second neural network model) and the clean speech cosine phase component. It measures the accuracy of the network's prediction of the cosine phase component.

[0123] Unit circle constraint loss: To ensure that the predicted sinusoidal and cosine phase masks conform to physical laws (i.e., lie on the unit circle), an auxiliary loss term is added: the square of the difference between the sum of the squares of the sine and cosine values ​​at each frequency point and 1, and then the average is calculated. This loss term ensures that the model output satisfies the physical constraint that the sum of the squares of the sine and cosine equals 1.

[0124] Total loss function: The three loss terms mentioned above are summed with certain weights to obtain the final total loss used for backpropagation. For example, the weights of the sine loss and cosine loss are each set to 1, and the weight of the unit circle constraint loss is set to 0.1 as an auxiliary term.

[0125] 4. Optimizer Configuration

[0126] 4.1 Optimizer Selection: The Adam optimizer is used to update the network parameters. This is an adaptive learning rate optimization algorithm that can dynamically adjust the learning rate of each parameter based on the first and second moments of the gradient, and it performs stably in speech processing tasks.

[0127] 4.2 Learning Rate Setting

[0128] The initial learning rate is typically set to 0.001. A learning rate that is too large may lead to unstable training, while a rate that is too small will result in slow convergence. Simultaneously, a weight decay coefficient of 0.00001 is set to act as L2 regularization, preventing the model from overfitting.

[0129] 4.3 Learning Rate Scheduling

[0130] As training progresses, if the validation set loss stops decreasing for several consecutive epochs, the learning rate needs to be reduced. Typically, this is set to multiply the learning rate by 0.5 every 5 epochs if the loss doesn't improve. Alternatively, a cosine annealing strategy can be used, allowing the learning rate to decay periodically according to a cosine function.

[0131] 4.4 Gradient clipping

[0132] To prevent gradient explosion, a gradient clipping threshold of 1.0 is set. That is, after backpropagation, if the norm of the gradient exceeds 1.0, the gradient is scaled down to a norm of 1.0.

[0133] 5. Training Cycle Execution Phase

[0134] 5.1 Batch Training

[0135] The training data is divided into multiple batches according to a set batch size (usually 32). Each batch contains the phase spectra of multiple speech segments and their corresponding training targets. The data is then fed into the model batch by batch for training.

[0136] 5.2 Forward Propagation

[0137] For each batch, the phase spectrum of the noisy speech is simultaneously input into the first neural network model and the second neural network model. The first neural network model outputs sinusoidal intermediate features, and the second neural network model outputs cosine intermediate features. Then, sine and cosine function transforms are applied to these two intermediate features respectively to obtain the predicted sinusoidal phase mask and cosine phase mask.

[0138] 5.3 Loss Calculation

[0139] Based on the sinusoidal and cosine phase masks predicted for the current batch, and the sinusoidal and cosine phase components of the clean speech corresponding to that batch, the total loss value is calculated according to the designed loss function.

[0140] 5.4 Backpropagation

[0141] The loss value is backpropagated to calculate the gradient of each network parameter. This process relies on the automatic differentiation mechanism of deep learning frameworks such as PyTorch or TensorFlow.

[0142] 5.5 Parameter Update

[0143] The optimizer updates the network parameters based on the calculated gradients. The Adam optimizer combines the current gradient with historical gradient information to calculate the direction and magnitude of adjustment for each parameter.

[0144] 5.6 Gradient clipping

[0145] Before updating the parameters, the gradient is clipped to ensure that the gradient norm does not exceed a preset threshold, thus preventing training instability caused by gradient explosion.

[0146] 5.7 Iterative Repetition

[0147] Repeat the above steps to iterate through all batches in the training set, completing one training epoch. Then, shuffle the data order and begin the next training epoch.

[0148] 6. Verification and Monitoring Phase

[0149] 6.1 Periodic verification

[0150] After each training epoch, the model performance is evaluated on the validation set. At this point, the network is set to evaluation mode, and gradient calculation is not performed; only forward propagation and loss calculation are performed. The validation loss is used to monitor whether the model is overfitting.

[0151] 6.2 Multi-indicator monitoring

[0152] In addition to the total loss value, several refined metrics need to be monitored: the mean square error of the sinusoidal phase component, the mean square error of the cosine phase component, the unit circle constraint error, and the phase reconstruction error calculated by combining the sinusoidal and cosine phase masks. These metrics reflect the model's learning performance from different perspectives.

[0153] 6.3 Early Termination Judgment

[0154] If the validation set loss stops decreasing for several consecutive rounds (e.g., 10 rounds) or even starts to increase, it indicates that the model is overfitting. In this case, training should be stopped early, and the model parameters with the lowest validation set loss should be saved.

[0155] 6.4 Model Saving

[0156] After each round, if the current validation set loss is lower than the historical best value, the current model parameters are saved. This way, the model obtained after training is the one that performed best throughout the entire training process.

[0157] 7. Post-training processing

[0158] After training, the model is finally evaluated on a test set that has never been used before. The metrics calculated at this point reflect the model's expected performance in real-world application scenarios.

[0159] Save the trained model parameters, optimizer status, training history, etc., for easy deployment or continued training.

[0160] S104. Process the initial phase spectrum based on the phase mask to obtain the target phase spectrum.

[0161] S105. Based on the target phase spectrum and the initial amplitude spectrum, determine the target time-domain speech signal.

[0162] The target time-domain speech signal can be obtained by performing an inverse Fourier transform on the target phase spectrum and the initial amplitude spectrum. Alternatively, the initial amplitude spectrum can be processed first, and then an inverse Fourier transform can be performed on the processed amplitude spectrum and the target phase spectrum to obtain the target time-domain speech signal.

[0163] The inverse Fourier transform can include IFFT, inverse short-time Fourier transform (ISTFT), or inverse discrete Fourier transform (IDFT), and this application does not limit it.

[0164] exist Figure 1 In the illustrated embodiment, the phase angle in the audio domain features is directly corrected by constructing a phase mask, which significantly improves the accuracy of phase perception in frequency domain speech processing, thereby reducing phase mismatch and optimizing speech quality during time domain reconstruction.

[0165] exist Figure 1 Based on the illustrated embodiment, the following is combined with Figure 2 and Figure 3 The technical solution of this application is described in detail.

[0166] Figure 2 Flowchart of the speech processing method provided in the embodiments of this application Figure 2 ,like Figure 2 As shown, the method includes the following steps:

[0167] S201. Obtain the initial time-domain speech signal.

[0168] S202. Perform Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum.

[0169] S203. Input the initial phase spectrum into the first neural network model to obtain the sinusoidal intermediate features; perform a sinusoidal transformation on the sinusoidal intermediate features to obtain the sinusoidal phase mask.

[0170] S204. Input the initial phase spectrum into the second neural network model to obtain the cosine intermediate features; perform a cosine transform on the cosine intermediate features to obtain the cosine phase mask.

[0171] It should be noted that the execution process of S201 to S204 can be referred to the execution process of S101 to S103, and will not be repeated here.

[0172] S205. Perform a sine transform on the initial phase spectrum to obtain the initial sine phase component; and perform a cosine transform on the initial phase spectrum to obtain the initial cosine phase component.

[0173] For example, the initial phase spectrum can be split into initial sinusoidal phase components using sine operations. The initial phase spectrum can also be split into initial cosine phase components using cosine operations.

[0174] The initial phase spectrum, initial sinusoidal phase component, and initial cosine phase component satisfy the following formula:

[0175] sin_pha, cos_pha =math.sin(pha), math.cos(pha)

[0176] Where pha represents the initial phase spectrum, sin_pha represents the initial sine phase component, cos_pha represents the initial cosine phase component, math.sin() represents the sine operation, and math.cos() represents the cosine operation.

[0177] It should be noted that this application does not limit the execution order of S203, S204 and S205. For example, S203, S204 and S205 can be executed in parallel.

[0178] S206. Based on the sinusoidal phase mask and the cosine phase mask, process the initial sinusoidal phase component and the initial cosine phase component to obtain the target sinusoidal phase component and the target cosine phase component.

[0179] The target sinusoidal phase component can also be called the repaired sinusoidal phase component, and the target cosine phase component can also be called the repaired cosine phase component. This application does not limit this.

[0180] For example, the target sinusoidal phase component, the target cosine phase component, the sinusoidal phase mask, the cosine phase mask, the initial sinusoidal phase component, and the initial cosine phase component satisfy the following formula:

[0181] sin_pha_E = sin_pha × cos_mask + cos_pha × sin_mask

[0182] cos_pha_E = cos_pha × cos_mask - sin_pha × sin_mask

[0183] Wherein, sin_pha_E represents the target sinusoidal phase component, cos_pha_E represents the target cosine phase component, sin_mask represents the sinusoidal phase mask, cos_mask represents the cosine phase mask, sin_pha represents the initial sinusoidal phase component, and cos_pha represents the initial cosine phase component.

[0184] S207. Perform arctangent operation on the target sinusoidal phase component and the target cosine phase component to obtain the target phase spectrum.

[0185] The target phase spectrum can also be referred to as the repaired phase spectrum, and this application does not limit it.

[0186] For example, the target sinusoidal phase component, the target cosine phase component, and the target phase spectrum satisfy the following formula:

[0187] pha_E = math.atan2(sin_pha_E, cos_pha_E)

[0188] Where pha_E represents the target phase spectrum, sin_pha_E represents the target sinusoidal phase component, cos_pha_E represents the target cosine phase component, and math.atan2() represents the arctangent operation.

[0189] S208. Determine the target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum.

[0190] It should be noted that the execution process of S208 can be referred to the execution process of S105, and will not be repeated here.

[0191] exist Figure 2 In the illustrated embodiment, the phase angle in the audio domain features is directly corrected by constructing a sine phase mask and a cosine phase mask, which significantly improves the accuracy of phase perception in frequency domain speech processing, thereby reducing phase mismatch and optimizing speech quality during time domain reconstruction.

[0192] Figure 3 Flowchart of the speech processing method provided in the embodiments of this application Figure 3 ,like Figure 3 As shown, the method includes the following steps:

[0193] S301. Acquire the initial time-domain speech signal.

[0194] S302. Perform Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum.

[0195] S303. Input the initial phase spectrum into the first neural network model to obtain the sinusoidal intermediate feature; perform a sinusoidal transformation on the sinusoidal intermediate feature to obtain the sinusoidal phase mask.

[0196] S304. Input the initial phase spectrum into the second neural network model to obtain the cosine intermediate features; perform a cosine transform on the cosine intermediate features to obtain the cosine phase mask.

[0197] S305. Perform a sine transform on the initial phase spectrum to obtain the initial sine phase component; and perform a cosine transform on the initial phase spectrum to obtain the initial cosine phase component.

[0198] S306. The initial sinusoidal phase component and the initial cosine phase component are processed based on the sinusoidal phase mask and the cosine phase mask to obtain the target sinusoidal phase component and the target cosine phase component.

[0199] S307. Perform arctangent operation on the target sinusoidal phase component and the target cosine phase component to obtain the target phase spectrum.

[0200] It should be noted that the execution process of S301 to S307 can be referred to the execution process of S201 to S207, and will not be repeated here.

[0201] S308. Determine the amplitude mask based on the initial amplitude spectrum.

[0202] Amplitude mask refers to amplitude correction, used to map the amplitude spectrum of noisy speech to the amplitude spectrum of clean speech.

[0203] In one possible implementation, the amplitude mask can be determined based on the initial amplitude spectrum in the following manner:

[0204] The initial amplitude spectrum is input into the third neural network model to obtain the intermediate amplitude features; the intermediate amplitude features are then subjected to a nonlinear transformation to obtain the amplitude mask.

[0205] The third neural network model can be a multi-layer neural network model, which includes convolutional neural network models or fully connected neural network models.

[0206] The training process of the third neural network model can refer to the training process of the first or second neural network model, and will not be repeated here.

[0207] Nonlinear transformation is used to map the intermediate features of the amplitude to the interval [0,1].

[0208] For example, nonlinear transformation can refer to processing intermediate features of amplitude using activation functions such as exponential functions, sigmoid functions, or ReLU functions.

[0209] For example, the amplitude mask and the initial amplitude spectrum satisfy the following formula:

[0210] amp_mask = math.sigmoid(model_amp(amp))

[0211] Where amp_mask represents the amplitude mask, amp represents the initial amplitude spectrum, model_amp() represents the third neural network model, and math.sigmoid() represents the nonlinear transformation.

[0212] S309. Process the initial amplitude spectrum based on the amplitude mask to obtain the target amplitude spectrum.

[0213] The target amplitude spectrum can also be referred to as the repaired amplitude spectrum, and this application does not limit it.

[0214] For example, the amplitude mask can be multiplied element-wise with the initial amplitude spectrum to obtain the target amplitude spectrum.

[0215] It should be noted that the amplitude spectrum repair process can also be performed before the phase spectrum repair process, that is, S308 and S309 can also be performed before S303; or, the amplitude spectrum repair process and the phase spectrum repair process can be performed in parallel, that is, S308 to S309 can be performed in parallel with S303 to S307.

[0216] S310. Perform inverse Fourier transform on the target phase spectrum and the target amplitude spectrum to obtain the target time-domain speech signal.

[0217] exist Figure 3 In the illustrated embodiment, the phase angle in the audio domain features is directly corrected by constructing sine and cosine phase masks, significantly improving the accuracy of phase perception in frequency domain speech processing. This reduces phase mismatch during time domain reconstruction and optimizes speech quality. Simultaneously, the amplitude in the audio domain features is directly corrected by constructing amplitude masks, significantly improving the accuracy of amplitude perception in frequency domain speech processing and further optimizing speech quality.

[0218] Figure 4 This is a schematic diagram of the structure of the voice processing device provided in an embodiment of this application. Figure 4 As shown, the device 10 includes an acquisition module 11, a Fourier transform module 12, a determination module 13, a processing module 14, and an inverse Fourier transform module 15.

[0219] Acquisition module 11 is used to acquire the initial time-domain speech signal;

[0220] Fourier transform module 12 is used to perform Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum;

[0221] The first determining module 13 is used to determine the phase mask based on the initial phase spectrum;

[0222] Processing module 14 is used to process the initial phase spectrum based on the phase mask to obtain the target phase spectrum;

[0223] The second determining module 15 is used to determine the target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum.

[0224] In one possible implementation, the first determining module 13 is specifically used for:

[0225] The initial phase spectrum is input into the first neural network model to obtain the sinusoidal intermediate features;

[0226] A sinusoidal phase mask is obtained by performing a sinusoidal transform on the intermediate features of the sinusoidal shape.

[0227] The initial phase spectrum is input into the second neural network model to obtain the cosine intermediate features;

[0228] A cosine phase mask is obtained by performing a cosine transform on the intermediate features of the cosine matrix.

[0229] In one possible implementation, the first neural network model and the second neural network model are multi-layer neural network models, which include convolutional neural network models or fully connected neural network models.

[0230] In one possible implementation, the phase mask includes a sine phase mask and a cosine phase mask; the processing module 14 is specifically used for:

[0231] The initial phase spectrum is subjected to a sine transform to obtain the initial sinusoidal phase component; and the initial phase spectrum is subjected to a cosine transform to obtain the initial cosine phase component.

[0232] The initial sinusoidal phase component and the initial cosine phase component are processed based on the sinusoidal phase mask and the cosine phase mask to obtain the target sinusoidal phase component and the target cosine phase component.

[0233] The target phase spectrum is obtained by performing arctangent operation on the target sinusoidal phase component and the target cosine phase component.

[0234] In one possible implementation, the second determining module 15 is specifically used for:

[0235] Determine the amplitude mask based on the initial amplitude spectrum;

[0236] The target amplitude spectrum is obtained by processing the initial amplitude spectrum based on the amplitude mask;

[0237] The target phase spectrum and target amplitude spectrum are subjected to inverse Fourier transform to obtain the target time-domain speech signal.

[0238] In one possible implementation, the second determining module 15 is specifically used for:

[0239] The initial amplitude spectrum is input into the third neural network model to obtain intermediate amplitude features;

[0240] A nonlinear transformation is performed on the intermediate features of the amplitude to obtain the amplitude mask.

[0241] The voice processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0242] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 20 includes a transceiver 21, a memory 22, and a processor 23. The transceiver 21 may include a transmitter and / or a receiver. The transmitter may also be referred to as a transmitter, transmitter port, or transmitter interface, etc., and the receiver may also be referred to as a receiver, receiver port, or receiver interface, etc. Exemplarily, the transceiver 21, memory 22, and processor 23 are interconnected via a bus 24.

[0243] Memory 22 is used to store program instructions;

[0244] The processor 23 is used to execute the program instructions stored in the memory to cause the electronic device to perform any of the voice processing methods shown above.

[0245] Transceiver 21 is used to perform the sending and receiving functions of electronic devices.

[0246] In one possible implementation, the memory 22 may be the storage medium described above.

[0247] Electronic devices can include chips, modules, integrated development environments (IDEs), etc.

[0248] Figure 5 The electronic device shown in the embodiments can execute the technical solutions shown in the above method embodiments. Its implementation principle and beneficial effects are similar, and will not be repeated here.

[0249] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement any of the above-mentioned speech processing methods.

[0250] This application provides a computer program product, including a computer program that, when executed by a processor, can implement any of the above-described voice processing methods.

[0251] This application provides a chip on which a computer program is stored. When the computer program is executed by the chip, it implements the above-mentioned voice processing method.

[0252] In one possible implementation, the chip is a chip in a chip module.

[0253] The computer-readable storage medium and computer program product of the present application embodiments can execute the technical solutions shown in the above-described speech processing method embodiments. Their implementation principles and beneficial effects are similar and will not be described again here.

[0254] All or part of the steps in the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above-described method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), random access memory (RAM), flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.

[0255] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0256] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0257] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0258] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A speech processing method, characterized in that, include: Acquire the initial time-domain speech signal; Perform a Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum; Determine the phase mask based on the initial phase spectrum; The initial phase spectrum is processed based on the phase mask to obtain the target phase spectrum; The target time-domain speech signal is determined based on the target phase spectrum and the initial amplitude spectrum.

2. The method according to claim 1, characterized in that, Determining the phase mask based on the initial phase spectrum includes: The initial phase spectrum is input into the first neural network model to obtain the sinusoidal intermediate features; A sinusoidal phase mask is obtained by performing a sinusoidal transform on the sinusoidal intermediate features. The initial phase spectrum is input into the second neural network model to obtain the cosine intermediate features; A cosine transform is performed on the cosine intermediate features to obtain a cosine phase mask.

3. The method according to claim 2, characterized in that, The first neural network model and the second neural network model are multi-layer neural network models, which include convolutional neural network models or fully connected neural network models.

4. The method according to claim 2 or 3, characterized in that, The phase mask includes the sine phase mask and the cosine phase mask; The step of processing the initial phase spectrum based on the phase mask to obtain the target phase spectrum includes: The initial phase spectrum is subjected to a sine transform to obtain an initial sine phase component; and the initial phase spectrum is subjected to a cosine transform to obtain an initial cosine phase component. The initial sinusoidal phase component and the initial cosine phase component are processed based on the sinusoidal phase mask and the cosine phase mask to obtain the target sinusoidal phase component and the target cosine phase component. The target phase spectrum is obtained by performing an arctangent operation on the target sinusoidal phase component and the target cosine phase component.

5. The method according to any one of claims 1-4, characterized in that, The step of determining the target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum includes: Determine the amplitude mask based on the initial amplitude spectrum; The initial amplitude spectrum is processed based on the amplitude mask to obtain the target amplitude spectrum; The target phase spectrum and the target amplitude spectrum are subjected to inverse Fourier transform to obtain the target time-domain speech signal.

6. The method according to claim 5, characterized in that, Determining the amplitude mask based on the initial amplitude spectrum includes: The initial amplitude spectrum is input into the third neural network model to obtain intermediate amplitude features; A nonlinear transformation is performed on the intermediate amplitude feature to obtain the amplitude mask.

7. A voice processing device, characterized in that, The device includes: The acquisition module is used to acquire the initial time-domain speech signal; The Fourier transform module is used to perform Fourier transform on the initial time-domain speech signal to obtain the initial amplitude spectrum and the initial phase spectrum. The first determining module is used to determine the phase mask based on the initial phase spectrum; The processing module is used to process the initial phase spectrum based on the phase mask to obtain the target phase spectrum; The second determining module is used to determine the target time-domain speech signal based on the target phase spectrum and the initial amplitude spectrum.

8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.