Speech packet loss compensation method and device based on improved discrete cosine transform domain generative adversarial model

By improving the discrete cosine transform domain generative adversarial model and the dual-path recurrent convolution generator network, and combining a multi-scale discriminator and a dynamic loss function, the problem of poor performance in deep learning speech packet loss compensation is solved, and higher accuracy and naturalness of speech compensation are achieved.

CN120783770BActive Publication Date: 2025-11-11SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511286311.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-11
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing deep learning-based speech packet loss compensation schemes have room for improvement, as traditional methods struggle to effectively enhance the accuracy and naturalness of speech packet loss compensation.

Method used

An improved discrete cosine transform domain generative adversarial model is adopted. Speech signal features are extracted by improving the discrete cosine transform. A generative speech compensation network is constructed by combining a dual-path recurrent convolutional generator network and a multi-scale multi-period discriminator. The network is then optimized using a dynamic weighted combination loss function.

Benefits of technology

It significantly improves the accuracy and naturalness of voice packet loss compensation, enhances voice quality assessment indicators, and achieves better voice packet loss compensation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783770B_ABST
    Figure CN120783770B_ABST
Patent Text Reader

Abstract

This invention discloses a speech packet loss compensation method and device based on an improved discrete cosine transform domain generative adversarial model (GAN). The method includes: forming a training set using complete speech signals and corresponding lost speech signals; constructing a GAN based on the improved discrete cosine transform domain, specifically including: an improved discrete cosine transform domain feature extraction module, a generator network based on dual-path recurrent convolution, an improved inverse discrete cosine transform module, a discriminator network, and a loss function calculation module. The generator network specifically includes an encoder module, a dual-path long short-term memory (LSTM) network module, and a decoder module connected sequentially. The encoder module includes several encoder layers, and the decoder module includes several decoder layers. The encoder and decoder layers are connected using a SimAM attention module. The GAN is trained using the training set, and then packet loss compensation is performed based on the trained network. This invention provides better compensation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to speech processing technology, and more particularly to a speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model. Background Technology

[0002] In real-time voice communication, packet loss is a core challenge due to the unreliability of network transmission. Early voice packet loss compensation mainly relied on signal processing techniques, particularly limited compensation based on signal analysis. Traditional packet loss compensation (PLC) techniques can be divided into two categories: transmitter-side compensation and receiver-side compensation. Transmitter-side solutions primarily utilize forward error correction and interleaving techniques. Forward error correction achieves packet loss recovery by adding redundant information; interleaving reorders voice units for transmission, dispersing packet loss to reduce its perceptual impact. Receiver-side compensation focuses on error concealment algorithms, mainly including silence filling and waveform repetition algorithms.

[0003] The introduction of deep learning technology has broken through the bottlenecks of traditional algorithms. Its core lies in modeling the long-term dependencies and contextual features of speech through data-driven approaches. Early research mainly focused on recurrent neural network (RNN) models, predicting the content of lost frames by modeling the temporal correlation of speech signals. The introduction of generative adversarial networks (GANs) has further promoted the refinement of speech generation. By training a generator network to synthesize the time spectrum or waveform of lost frames, a discriminator network evaluates the coherence and naturalness of the generated speech. Through the introduction of deep learning schemes, speech packet loss compensation has evolved from the traditional passive filling paradigm to an active generation paradigm, showing broad development prospects. However, the effectiveness of current deep learning-based speech packet loss compensation schemes needs further improvement. Summary of the Invention

[0004] To address the problems existing in the prior art, the purpose of this invention is to provide a more effective speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model.

[0005] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0006] A speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model, characterized by comprising:

[0007] Step 1: Process the complete speech signal to obtain the lost speech signal, and combine several complete speech signals and corresponding lost speech signals to form sample pairs to form a training dataset;

[0008] Step 2: Construct an improved Generative Adversarial Network (GAN) based on the Discrete Cosine Transform domain, specifically including:

[0009] An improved discrete cosine transform domain feature extraction module is used to extract the time-frequency domain features of the lost speech signal by performing an improved discrete cosine transform on the lost speech signal.

[0010] A generator network based on dual-path recurrent convolution is used to compensate for lost speech signals by utilizing the time-frequency domain features of lost speech signals to generate pseudo-complete speech signals. The generator network specifically includes an encoder module, a dual-path long short-term memory network (LSTM) module, and a decoder module connected in sequence. The encoder module includes several encoder layers, and the decoder module includes several decoder layers. The encoder and decoder layers are connected by a SimAM attention module.

[0011] An improved inverse discrete cosine transform module is used to perform an improved inverse discrete cosine transform on the features of pseudo-complete speech signals to obtain pseudo-complete speech signals.

[0012] A discriminator network is used to distinguish between genuine and pseudo-genuine speech signals;

[0013] The loss function calculation module is used to combine the generator network loss, the temporal MAE loss, the improved discrete cosine transform loss, and the speech quality perception evaluation loss to calculate the total loss, thereby updating the parameters of the generator network.

[0014] Step 3: Input the training dataset into the improved discrete cosine transform domain generative adversarial network to train the network;

[0015] Step 4: Input the lost speech signal to be tested into the trained network to obtain the compensated complete speech signal.

[0016] Furthermore, the encoder module includes a cascaded first encoder layer and a second encoder layer. The first encoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a one-dimensional pixel inverse rearrangement layer, a PReLU activation function layer, and a residual connection operation connected in sequence. The second encoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a PReLU activation function layer, a dual-path feedforward sequence memory neural network layer, a Snake activation function layer, and a residual connection operation connected in sequence.

[0017] Furthermore, the decoder module includes a first decoder layer and a second decoder layer. The first decoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a PReLU activation function layer, a dual-path feedforward sequence memory neural network layer, a Snake activation function layer, a SimAM attention layer, and a residual connection operation connected in sequence. The second decoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a one-dimensional pixel rearrangement layer, and a PReLU activation function layer connected in sequence.

[0018] Furthermore, the dual-path long short-term memory network LSTM module specifically includes several intra-block LSTM modules and inter-block LSTM modules connected in sequence. The intra-block LSTM module includes a bidirectional long short-term memory network Bi-LSTM, a linear layer, a normalization layer, and a residual connection operation connected in sequence. The inter-block LSTM module includes a unidirectional LSTM and a linear layer connected in sequence.

[0019] Furthermore, the SimAM attention module is used to perform the following calculation process:

[0020] ,

[0021] In the formula, X is the input to the SimAM attention module. For output, This indicates point-by-point multiplication. for function, The weight matrix is ​​composed of the minimum energy value. The minimum energy value is obtained by taking the reciprocal of each point. for:

[0022] ,

[0023] in, express The t-th feature point, This represents the global mean of the channel containing the feature point. This represents the global variance of the channel containing the feature point. This is a hyperparameter.

[0024] Furthermore, the discriminator network includes a multi-scale discriminator network and a multi-period discriminator network. For the multi-scale discriminator, the input speech data is downsampled at different scales to generate multiple low-resolution scale signals. Each scale signal is input into an independent sub-discriminator network. These sub-discriminator networks are composed of stacked depth-separable convolutional layers. The higher the resolution of the sub-discriminator network, the smaller the convolutional kernel and the deeper the network. For the multi-period discriminator, the input speech data is divided into multiple sub-sequences according to different period lengths. Each period sub-sequence group is input into an independent sub-discriminator network. These sub-discriminators use asymmetric period-size convolutional kernels to perform temporal convolution operations.

[0025] Furthermore, the function for calculating the total loss is:

[0026] ,

[0027] In the formula, For the total loss, To represent the time-domain weighted MAE loss The weighting coefficients, Indicates the loss in perceived speech quality. The weighting coefficients, The amplitude MAE loss of the improved discrete cosine transform spectrum. The weighting coefficients, Represents the polarity MAE loss of the improved discrete cosine transform spectrum. The weighting coefficients, Indicate generator loss The weighting coefficients.

[0028] Furthermore, the generator loss is calculated using the following function:

[0029] ,

[0030] In the formula, For the discriminator network to process the pseudo-complete speech signal generated by the generator. The output, Represents the generator network function. It is the lost audio signal input to the generator network. Represents pseudo-complete speech signal The distribution This indicates that the mean value is being calculated.

[0031] Furthermore, the improved amplitude MAE loss of the discrete cosine transform spectrum... and the polarity MAE loss of the improved discrete cosine transform spectrum The calculation formula is as follows:

[0032] ,

[0033] in, This represents the improved discrete cosine transform spectrum corresponding to the complete speech signal. This represents the improved discrete cosine transform spectrum of the speech signal after compensation by a generator network based on dual-path circular convolution.

[0034] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described above.

[0035] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention transforms the speech signal from the time domain to the improved discrete cosine transform domain by introducing an improved discrete cosine transform. Compared with the traditional short-time Fourier transform, the discrete cosine transform has advantages in multiple dimensions, such as avoiding phase modeling, compressing feature dimensions, and being compatible with coding standards. Simultaneously, a dual-path CRN network is introduced as the generator for speech packet loss compensation, constructing a generative speech compensation network structure that takes into account both local details and long-range time-frequency dependencies. Then, a lightweight multi-scale and multi-period discriminator is integrated, and the naturalness of the compensated speech is improved through adversarial training. The compensation accuracy is enhanced by combining a dynamically weighted combined loss function with joint optimization of time-domain MAE, PMSQE, and improved discrete cosine transform spectral constraints. This invention achieves significant improvements in all evaluation indicators of speech packet loss compensation, resulting in better speech packet loss compensation performance. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the speech packet loss compensation method based on the improved discrete cosine transform domain generative adversarial model provided by the present invention;

[0037] Figure 2 This is a schematic diagram of the generator network structure based on dual-path CRN provided by the present invention;

[0038] Figure 3 This is a schematic diagram of the dual-path LSTM module structure provided by the present invention;

[0039] Figure 4 This is a schematic diagram of the dual-path LSTM module provided by the present invention;

[0040] Figure 5 This is a schematic diagram of the principle of the multi-scale, multi-period discriminator provided by the present invention. Detailed Implementation

[0041] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0042] Example 1

[0043] This invention provides a speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model, such as... Figure 1 As shown, it includes:

[0044] Step 1: Process the complete speech signal to obtain the lost speech signal, and combine several complete speech signals and corresponding lost speech signals to form sample pairs to form a training dataset.

[0045] In the specific implementation, the dataset from the INTERSPEECH 2022 Audio Deep Packet LossConcealment Challenge was selected as the first type of dataset, specifically designed for studying the packet loss compensation algorithm PLC problem in voice communication. The dataset consists of three parts: a training set, a validation set, and a blind test set. Each sample in the dataset consists of three parts: a complete audio file: the original speech segment unaffected by packet loss; a packet-loss audio file: audio generated from a real packet-loss audio track, with the packet loss portion set to zero; and a packet-loss audio track file: the packet loss status is marked in 20-millisecond time units ("1" indicates loss, "0" indicates normal), providing the model with accurate contextual information.

[0046] The second dataset models each packet loss event as an independent event and strictly limits the maximum continuous packet loss duration to 160ms. On the complete audio files of the first dataset, random frame masking is performed using a Markov model to simulate random packet loss (10%-50%) and Wi-Fi packet loss patterns, creating artificially synthesized custom-designed lost-packet audio. This scheme systematically constructs a training dataset with short-term burst packet loss as its core feature through the artificial degradation processing of the original clean audio.

[0047] All audio in the dataset is single-channel audio with a sampling rate of 16000Hz. The training set contains 23,184 audio tracks, each approximately 10 seconds long, totaling about 64 hours. The blind test set contains 966 real-world call audio tracks to test the model's performance under various packet loss conditions.

[0048] The first and second datasets are combined to form the training dataset for model training.

[0049] Step 2: Construct a generative adversarial network based on the improved discrete cosine transform domain.

[0050] like Figure 2 As shown, the improved discrete cosine transform domain generative adversarial network specifically includes:

[0051] An improved Discrete Cosine Transform (MDCT) domain feature extraction module is used to extract the improved MDCT of the lost speech signal as a time-frequency domain feature by performing an improved MDCT on the lost speech signal.

[0052] A generator network based on dual-path recurrent convolution is used to compensate for lost speech signals by utilizing the time-frequency domain features of lost speech signals to generate pseudo-complete speech signals. The generator network specifically includes an encoder module, a dual-path long short-term memory network (LSTM) module, and a decoder module connected in sequence. The encoder module includes several encoder layers, and the decoder module includes several decoder layers. The encoder and decoder layers are connected by a SimAM attention module.

[0053] An improved Inverse Discrete Cosine Transform (iMDCT) module is used to perform an improved inverse discrete cosine transform on the features of pseudo-complete speech signals to obtain pseudo-complete speech signals;

[0054] A discriminator network is used to distinguish between genuine and pseudo-genuine speech signals;

[0055] The loss function calculation module combines the generator network loss, the time-domain weighted average absolute error loss, the improved discrete cosine transform loss, and the speech quality perception evaluation loss to calculate the total loss, thereby updating the parameters of the generator network.

[0056] like Figure 2 As shown, the encoder module includes a cascaded first encoder layer and a second encoder layer. The first encoder layer includes a two-dimensional convolutional layer (2DConv with 32 / 64 channels and a kernel size of (5,3)), a two-dimensional batch normalization layer (BatchNorm), a one-dimensional pixel unshuffle layer (PixelUnshuffle with a downsampling factor of (1,2)), a PReLU activation function layer, and a residual connection operation, all connected in sequence. The second encoder layer includes a two-dimensional convolutional layer (2dConv with 32 / 64 channels and a kernel size of (5,3)), a two-dimensional batch normalization layer (BatchNorm), a PReLU activation function layer, a dual-path feedforward sequence memory neural network layer (FSMN), a Snake activation function layer, and a residual connection operation, all connected in sequence.

[0057] like Figure 2 As shown, the decoder module includes a first decoder layer and a second decoder layer. The first decoder layer includes a two-dimensional convolutional layer (2DConv with 64 / 32 channels and a kernel size of (5,3)) connected in sequence, a two-dimensional batch normalization layer (BatchNorm), a PReLU activation function layer, a dual-path FSMN layer, and so on. The second decoder layer consists of a Snake activation function layer with a value of 5, a SimAM attention layer, and a residual connection operation. The second decoder layer includes a two-dimensional convolutional layer (2dConv with 64 / 32 channels and a kernel size of (5,3)), a two-dimensional batch normalization layer (BatchNorm), a one-dimensional pixel rearrangement layer (PixelShuffle with an upsampling factor of (1,2)) and a PReLU activation function layer, which are connected in sequence.

[0058] like Figure 3 and Figure 4 As shown, the dual-path Long Short-Term Memory (LSTM) network module specifically includes several sequentially connected intra-block LSTM modules and inter-block LSTM modules. That is, the intra-block LSTM modules and inter-block LSTM modules are connected as a sub-module, specifically several sequentially connected sub-modules. The intra-block LSTM module includes a sequentially connected bidirectional long short-term memory network (Bi-LSTM), a linear layer, layer normalization, and residual connection operations. The inter-block LSTM module includes a sequentially connected unidirectional LSTM and a linear layer. The intra-block LSTM path uses a bidirectional long short-term memory network (Bi-LSTM with 64 channels) as the core computational unit, performing serialization processing on the single-frame spectrum along the frequency dimension. The inter-block LSTM path uses a unidirectional LSTM with 64 channels as the core computational unit, recombining the feature values ​​at the same frequency position in each time frame into a time series along the time dimension.

[0059] The SimAM attention module is used to perform the following calculation process:

[0060] ,

[0061] In the formula, As input to the SimAM attention module, The number of channels is 32 / 64. For spatial dimensions, For output, This is the weight matrix. for function, For the minimum energy value The minimum energy value of the matrix formed by taking the reciprocals of each point. for:

[0062] ,

[0063] in, This represents the global mean of the channel containing the feature point. This represents the global variance of the channel containing the feature point. For hyperparameters, express The t-th feature point, i.e., the part in the t-th spatial dimension. The SimAM attention module will calculate the weight matrix. The feature is enhanced by multiplying the input X point by point.

[0064] The discriminator network is used to distinguish between genuine and fake speech signals, including complete speech signals, speech signals with packet loss, and speech signals compensated for by the generator. Its goal is to correctly differentiate between real and generated data. Its output is a score for the authenticity of the input speech signal; a higher score indicates a greater probability that the input signal is a genuine, complete speech signal. This score is used to calculate the generator loss and discriminator loss in the generative adversarial loss function. Figure 5 As shown, the discriminator network includes a multi-scale discriminator network and a multi-period discriminator network. For the multi-scale discriminator, the input speech data is downsampled at different scales (scale values ​​are 2, 4, and 8) to generate multiple low-resolution scale signals. The signal at each scale is input into an independent sub-discriminator network. These sub-discriminator networks are composed of stacked depthwise separable convolutional layers (1DConv). The higher the resolution of the sub-discriminator network, the smaller the convolutional kernel and the deeper the network. For the multi-period discriminator, the input speech data is divided into multiple sub-sequences according to different period lengths (period values ​​are 2, 3, and 5). The sub-sequence group of each period is input into an independent sub-discriminator network. These sub-discriminators use asymmetric period-size convolutional kernels to perform temporal convolution operations.

[0065] The function for calculating the total loss is:

[0066] ,

[0067] In the formula, For the total loss, To represent the time-domain weighted MAE loss The weighting coefficients, Indicates the loss in perceived speech quality. The weighting coefficients, The amplitude MAE loss of the improved discrete cosine transform spectrum. The weighting coefficients, Represents the polarity MAE loss of the improved discrete cosine transform spectrum. The weighting coefficients, Indicate generator loss The weighting coefficients. In this embodiment, the weighting coefficients... , 0.75, , , .

[0068] The functions for calculating the generator loss and discriminator loss are:

[0069] ,

[0070] ,

[0071] In the formula, For the discriminator to detect real, complete speech signals The output, For the discriminator network to process the pseudo-complete speech signal generated by the generator. The output, Represents the generator network function. It is the lost audio signal input to the generator network. Represents pseudo-complete speech signal The distribution Represents a complete speech signal x The distribution This indicates that the mean value is being calculated.

[0072] It is the temporal loss MAE loss function calculated by weighting the lost frames labeled in the dataset, where lost frames have a larger weight and non-lost frames have a smaller weight. It is a loss function based on Perceptual Speech Quality Assessment (PESQ).

[0073] and The calculation formula is as follows:

[0074] ,

[0075] in, This represents the improved discrete cosine transform spectrum corresponding to the complete speech signal. This represents the improved discrete cosine transform spectrum of the speech signal after compensation by a generator network based on dual-path circular convolution.

[0076] Step 3: Input the training dataset into the improved discrete cosine transform domain generative adversarial network to train the network.

[0077] Step 4: Input the lost speech signal to be tested into the trained network to obtain the compensated complete speech signal.

[0078] To verify the effectiveness of this invention, the method of this embodiment was simulated. In the Audio Deep Packet Loss Concealment Challenge, the organizers provided two key no-reference quality evaluation metrics, DNSMOS (Deep Noise Suppression Mean Opinion Score) and PLCMOS (Packet Loss Concealment Mean Opinion Score), to evaluate packet loss compensation performance. DNSMOS constructs a deep neural network to establish a mapping relationship between human subjective ratings and deep features of the speech signal. Its training data covers noise-suppressed speech segments and corresponding P.835 subjective ratings, which include three dimensions: speech quality, background noise, and overall quality. PLCMOS focuses on quality evaluation in packet loss compensation scenarios. Since packet loss can cause sudden interruptions or phase discontinuities in the speech signal, traditional metrics based on complete reference signals are difficult to apply. The model architecture of PLCMOS is similar to that of DNSMOS, but its training data is specifically constructed for packet loss scenarios. This invention compares the performance of two types of datasets, different combinations of loss functions, and whether or not generative adversarial training is adopted, thereby analyzing the impact of different dataset generation methods, different combinations of loss functions, and generative adversarial training on speech packet loss compensation algorithms.

[0079] The final performance evaluation is as follows:

[0080] A. The performance comparison of the two dataset generation schemes is shown in Table 1:

[0081] Table 1 Performance evaluation of different dataset generation methods

[0082]

[0083] B. The impact of combining loss functions is shown in Table 2:

[0084] Table 2. Impact of Loss Function Combination on Speech Packet Loss Compensation Algorithm

[0085]

[0086] C. The impact of generative adversarial training is shown in Table 3:

[0087] Table 3. Impact of Generative Adversarial Training on Speech Packet Loss Compensation Algorithms

[0088]

[0089] Example 2

[0090] This invention provides a computer device that provides services for implementing the method described in Embodiment 1. The device may include: a memory storing a computer-executable program; a processor coupled to the memory; and the processor calling the computer-executable program stored in the memory to execute the steps of the method described in Embodiment 1.

[0091] The memory may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the memory may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). A program / utility having a set (at least one) of program modules may be stored in, for example, memory. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The computer-executable program of the program modules typically performs the functions and / or methods described in the embodiments of the present invention.

[0092] The processor executes various functional applications and data processing by running programs stored in memory, such as the method provided in Embodiment 1 of the present invention.

[0093] The code of a computer executable program can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages.

[0094] It should be understood that the embodiments and descriptions above are only the principles, main features and advantages of the present invention. Various changes and modifications can be made to the present invention without departing from the spirit and scope of the invention, and all such changes and modifications fall within the protection scope of the present invention.

Claims

1. A speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model, characterized in that, include: Step 1: Process the complete speech signal to obtain the lost speech signal, and combine several complete speech signals and corresponding lost speech signals to form sample pairs to form a training dataset; Step 2: Construct an improved Generative Adversarial Network (GAN) based on the Discrete Cosine Transform domain, specifically including: An improved discrete cosine transform domain feature extraction module is used to extract the time-frequency domain features of the lost speech signal by performing an improved discrete cosine transform on the lost speech signal. A generator network based on dual-path recurrent convolution is used to compensate for lost speech signals by utilizing the time-frequency domain features of lost speech signals to generate pseudo-complete speech signals. The generator network specifically includes an encoder module, a dual-path long short-term memory network (LSTM) module, and a decoder module connected in sequence. The encoder module includes several encoder layers, and the decoder module includes several decoder layers. The encoder and decoder layers are connected by a SimAM attention module. An improved inverse discrete cosine transform module is used to perform an improved inverse discrete cosine transform on the features of pseudo-complete speech signals to obtain pseudo-complete speech signals. A discriminator network is used to distinguish between genuine and pseudo-genuine speech signals; The loss function calculation module is used to combine the generator network loss, the temporal MAE loss, the improved discrete cosine transform loss, and the speech quality perception evaluation loss to calculate the total loss, thereby updating the parameters of the generator network. Step 3: Input the training dataset into the improved discrete cosine transform domain generative adversarial network to train the network; Step 4: Input the lost speech signal to be tested into the trained network to obtain the compensated complete speech signal.

2. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 1, characterized in that: The encoder module includes a cascaded first encoder layer and a second encoder layer. The first encoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a one-dimensional pixel inverse rearrangement layer, a PReLU activation function layer, and a residual connection operation connected in sequence. The second encoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a PReLU activation function layer, a dual-path feedforward sequence memory neural network layer, a Snake activation function layer, and a residual connection operation connected in sequence.

3. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 1, characterized in that: The decoder module includes a first decoder layer and a second decoder layer. The first decoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a PReLU activation function layer, a dual-path feedforward sequence memory neural network layer, a Snake activation function layer, a SimAM attention layer, and a residual connection operation connected in sequence. The second decoder layer includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, a one-dimensional pixel rearrangement layer, and a PReLU activation function layer connected in sequence.

4. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 1, characterized in that: The dual-path long short-term memory network LSTM module specifically includes several intra-block LSTM modules and inter-block LSTM modules connected in sequence. The intra-block LSTM module includes a bidirectional long short-term memory network Bi-LSTM, a linear layer, a normalization layer, and a residual connection operation connected in sequence. The inter-block LSTM module includes a unidirectional LSTM and a linear layer connected in sequence.

5. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 1, characterized in that: The SimAM attention module is used to perform the following calculation process: , In the formula, X is the input to the SimAM attention module. For output, This indicates point-by-point multiplication. for function, The weight matrix is ​​composed of the minimum energy value. The minimum energy value is obtained by taking the reciprocal of each point. for: , in, express The t-th feature point, This represents the global mean of the channel containing the feature point. This represents the global variance of the channel containing the feature point. This is a hyperparameter.

6. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 1, characterized in that: The discriminator network includes a multi-scale discriminator network and a multi-period discriminator network. For the multi-scale discriminator, the input speech data is downsampled at different scales to generate multiple low-resolution signals. Each scale signal is input into an independent sub-discriminator network. These sub-discriminator networks are composed of stacked depthwise separable convolutional layers. The higher the resolution of the sub-discriminator network, the smaller the convolutional kernel and the deeper the network. For the multi-period discriminator, the input speech data is divided into multiple sub-sequences according to different period lengths. Each period sub-sequence group is input into an independent sub-discriminator network. These sub-discriminators use asymmetric period-size convolutional kernels to perform temporal convolution operations.

7. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 1, characterized in that: The function for calculating the total loss is: , In the formula, For the total loss, To represent the time-domain weighted MAE loss The weighting coefficients, Indicates the loss in perceived speech quality. The weighting coefficients, The amplitude MAE loss of the improved discrete cosine transform spectrum. The weighting coefficients, Represents the polarity MAE loss of the improved discrete cosine transform spectrum. The weighting coefficients, Indicate generator loss The weighting coefficients.

8. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 7, characterized in that: The generator loss is calculated using the following function: , In the formula, For the discriminator network to process the pseudo-complete speech signal generated by the generator. The output, Represents the generator network function. It is the lost audio signal input to the generator network. Represents pseudo-complete speech signal The distribution, This indicates that the mean value is being calculated.

9. The speech packet loss compensation method based on an improved discrete cosine transform domain generative adversarial model according to claim 7, characterized in that: The amplitude MAE loss of the improved discrete cosine transform spectrum and the polarity MAE loss of the improved discrete cosine transform spectrum The calculation formula is as follows: , in, This represents the improved discrete cosine transform spectrum corresponding to the complete speech signal. This represents the improved discrete cosine transform spectrum of the speech signal after compensation by a generator network based on dual-path circular convolution.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Audio signal processing method, and audio generation model training method and device

    CN114866856A

  • Audio processing method and device, storage medium and electronic equipment

    CN118136030A