A method for compensating for voice packet loss, a method for voice communication and a device
By using pre-trained target generation adversarial network to compensate voice data for packet loss, the problem of difficulty in dealing with long-term, continuous, and burst packet loss is solved, and efficient voice packet loss compensation and audio quality improvement is achieved.
Patent Information
- Application Number
- CN202210617394.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The prior art is difficult to effectively handle long-term, continuous, and burst voice packet loss, especially in real-time communication technologies such as VOIP, resulting in poor audio quality.
The pre-trained target generation adversarial network is used to compensate for packet loss for voice data, and the unlost speech frames before the lost speech frame are sorted and reconstructed, so as to realize the processing of long-term packet loss.
It improves the real-time and quality of voice packet loss compensation, can effectively handle long-term, continuous and sudden packet loss situations, and improves audio quality.
Smart Images

Figure CN115171705B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice processing, and in particular to a method for compensating for voice packet loss, a method for voice communication and a device. Background Art
[0002] With the development of Internet technology, real-time communication technology (RTC) has been widely used, such as in live broadcast, online education, audio and video conferencing, interactive games, etc. For audio links, real-time communication technology mainly includes acquisition, pre-processing, encoding, jitter elimination, decoding, packet loss compensation, mixing, playback and other links. In communication methods such as VOIP (Voice Over Internet Phone) that use real-time communication technology, the audio data is encoded and compressed, and then transmitted in frames on the network. Since the data packet switching technology based on the IP protocol provides a "best-effort service", it will inevitably cause packet delay and packet loss, which will cause poor audio quality. Therefore, the packet loss compensation link is particularly important.
[0003] In the Packet Loss Concealment (PLC) technology, the waveform of the lost packet is predicted by the parameters of the normal received packet, which can include the compensation method based on the sender and the compensation method based on the receiver. The compensation method based on the sender is to use the coding redundant information to restore the content of the lost packet, and the compensation method based on the receiver is to use the decoding parameter information before the packet loss to reconstruct the voice signal. However, the above method can usually only handle short-term data loss, such as packet loss less than 40ms, and is difficult to apply to long-term, continuous, and sudden packet loss. Summary of the invention
[0004] In view of the above problems, a method for voice packet loss compensation, a method and apparatus for voice call are proposed to overcome the above problems or at least partially solve the above problems, including:
[0005] A method for voice packet loss compensation, the method comprising:
[0006] Get a pre-trained target generative adversarial network;
[0007] Acquire voice data and use a target-generated adversarial network to compensate for packet loss of the voice data;
[0008] In the process of packet loss compensation, for a first speech frame with data loss in the speech data, a second speech frame sorted before the first speech frame is used for reconstruction in the target generative adversarial network; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0009] Optionally, for a first speech frame with data loss in the speech data, before reconstructing the first speech frame in the target generative adversarial network using a second speech frame sorted before the first speech frame, the method further includes:
[0010] A mask is set for a voice frame in the voice data to identify whether data loss occurs therein.
[0011] Optionally, the target generative adversarial network has a generator and a discriminator. The target generative adversarial network is trained in a confrontation manner between the generator and the discriminator. The training process of the target generative adversarial network includes:
[0012] In the process of training the target generative adversarial network, the generator is used to compensate for packet loss of sample data with data loss, and the generator is used to identify the sample data after packet loss compensation, so as to adjust the generator according to the identification result.
[0013] Optionally, the generator has an encoder and a decoder using a U_net structure, the encoder is used to extract speech features, and the decoder is used to reconstruct according to the speech features.
[0014] Optionally, in the process of training the target generative adversarial network, the encoder is trained using a semi-supervised learning approach.
[0015] Optionally, the loss function of the target generative adversarial network is composed of multiple losses, including:
[0016] Generative adversarial loss for target generative adversarial networks, loss for time domain waveforms, short-time Fourier transform loss for multi-resolution, and consistency loss for semi-supervised learning.
[0017] Optionally, the discriminator is composed of a combination of multiple discriminators, and the multiple discriminators include:
[0018] Multi-cycle discriminator, multi-scale discriminator, multi-dilation discriminator.
[0019] Optionally, a bottleneck layer connection is adopted between the encoder and the decoder, the encoder and the decoder have the same number of multi-level processing units, and inter-layer skip connections are arranged between processing units at the same level.
[0020] A method for voice communication, the method comprising:
[0021] During a voice call, voice data is acquired and a pre-trained target-generated adversarial network is used to compensate for packet loss in the voice data.
[0022] In the process of packet loss compensation, for a first speech frame with data loss in the speech data, a second speech frame sorted before the first speech frame is used for reconstruction in the target generative adversarial network; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0023] A device for voice packet loss compensation, the device comprising:
[0024] A target generative adversarial network acquisition module is used to acquire a pre-trained target generative adversarial network;
[0025] A first packet loss compensation module is used to obtain voice data and use a target generative adversarial network to perform packet loss compensation on the voice data;
[0026] The first speech frame reconstruction module is used to reconstruct the first speech frame with data loss in the speech data by using the second speech frame sorted before the first speech frame in the target generative adversarial network during the process of packet loss compensation; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0027] A device for voice communication, comprising:
[0028] The second packet loss compensation module is used to obtain voice data during the voice call and use a pre-trained target generative adversarial network to perform packet loss compensation on the voice data;
[0029] The second speech frame reconstruction module is used to reconstruct the first speech frame with data loss in the speech data by using the second speech frame sorted before the first speech frame in the target generative adversarial network during the process of packet loss compensation; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0030] An electronic device includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the method for compensating for voice packet loss as described above is implemented, or the method for voice call as described above is implemented.
[0031] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for compensating for voice packet loss or the method for voice call is implemented.
[0032] The embodiments of the present invention have the following advantages:
[0033] In an embodiment of the present invention, a pre-trained target generative adversarial network is obtained, and when voice data is obtained, the target generative adversarial network is used to perform packet loss compensation on the voice data. In the process of packet loss compensation, for a first voice frame with data loss in the voice data, a second voice frame sorted before the first voice frame is used in the target generative adversarial network for reconstruction, and the second voice frame is a voice frame with no data loss. This realizes voice packet loss compensation using a generative adversarial network, and not only reconstructs by using voice frames with no data loss, but avoids the impact on voice quality when a large number of voice frames with data loss but reconstructed are used for packet loss compensation. This can be applied to long-term, continuous, and sudden packet loss situations, and reconstructs by using voice frames sorted in front, without considering voice frames sorted in the back, and can process voice frames with data loss in parallel, thereby improving the real-time performance of packet loss compensation. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0035] Figure 1a is a schematic diagram of the architecture of a real-time communication system provided by an embodiment of the present invention;
[0036] Figure 1b is a schematic diagram of an audio link architecture provided by an embodiment of the present invention;
[0037] Figure 2 This is a flow chart of the steps of voice packet loss compensation provided by one embodiment of the present invention;
[0038] Figure 3 It is a schematic diagram of a framework of a target generative adversarial network provided by an embodiment of the present invention;
[0039] Figure 4a is a schematic diagram of POLQA scores under different packet loss rates provided by an embodiment of the present invention;
[0040] Figure 4b is a schematic diagram of PESQ scores under different packet loss rates provided by an embodiment of the present invention;
[0041] Figure 4c is a schematic diagram of an average STOI score under different packet loss rates provided by an embodiment of the present invention;
[0042] Figure 5ais a schematic diagram of an average POLQA score under different packet loss rates provided by an embodiment of the present invention;
[0043] Figure 5b is a schematic diagram of an average PESQ score under different packet loss rates provided by an embodiment of the present invention;
[0044] Figure 5c is a schematic diagram of an average STOI score under different packet loss rates provided by an embodiment of the present invention;
[0045] Figure 6 is a flowchart of a method for voice communication provided by an embodiment of the present invention;
[0046] Figure 7 It is a structural block diagram of a device for voice packet loss compensation provided by an embodiment of the present invention;
[0047] Figure 8 The present invention is a block diagram of a voice communication device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] The embodiments of the present invention can be applied to communication scenarios. In a practical application, they are particularly applicable to communication scenarios based on real-time communication technology. Real-time communication technology refers to communication technology that can send and receive text, audio, and video in real time. It is applicable to scenarios such as live broadcast, on-demand, video conferencing, online classrooms, online chat rooms, and game interactions, and realizes real-time transmission of pure audio data, video data, and the like. The embodiments of the present invention can be specifically applied to communication scenarios such as live broadcast, on-demand, video conferencing, online classrooms, online chat rooms, and game interactions based on real-time communication technology.
[0050] See also Figure 1a , shows a schematic diagram of the architecture of a real-time communication system to which the embodiment of the present invention can be applied, which may include a server 100 and multiple clients 200. Multiple clients 200 may establish communication connections through the server 100. In the real-time communication scenario, the server 100 is used to provide real-time communication services between multiple clients 200. Multiple clients 200 may respectively serve as senders or receivers to achieve real-time communication through the server 100.
[0051] The user can interact with the server 100 through the client 200 to receive data sent by other clients 200, or send data to other clients 200, etc. In a real-time communication scenario, the user can publish a data stream to the server 100 through the client 200, and the server 200 pushes the data stream to the client that subscribes to the data stream. The data stream can be, for example, media data such as an audio stream and a video stream. For example, in a live broadcast scenario, the anchor user can collect media data in real time through the client and send it to the server. The media data of different anchor users are distinguished by the live broadcast room. The server can push the media data of the anchor user to the viewing users who enter the corresponding live broadcast room of the anchor user. For example, in a conference scenario, the participating users can collect media data in real time through the client and send it to the server. The server can push the media data sent by each client to the clients of other participating users, etc.
[0052] The data transmitted by the client 200 may need to be encoded, transcoded, compressed, etc. before being published to the server 100. The client 200 and the server 100 are connected via a network, which provides a medium for the communication link between the client and the server. The network may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0053] Among them, the client 200 can be a browser, an application (APP, Application), or a web application such as H5 (HyperText Markup Language5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 200 can be based on the software development kit (SDK, Software Development Kit) of the corresponding service provided by the server, such as developed based on RTC SDK, etc. The client 200 can be deployed in an electronic device and needs to rely on the device to run or some apps in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0054] Among them, the server 100 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that provide support for models used on clients, and servers that process data sent by clients. The server 100 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0055] It should be noted that the voice packet loss compensation method and the voice call method provided in the embodiments of the present invention are generally executed by the server, and the corresponding voice packet loss compensation device and the voice call device are generally set in the server. However, in other embodiments of the present invention, the client may also have similar functions to the server, so as to execute the voice packet loss compensation method and the voice call method provided in the embodiments of the present invention. In other embodiments, the voice packet loss compensation method and the voice call method provided in the present invention may also be jointly executed by the client and the server.
[0056] For the audio link, it mainly includes acquisition, pre-processing, encoding, jitter elimination, decoding, packet loss compensation, mixing, playback and other links, such as Figure 1b The figure shows the architecture diagram of the audio link, which can be divided into the process of audio sending (stream pushing) and the process of audio receiving (stream pulling).
[0057] The process of audio transmission may include acquisition, pre-processing, encoding and other links. Specifically, the audio signal may be acquired through an acquisition module, such as a microphone, and then the analog signal may be converted into a data signal, and then the audio signal may be pre-processed.
[0058] Among them, the pre-processing can include three parts: acoustic echo cancellation (AEC, Acoustic Echo Canceller), automatic noise control (ANS, Automatic Noise Suppression), automatic gain control (AGC, Automatic Gain Control), which can perform acoustic echo cancellation, automatic noise control, and automatic gain control on the audio signal in turn.
[0059] After the audio signal is pre-processed, audio encoding can be performed, that is, the audio signal is compressed and encoded, and then the compressed and encoded audio signal is packaged and sent to a network server through the network.
[0060] The process of audio reception may include jitter elimination, decoding, packet loss compensation, mixing, playback and the like. Specifically, the audio data packet may be firstly subjected to jitter elimination, such as by using a jitter buffer for jitter elimination, and then the audio data packet may be subjected to audio decoding.
[0061] For the decoded audio data packets, if there is data loss in the voice frames, packet loss compensation can be performed on the voice frames with data loss. After packet loss compensation, multiple audio streams can be mixed (MIX) and then played through a playback module, such as a speaker.
[0062] In packet loss compensation technology, the waveform of the lost packet is predicted by the parameters of the normal received packet, which can include the compensation method based on the sender and the compensation method based on the receiver. The compensation method based on the sender uses the coding redundant information to restore the content of the lost packet, and the compensation method based on the receiver uses the decoding parameter information before the packet loss to reconstruct the voice signal. However, the above method can usually only handle short-term data loss, such as packet loss of less than 40ms, and is difficult to apply to long-term, continuous, and sudden packet loss.
[0063] In the field of deep learning, there are generative adversarial networks (GAN), recurrent neural networks (RNN), autoencoder networks (AutoEncoder), etc., which have great advantages in generating high-quality speech. Applying them in the packet loss compensation link can achieve good results. The working framework of the packet loss compensation algorithm based on deep learning can include an offline processing framework and a real-time processing framework.
[0064] For the offline processing framework, in addition to using historical non-lost frames, it is also possible to use a broader context including future frames, which is not suitable for real-time streaming fast processing. For example, if the jth frame is lost, the historical jmth frame and the future j+mth frame need to be sent to the deep learning network to generate the jth frame speech signal.
[0065] For the real-time processing framework, an algorithm is used for post-processing, and only the historical frames that have not been lost are used. For example, if the jth frame is lost, the jmth frame of the history needs to be sent to the deep learning network to generate the jth frame of speech signal. Specifically, the real-time processing framework can use a recurrent neural network or a generative adversarial network.
[0066] The method based on recurrent neural network uses the effective information of the previous frame to recursively infer the current frame, which will cause the context of the frame to be predicted to contain many reconstructed frames instead of the original frame, resulting in a mismatch between training and reasoning. Especially when there is long-term, continuous, and burst packet loss, the energy of the generated waveform will be greatly attenuated, and the voice quality needs to be further improved. In addition, the phase of the compensated speech signal and the real speech signal is discontinuous, and a smoothing operation is required, which will reduce the quality of the speech signal. The methods based on generative adversarial networks have high computational complexity in terms of the number of parameters and inference delay, making them difficult to use for real-time processing.
[0067] In an embodiment of the present invention, a pre-trained target generative adversarial network is obtained, and when voice data is obtained, the target generative adversarial network is used to perform packet loss compensation on the voice data. In the process of packet loss compensation, for a first voice frame with data loss in the voice data, a second voice frame sorted before the first voice frame is used in the target generative adversarial network for reconstruction, and the second voice frame is a voice frame with no data loss. This realizes voice packet loss compensation using a generative adversarial network, and by using voice frames with no data loss for reconstruction, the impact on voice quality when a large number of voice frames with data loss but reconstructed are used for packet loss compensation is avoided. This can be applied to long-term, continuous, and sudden packet loss situations. By using the voice frames sorted first for reconstruction, there is no need to consider the voice frames sorted later, and the voice frames with data loss can be processed in parallel, thereby improving the real-time performance of packet loss compensation.
[0068] The following further describes the embodiments of the present invention:
[0069] Reference Figure 2 , shows a flowchart of a method for voice packet loss compensation provided by an embodiment of the present invention, which may specifically include the following steps:
[0070] Step 201, obtaining a pre-trained target generative adversarial network.
[0071] Among them, the target generative adversarial network can have a generator and a discriminator.
[0072] For the generator:
[0073] The generator has an encoder and a decoder using a U_net structure. Through the U_net structure (which is causal), the real-time performance of packet loss compensation of the target generative adversarial network can be guaranteed.
[0074] Among them, the encoder can be used to extract speech features, and the decoder can be used to reconstruct according to the speech features. In the encoder, the dimension of the speech feature mapping can be reduced by downsampling, such as mapping a 16KHZ waveform to 50HZ, thereby reducing the number of parameters and the amount of calculation. After dimensionality reduction, data training and feature extraction can be more effective and intuitive. In the decoder, the dimension of the speech features can be increased by upsampling, and the speech features can be restored to the same dimension as the speech data.
[0075] In one embodiment of the present invention, a bottleneck layer (Bottleneck) can be used between the encoder and the decoder to connect. The bottleneck layer can be composed of 2 layers of 1-D causal convolution, which can model the temporal correlation, thereby improving the network's ability to learn temporal correlation and enhance feature correlation.
[0076] In one embodiment of the present invention, the encoder and the decoder may have the same number of multi-level processing units, and inter-layer skip connections may be set between the processing units at the same level, thereby allowing phase or alignment information to pass through, ensuring that the low-dimensional features of the input audio are not lost. In one example, each processing unit may have multiple residual units, such as 3 residual units, and each residual unit alternately uses 1-D dilated convolution and 1-D convolution.
[0077] like Figure 3 In the framework of the target generative adversarial network shown in the figure, there may be four processing units EncoderBlock1-EncoderBlock4 and DecoderBlock1-DecoderBlock4 in the encoder and decoder respectively. The encoder and decoder are connected through a bottleneck layer, and inter-layer jump connections are set between the processing units.
[0078] For the discriminator:
[0079] In order to maximize the ability of the discriminator in the target generative adversarial network to distinguish synthetic or real audio, the discriminator can be composed of multiple discriminators, which can then recognize speech signals from different angles, such as Figure 3 As shown in , the various discriminators may include: Multi-Period Discriminator (MPD), Multi-Scale Discriminator (MSD), and Multi-Dilation Discriminator (MDD).
[0080] Among them, the multi-period discriminator can fold a mono audio sequence into a two-channel audio with different fixed lengths, and then perform 2-D convolution on the folded data, but the folded data on each channel mixes artifacts of different frequencies.
[0081] The multi-scale discriminator can halve the length of the speech sequence through average pooling operation, then perform convolution operations on speech signals of different scales, and finally flatten and output them.
[0082] The multi-period discriminator can fold monophonic audio into multi-channel audio by wavelet transform, and then apply 1-D dilated convolution. Then each channel in the folded data contains few or even no artifacts of other frequencies, ensuring the stability and accuracy of the discrimination.
[0083] In one embodiment of the present invention, the target generative adversarial network can be trained by using a generator and a discriminator adversarial method. Accordingly, in step 201, the training process of the target generative adversarial network may include: in the process of training the target generative adversarial network, using the generator to compensate for packet loss of sample data with data loss, and using the generator to identify the sample data after packet loss compensation, so as to adjust the generator according to the identification result.
[0084] In a specific implementation, the generator can perform packet loss compensation on sample data with data loss in parallel to obtain sample data after packet loss compensation, and then input it into the discriminator. The discriminator can identify the sample data after packet loss compensation input by the generator, and then guide the generator to learn according to the identification result, so that the generator can synthesize samples that are close to the real ones, so that the discriminator cannot distinguish between the real and generated samples.
[0085] In one embodiment of the present invention, in order to further improve the ability of the encoder to extract global features, in the process of training the target generative adversarial network, the encoder can be trained in a semi-supervised learning (Mean Teacher) manner. In the process of semi-supervised learning, the encoder can have two models (teacher model and student model). The teacher model is first used to encode sample data without data loss to generate a learning target for the student model. The weights in the teacher model can be used as the exponential moving average (EMA) of the weights in the student model. The student model encodes sample data with data loss and predicts a complete data representation.
[0086] In one embodiment of the present invention, the loss function of the target generative adversarial network can be composed of multiple losses, and the multiple losses can include: the generative adversarial loss of the target generative adversarial network, the loss of the time domain waveform, the multi-resolution short-time Fourier transform (STFT, Short-time Fourier Transform) loss, and the consistency loss of semi-supervised learning. Through the above loss function, the time-frequency distribution of the real speech waveform can be effectively captured, and the entire network can be easily trained even with a small number of parameters, and the reasoning time can be effectively reduced and the perceived quality of the synthesized speech can be improved.
[0087] Among them, since the target generative adversarial network is trained in a generator-discriminator adversarial manner, the generative adversarial loss of the target generative adversarial network can be calculated based on the real lossless audio signal and the lossy audio signal.
[0088] The loss of the time domain waveform can be obtained by calculating the L1 distance between the true waveform and the generated waveform.
[0089] The multi-resolution short-time Fourier transform loss can be obtained based on the spectral convergence loss and the logarithmic STFT amplitude spectrum loss.
[0090] Since the encoder can be trained using semi-supervised learning, the consistency loss of semi-supervised learning can be obtained based on the L2 distance between the output of the teacher model and the output of the student model.
[0091] Step 202: Acquire voice data and use a target generative adversarial network to perform packet loss compensation on the voice data.
[0092] For voice data, it can be voice data obtained during real-time communication, such as voice data obtained in VOIP communication. Since the obtained voice data may have data packet loss, a target generative adversarial network can be used to compensate for packet loss of the voice data to reconstruct voice frames with data loss.
[0093] Step 203, in the process of packet loss compensation, for a first speech frame with data loss in the speech data, a second speech frame sorted before the first speech frame is used in the target generative adversarial network for reconstruction; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0094] In order to be compatible with processing real-time audio streams on low-power mobile devices, the target generative adversarial network needs to be small in size and have low CPU occupancy. The convolutional layer in the target generative adversarial network can be causal, that is, it only relies on historical information but not future information. In the process of packet loss compensation, since the voice frames in the voice data have a certain time order, the voice frames sorted before the first voice frame can be used to reconstruct the voice frames with data loss. By reconstructing with the voice frames sorted in front, the voice frames with data loss can be processed in parallel, and the smoothness and continuity of the audio before and after the packet loss can be guaranteed without a separate smoothing operation, which not only ensures the continuity of the phase, but also does not produce any delay in the overall operation, thereby improving the real-time performance of packet loss compensation.
[0095] Moreover, in order to improve the quality of reconstructed speech, only speech frames without data loss can be used to reconstruct speech frames with data loss, that is, the original speech frames in the speech data are used instead of the reconstructed speech frames. This can be applied to long-term, continuous, and sudden packet loss situations.
[0096] In an embodiment of the present invention, before step 203, the following steps may also be included:
[0097] A mask is set for a voice frame in the voice data to identify whether data loss occurs therein.
[0098] For speech frames in speech data, the target generative adversarial network can determine whether there is data loss. If there is no data loss, it can be masked to 0, and it can be directly output without processing, and can be saved in the sliding buffer for subsequent use as a basis for speech frame reconstruction. If there is data loss, it can be masked to 1 to cover its sampling points to indicate that it needs to be reconstructed.
[0099] In one example, during the process of training the target generative adversarial network, a mask may also be set for the speech frames in the input sample data to identify whether there is data loss.
[0100] In an embodiment of the present invention, a pre-trained target generative adversarial network is obtained, and when voice data is obtained, the target generative adversarial network is used to perform packet loss compensation on the voice data. In the process of packet loss compensation, for a first voice frame with data loss in the voice data, a second voice frame sorted before the first voice frame is used in the target generative adversarial network for reconstruction, and the second voice frame is a voice frame with no data loss. This realizes voice packet loss compensation using a generative adversarial network, and not only reconstructs by using voice frames with no data loss, but avoids the impact on voice quality when a large number of voice frames with data loss but reconstructed are used for packet loss compensation. This can be applied to long-term, continuous, and sudden packet loss situations, and reconstructs by using voice frames sorted in front, without considering voice frames sorted in the back, and can process voice frames with data loss in parallel, thereby improving the real-time performance of packet loss compensation.
[0101] In order to demonstrate the effect of the embodiment of the present invention, the embodiment of the present invention was experimented by using three measurement indicators:
[0102] 1. Mean Opinion Score (MOS) is a method for evaluating the quality of sentences read aloud by male and female speakers by a large number of listeners. The listeners rate each sentence according to the following criteria: 1 for very poor; 2 for poor; 3 for fair; 4 for good; 5 for very good. After summarizing the scores, an average is taken. The mean opinion score ranges from 1 to 5, and the higher the score, the better the voice quality.
[0103] 2. Perceptual Evaluation of Speech Quality (PESQ), whose annotated code name in the International Telecommunication Union is ITU-TP.862, provides a predicted value of a subjective mean opinion score for objective speech quality evaluation by using the perceptual objective listening quality analysis (POLQA) algorithm, and can be mapped to a mean opinion score scale range. The perceptual objective listening quality analysis score ranges from -0.5 to 4.5. The higher the score, the better the speech quality.
[0104] 3. Short Time Objective Intelligibility (STOI) measurement, the score is between 0-1. The larger the value, the higher the speech intelligibility and the better the speech quality.
[0105] Under the same experimental conditions, Figure 4a , Figure 4b , Figure 4cThe POLQA, PESQ, and STOI scores of the algorithm proposed in the present invention and the online neteq algorithm and Lossy algorithm under different packet loss rates. The horizontal axis is the different packet loss rates, and the vertical axis is the score. Figure 5a , Figure 5b , Figure 5c The average scores of POLQA, PESQ, and STOI of the algorithm proposed in the present invention and the online neteq algorithm and Lossy algorithm under different packet loss rates. It can be seen that the algorithm proposed in the present invention is superior to the online neteq algorithm and Lossy algorithm in POLQA, PESQ, and STOI evaluation, among which the perceived objective hearing quality analysis score is improved by an average of 0.5 points, the speech quality perception evaluation is improved by an average of 0.53 points, and the short-time target intelligibility measurement is improved by an average of 0.17 points. Moreover, the index of the embodiment of the present invention when the packet loss is 30% is better than the index of the traditional algorithm when the packet loss is 20%, that is, the embodiment of the present invention can keep the speech quality unchanged when the packet loss rate increases by 10%-15%.
[0106] Reference Figure 6 , shows a flowchart of a method for voice communication provided by an embodiment of the present invention, which may specifically include the following steps:
[0107] Step 601, during a voice call, voice data is acquired, and a pre-trained target generative adversarial network is used to compensate for packet loss of the voice data.
[0108] Step 602, in the process of packet loss compensation, for a first speech frame with data loss in the speech data, a second speech frame sorted before the first speech frame is used in the target generative adversarial network for reconstruction; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0109] In an embodiment of the present invention, before step 602, the following steps may also be included:
[0110] A mask is set for a voice frame in the voice data to identify whether data loss occurs therein.
[0111] In one embodiment of the present invention, the target generative adversarial network may have a generator and a discriminator. The target generative adversarial network may be trained in a generator-discriminator confrontation manner. The training process of the target generative adversarial network may include:
[0112] In the process of training the target generative adversarial network, the generator is used to compensate for packet loss of sample data with data loss, and the generator is used to identify the sample data after packet loss compensation, so as to adjust the generator according to the identification result.
[0113] In one embodiment of the present invention, the generator may have an encoder and a decoder using a U_net structure, the encoder may be used to extract speech features, and the decoder may be used to reconstruct based on the speech features.
[0114] In one embodiment of the present invention, during the process of training the target generative adversarial network, the encoder may be trained in a semi-supervised learning manner.
[0115] In one embodiment of the present invention, the loss function of the target generative adversarial network may be composed of multiple losses, and the multiple losses may include:
[0116] Generative adversarial loss for target generative adversarial networks, loss for time domain waveforms, short-time Fourier transform loss for multi-resolution, and consistency loss for semi-supervised learning.
[0117] In an embodiment of the present invention, the discriminator may be composed of a combination of multiple discriminators, and the multiple discriminators may include:
[0118] Multi-cycle discriminator, multi-scale discriminator, multi-dilation discriminator.
[0119] In one embodiment of the present invention, a bottleneck layer connection may be adopted between the encoder and the decoder, the encoder and the decoder may have the same number of multi-level processing units, and inter-layer jump connections may be provided between processing units at the same level.
[0120] It should be noted that, for the embodiment of the method for voice call, its specific content can refer to the description of the embodiment of the method for voice packet loss compensation in the above text.
[0121] For the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0122] Reference Figure 7 , shows a schematic diagram of the structure of a voice packet loss compensation device provided by an embodiment of the present invention, which may specifically include the following modules:
[0123] The target generative adversarial network acquisition module 701 can be used to acquire a pre-trained target generative adversarial network.
[0124] The first packet loss compensation module 702 can be used to obtain voice data and use a processing unit target generation adversarial network to perform packet loss compensation on the processing unit voice data.
[0125] The first speech frame reconstruction module 703 can be used to reconstruct the first speech frame with data loss in the speech data of the processing unit by using the second speech frame sorted before the first speech frame of the processing unit in the target generative adversarial network of the processing unit during the process of packet loss compensation; wherein the second speech frame of the processing unit is a speech frame other than the speech frame with data loss.
[0126] In one embodiment of the present invention, the method may further include:
[0127] The first mask setting module can be used to set a mask for the voice frame in the voice data of the processing unit to identify whether there is data loss.
[0128] In one embodiment of the present invention, the processing unit target generation adversarial network may have a generator and a discriminator. The processing unit target generation adversarial network may be trained in a confrontational manner between the processing unit generator and the processing unit discriminator. The training process of the processing unit target generation adversarial network may include:
[0129] In the process of training the processing unit target generative adversarial network, the processing unit generator is used to compensate for packet loss of sample data with data loss, and the processing unit generator is used to identify the sample data after packet loss compensation, so as to adjust the processing unit generator according to the identification result.
[0130] In one embodiment of the present invention, the processing unit generator may have an encoder and a decoder using a U_net structure, the processing unit encoder may be used to extract speech features, and the processing unit decoder may be used to reconstruct based on the speech features.
[0131] In one embodiment of the present invention, during the process of training the processing unit target generative adversarial network, the processing unit encoder may be trained in a semi-supervised learning manner.
[0132] In one embodiment of the present invention, the loss function of the processing unit target generative adversarial network may be composed of multiple losses, and the multiple losses of the processing unit may include:
[0133] The processing unit targets the generative adversarial loss of the generative adversarial network, the loss of the time domain waveform, the short-time Fourier transform loss of multiple resolutions, and the consistency loss of semi-supervised learning.
[0134] In one embodiment of the present invention, the processing unit identifier may be composed of a plurality of identifiers, and the processing unit multiple identifiers may include:
[0135] Multi-cycle discriminator, multi-scale discriminator, multi-dilation discriminator.
[0136] In one embodiment of the present invention, a bottleneck layer connection may be adopted between the processing unit encoder and the processing unit decoder, the processing unit encoder and the processing unit decoder may have the same number of multi-level processing units, and inter-layer jump connections may be set between processing units at the same level.
[0137] In an embodiment of the present invention, a pre-trained target generative adversarial network is obtained, and when voice data is obtained, the target generative adversarial network is used to perform packet loss compensation on the voice data. In the process of packet loss compensation, for a first voice frame with data loss in the voice data, a second voice frame sorted before the first voice frame is used in the target generative adversarial network for reconstruction, and the second voice frame is a voice frame with no data loss. This realizes voice packet loss compensation using a generative adversarial network. By using voice frames with no data loss for reconstruction, it can be applied to long-term, continuous, and sudden packet loss situations. By using voice frames sorted in front for reconstruction, voice frames with data loss can be processed in parallel, thereby improving the real-time performance of packet loss compensation.
[0138] Reference Figure 8 , shows a structural block diagram of a voice call device provided by an embodiment of the present invention, which may specifically include the following modules:
[0139] The second packet loss compensation module 801 can be used to obtain voice data during a voice call and use a pre-trained target generative adversarial network to perform packet loss compensation on the voice data;
[0140] The second packet loss compensation module 802 can be used to reconstruct the first speech frame with data loss in the speech data by using the second speech frame sorted before the first speech frame in the target generative adversarial network during the packet loss compensation process; wherein the second speech frame is a speech frame other than the speech frame with data loss.
[0141] In one embodiment of the present invention, the method may further include:
[0142] The second mask setting module can be used to set a mask for the voice frame in the voice data to identify whether there is data loss.
[0143] In one embodiment of the present invention, the target generative adversarial network may have a generator and a discriminator. The target generative adversarial network may be trained in a generator-discriminator confrontation manner. The training process of the target generative adversarial network may include:
[0144] In the process of training the target generative adversarial network, the generator is used to compensate for packet loss of sample data with data loss, and the generator is used to identify the sample data after packet loss compensation, so as to adjust the generator according to the identification result.
[0145] In one embodiment of the present invention, the generator may have an encoder and a decoder using a U_net structure, the encoder may be used to extract speech features, and the decoder may be used to reconstruct based on the speech features.
[0146] In one embodiment of the present invention, during the process of training the target generative adversarial network, the encoder may be trained in a semi-supervised learning manner.
[0147] In one embodiment of the present invention, the loss function of the target generative adversarial network may be composed of multiple losses, and the multiple losses may include:
[0148] Generative adversarial loss for target generative adversarial networks, loss for time domain waveforms, short-time Fourier transform loss for multi-resolution, and consistency loss for semi-supervised learning.
[0149] In an embodiment of the present invention, the discriminator may be composed of a combination of multiple discriminators, and the multiple discriminators may include:
[0150] Multi-cycle discriminator, multi-scale discriminator, multi-dilation discriminator.
[0151] In one embodiment of the present invention, a bottleneck layer connection may be adopted between the encoder and the decoder, the encoder and the decoder may have the same number of multi-level processing units, and inter-layer jump connections may be provided between processing units at the same level.
[0152] An embodiment of the present invention also provides an electronic device, which may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the method for voice packet loss compensation as described above is implemented, or the method for voice call as described above is implemented.
[0153] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for compensating for voice packet loss as described above is implemented, or the method for voice call as described above is implemented.
[0154] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0155] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0156] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0157] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0158] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0160] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0161] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0162] The above is a detailed introduction to a method for voice packet loss compensation, a method and a device for voice calls. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for voice packet loss compensation, characterized in that: The method comprises: Obtain a pre-trained target generative adversarial network; wherein the target generative adversarial network has a generator, and the generator has an encoder; the encoder has a teacher model and a student model, the teacher model is used to encode speech data without data loss and generate a learning target of the student model; the student model is used to encode speech data with data loss and predict a complete data representation; Acquire voice data, and use the target generative adversarial network to perform packet loss compensation on the voice data; In the process of packet loss compensation, for a first speech frame with data loss in the speech data, a second speech frame sorted before the first speech frame is used in the target generative adversarial network for reconstruction; wherein the second speech frame is a speech frame other than the speech frame with data loss; Wherein, for a first speech frame with data loss in the speech data, before reconstructing using a second speech frame sorted before the first speech frame in the target generative adversarial network, the method further includes: A mask is set for the voice frames in the voice data to identify whether data loss exists therein.
2. The method according to claim 1, characterized in that The target generative adversarial network has a discriminator, and the target generative adversarial network is trained in a confrontation manner between the generator and the discriminator. The training process of the target generative adversarial network includes: In the process of training the target generative adversarial network, the generator is used to perform packet loss compensation on sample data with data loss, and the generator is used to identify the sample data after packet loss compensation, so as to adjust the generator according to the identification result.
3. The method according to claim 2, characterized in that The generator has an encoder and a decoder using a U_net structure, the encoder is used to extract speech features, and the decoder is used to reconstruct according to the speech features.
4. The method according to claim 3, characterized in that In the process of training the target generative adversarial network, the encoder is trained using a semi-supervised learning method.
5. The method according to claim 4, characterized in that The loss function of the target generative adversarial network is composed of multiple losses, and the multiple losses include: The target generates adversarial losses of the adversarial network, time domain waveform losses, multi-resolution short-time Fourier transform losses, and consistency losses of semi-supervised learning.
6. The method according to claim 2, characterized in that The discriminator is composed of a plurality of discriminators, and the plurality of discriminators include: Multi-cycle discriminator, multi-scale discriminator, multi-dilation discriminator.
7. The method according to claim 3, characterized in that The encoder and the decoder are connected by a bottleneck layer. The encoder and the decoder have the same number of multi-level processing units, and inter-layer jump connections are set between processing units at the same level.
8. A method for voice communication, characterized in that: The method comprises: During a voice call, voice data is acquired, and a pre-trained target generative adversarial network is used to perform packet loss compensation on the voice data; wherein the target generative adversarial network has a generator, and the generator has an encoder; the encoder has a teacher model and a student model, the teacher model is used to encode voice data without data loss and generate a learning target of the student model; the student model is used to encode voice data with data loss and predict a complete data representation; In the process of packet loss compensation, for a first speech frame with data loss in the speech data, a second speech frame sorted before the first speech frame is used in the target generative adversarial network for reconstruction; wherein the second speech frame is a speech frame other than the speech frame with data loss; Wherein, for a first speech frame with data loss in the speech data, before reconstructing using a second speech frame sorted before the first speech frame in the target generative adversarial network, the method further includes: A mask is set for the voice frames in the voice data to identify whether data loss exists therein.
9. A device for voice packet loss compensation, characterized in that: The device comprises: A target generative adversarial network acquisition module is used to acquire a pre-trained target generative adversarial network; wherein the target generative adversarial network has a generator, and the generator has an encoder; the encoder has a teacher model and a student model, the teacher model is used to encode speech data without data loss and generate a learning target for the student model; the student model is used to encode speech data with data loss and predict a complete data representation; A first packet loss compensation module, used to obtain voice data and use the target generative adversarial network to perform packet loss compensation on the voice data; A first voice frame reconstruction module is used to reconstruct, in the process of packet loss compensation, a first voice frame with data loss in the voice data by using a second voice frame sorted before the first voice frame in the target generative adversarial network; wherein the second voice frame is a voice frame other than the voice frame with data loss; The first mask setting module can be used to set a mask for the voice frame in the voice data of the processing unit to identify whether there is data loss.
10. A device for voice communication, characterized in that: The device comprises: A second packet loss compensation module is used to obtain voice data during a voice call, and use a pre-trained target generative adversarial network to perform packet loss compensation on the voice data; wherein the target generative adversarial network has a generator, and the generator has an encoder; the encoder has a teacher model and a student model, the teacher model is used to encode voice data without data loss and generate a learning target of the student model; the student model is used to encode voice data with data loss and predict a complete data representation; A second voice frame reconstruction module is used to reconstruct, in the process of packet loss compensation, a first voice frame with data loss in the voice data by using a second voice frame sorted before the first voice frame in the target generative adversarial network; wherein the second voice frame is a voice frame other than the voice frame with data loss; The second mask setting module can be used to set a mask for the voice frame in the voice data to identify whether there is data loss.
11. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the method for voice packet loss compensation as described in any one of claims 1 to 7 is implemented, or the method for voice call as described in claim 8 is implemented.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for voice packet loss compensation according to any one of claims 1 to 7 is implemented, or the method for voice call according to claim 8 is implemented.
Citation Information
Patent Citations
Method and apparatus for packet loss concealment using generative adversarial network
US20190051310A1