Voice stream packet loss compensation method and its device, equipment, medium, and product

By identifying missing voice frames in the voice stream and generating subsequent voice frames using the conditions of the vocoder and the recurrent network, the problem of call quality degradation caused by voice packet loss in network voice calls is solved, and high-quality voice stream recovery is achieved.

CN115171707BActive Publication Date: 2025-05-09BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210804024.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-05-09
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

In online voice call scenarios, voice packet loss often leads to a decline in call quality, especially in scenarios where packet loss exceeds 40ms, the synthesized audio is prone to problems such as duplicate sounds or mechanical sounds.

Method used

By determining the current speech frame of the subsequent speech frame missing in the speech stream, the acoustic characteristics of the speech frame sequence are extracted, and the global feature information and local feature information are extracted respectively using the preset conditional network and cyclic network in the vocoder, the comprehensive feature information is constructed, and the subsequent speech frame is generated based on this information.

Benefits of technology

It effectively avoids the repetitive sound, mechanical sound and other phenomena of voice after packet loss compensation, and improves the restore degree and subjective quality score of the voice stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171707B_ABST
    Figure CN115171707B_ABST
Patent Text Reader

Abstract

The present application relates to a method for compensating for packet loss in a voice stream and its apparatus, equipment, medium, and product. The method includes: determining a current voice frame in a voice stream that lacks subsequent voice frames; obtaining a voice frame sequence including the current voice frame from the voice stream, and extracting acoustic features of the voice frame sequence; using a conditional network in a preset vocoder to extract global feature information and local feature information of the acoustic features, respectively, to construct comprehensive feature information; using a cyclic network in the vocoder, taking the global feature information as a reference, and generating the subsequent voice frames according to the comprehensive feature information. The present application can realize packet loss compensation for a voice stream, and generate missing subsequent voice frames for the voice stream. The generated subsequent voice frames have a high degree of restoration, and can effectively avoid the existence of repeated sounds, mechanical sounds, and other phenomena in the voice after packet loss compensation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio transmission technology, and in particular to a voice stream packet loss compensation method and its device, equipment, medium, and product. Background Art

[0002] In the scenario of network voice calls, voice packet loss often occurs during voice transmission, affecting the call quality and needs to be prevented or handled through technical means.

[0003] Traditional packet loss compensation methods use signal processing technology to find the segment that best matches the waveform before packet loss, use waveform replacement, linear predictive coding or other difference methods to predict the lost segment, and then use overlap-addition, fade-in and fade-out methods to fuse it with the received voice to complete signal reconstruction.

[0004] When using traditional packet loss compensation methods, for voice data, when the packet loss exceeds 40ms, the synthesized audio will have repeated sounds, mechanical sounds, etc., so it is necessary to find other effective methods. Summary of the invention

[0005] The purpose of the present application is to solve the above-mentioned problem and to provide a voice stream packet loss compensation method and its corresponding device, equipment, non-volatile readable storage medium, and computer program product.

[0006] According to one aspect of the present application, a method for compensating for voice stream packet loss is provided, comprising the following steps:

[0007] Determining a current speech frame in a speech stream that lacks subsequent speech frames;

[0008] Acquire a speech frame sequence including a current speech frame from a speech stream, and extract acoustic features of the speech frame sequence;

[0009] Using a conditional network in a preset vocoder, respectively extracting global feature information and local feature information of the acoustic feature to construct comprehensive feature information;

[0010] The subsequent speech frame is generated according to the comprehensive feature information by using a recurrent network in the vocoder and taking the global feature information as a reference.

[0011] According to another aspect of the present application, a voice stream packet loss compensation device is provided, comprising:

[0012] A current frame processing module configured to determine a current speech frame in a speech stream that lacks subsequent speech frames;

[0013] A sequence processing module is configured to obtain a speech frame sequence including a current speech frame from a speech stream and extract acoustic features of the speech frame sequence;

[0014] A feature construction module is configured to use a conditional network in a preset vocoder to extract global feature information and local feature information of the acoustic feature respectively to construct comprehensive feature information;

[0015] The speech frame generation module is configured to use a recurrent network in the vocoder, take the global feature information as a reference, and generate the subsequent speech frame according to the comprehensive feature information.

[0016] According to another aspect of the present application, a voice stream packet loss compensation device is provided, comprising a central processing unit and a memory, wherein the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the voice stream packet loss compensation method described in the present application.

[0017] According to another aspect of the present application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the voice stream packet loss compensation method in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the method are executed.

[0018] According to another aspect of the present application, a computer program product is provided, including a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the steps of the method described in any one of the embodiments of the present application are implemented.

[0019] Compared with the prior art, after determining that a current speech frame lacks a subsequent speech frame, the present application utilizes the acoustic features of a speech frame sequence composed of the current speech frame and its preceding multiple speech frames, and extracts the global feature information and the local feature information therein respectively with the help of a conditional network of a vocoder to construct comprehensive feature information, and then generates the subsequent speech frames according to the comprehensive feature information with the help of a recurrent network in the vocoder, with reference to the global feature information. Since the conditional network extracts significant features of different scales in the acoustic features in the speech frame sequence to obtain comprehensive feature information, it is possible to effectively obtain information such as phonemes and rhythms that change slowly in the speech information. The recurrent network can generate subsequent speech frames according to the comprehensive feature information under the condition of referring to the global feature information. The subsequent speech frames effectively inherit the phonemes, rhythms and other information of their preceding speech frames in time continuity. Therefore, the generated subsequent speech frames have a high degree of restoration, and can effectively avoid the existence of repeated sounds, mechanical sounds and the like in the speech after packet loss compensation. The compensated speech stream can obtain a higher subjective quality score. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A schematic diagram of the network architecture corresponding to the voice call environment used in this application;

[0022] Figure 2 A schematic diagram of the network architecture of the vocoder used in the voice stream packet loss compensation method of the present application;

[0023] Figure 3 A flowchart of an embodiment of a voice stream packet loss compensation method of the present application;

[0024] Figure 4 A flowchart of another embodiment of the voice stream packet loss compensation method of the present application;

[0025] Figure 5 A schematic diagram of a process of connecting subsequent voice frames to a voice stream in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of a process for obtaining acoustic features of a speech frame sequence in an embodiment of the present application;

[0027] Figure 7 A schematic diagram of a process for obtaining comprehensive feature information based on acoustic features in an embodiment of the present application;

[0028] Figure 8 A schematic diagram of a process for generating subsequent speech frames according to comprehensive feature information in an embodiment of the present application;

[0029] Fig. 9 A schematic diagram of a process of training a vocoder in an embodiment of the present application;

[0030] Fig.10 A schematic diagram of a process of training a vocoder using a mask in an embodiment of the present application;

[0031] Fig.11 This is a principle block diagram of the voice stream packet loss compensation device of the present application;

[0032] Fig.12 This is a structural diagram of a voice stream packet loss compensation device used in this application. DETAILED DESCRIPTION

[0033] The models referenced or may be referenced in this application, including traditional machine learning models or deep learning models, can be deployed on a remote server and remotely called on the client, or deployed on a client with sufficient device capabilities and directly called, unless explicitly specified. In some embodiments, when it runs on the client, its corresponding machine intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.

[0034] Those skilled in the art should be aware that, although the various methods of the present application are described based on the same concept and thus present commonality to each other, unless otherwise specified, these methods can be independently executed. Similarly, for each embodiment disclosed in the present application, they are all proposed based on the same inventive concept, therefore, concepts with the same expression, and concepts that are appropriately changed for convenience despite different expressions, should be understood as equivalent.

[0035] See also Figure 1 The network architecture adopted in the exemplary application scenario of the present application can be used to deploy a voice call service. The voice call service supports real-time voice communication. The encoding and decoding processing process of the voice stream of the voice call service can be implemented by running a computer program product obtained by any embodiment of the present application. Figure 1 The application server 81 shown can be used to support the operation of the voice call service, and the media server 82 can be used to process the decoding process of the voice stream pushed by each user to achieve relaying. The terminal devices such as the computer 83 and the mobile phone 84 are generally provided to the terminal user as clients and can be used to send or receive voice streams. In addition, when it is necessary to encode the voice stream on the terminal device, the computer program product obtained by each embodiment of the present application can also be deployed in the terminal device, so that the method of any embodiment of the present application can be applied to compensate for the packet loss of the received or sent voice stream. The application scenarios disclosed above are for illustrative purposes only. The voice stream packet loss compensation method of the present application is applicable to all scenarios where packet loss compensation for voice streams is required. For example, it can also be used in a live network service scenario to compensate for the packet loss of voice streams in live broadcast streams.

[0036] Figure 2 The network architecture of an exemplary vocoder used in the present application is shown. The vocoder includes a conditional network and a recurrent network. The conditional network is mainly used to extract features from voice data in a voice stream, and the recurrent network is mainly used to generate subsequent voice frames required for packet loss compensation for the voice data.

[0037] The conditional network mainly consists of a residual network responsible for extracting the global feature information of the speech data, and an upsampling network responsible for extracting the local feature information corresponding to multiple scales of the speech data using multiple different scaling systems. The global feature information and the local feature information are then spliced ​​into comprehensive feature information through a splicing layer to extract the deep speech information of the speech data, so that the comprehensive feature information can represent the features of the slowly changing information such as phonemes and rhythms in the speech data. The global feature information obtained by the residual network is further divided into multiple outputs to provide reference information for the process of the recurrent network processing the comprehensive feature information.

[0038] The recurrent network is implemented using RNN (Recurrent Convolutional Network), which uses two unidirectional gated recurrent units (GRU) internally. Each gated recurrent unit is responsible for processing comprehensive feature information, obtaining corresponding deep semantic information and splicing it with the global feature information output by the residual network before inputting it into the next node for processing. A classification network is set at the end of the recurrent network to implement classification mapping based on the combined information of the output processed by the gated recurrent unit and the global feature information, so as to restore subsequent speech frames.

[0039] In one embodiment, the vocoder can be modified from a basic model of WaveRNN (Wave Recurrent Neural Network) or its variant model such as SC-WaveRNN (Speaker Conditional WaveRNN). WaveRNN is a recurrent network model suitable for processing audio data in sequence form. The original design goal of WaveRNN is to maintain high-speed sequence generation. The author uses simplified models, sparsification, parallel sequence generation and other technologies to significantly improve the sequence generation speed. Its good performance can even achieve real-time speech synthesis on the CPU.

[0040] The vocoder can be trained to a converged state in advance for use. The network structure of the vocoder has reduced overall complexity, fewer parameters, and uses global feature information to provide temporary context, making it easier to be trained to convergence and making more accurate predictions.

[0041] See also Figure 3 According to one aspect of the present application, a voice stream packet loss compensation method is provided, in one embodiment of which, the method comprises the following steps:

[0042] Step S1100, determining a current voice frame in a voice stream that lacks subsequent voice frames;

[0043] During a voice call, the voice stream will pass through multiple device nodes in the entire transmission channel from the sender's terminal device to the server and then to the receiver's terminal device. Voice stream packet loss compensation can be performed in any set node and the voice stream packet loss compensation method of the present application can be applied. Specifically, for the server, it receives the voice stream submitted by the sender, and after decoding the voice stream, it can obtain the voice frames therein. By detecting whether the timestamps carried by the received voice frames are continuous within the preset time range, it can be detected whether the voice frames are lost, that is, the phenomenon of packet loss occurs, and packet loss compensation can be implemented. For terminal devices, especially the receiving terminal devices, it is the same as the server side. By detecting whether the voice frames are lost in the received voice stream, it is decided whether to implement packet loss compensation for the voice stream.

[0044] The voice stream is composed of multiple voice frames. In one embodiment, when a voice frame is detected and confirmed to be lost, the previous voice frame with a timestamp earlier than that of the voice frame can be determined as the current voice frame. For the current voice frame, the next voice frame that is sequentially consecutive to it is missing, and what needs to be compensated is this missing voice frame, that is, the missing subsequent voice frame. The current voice frame and its subsequent voice frame are continuous in the timing determined according to the timestamp. In another embodiment, in the voice stream, after the current voice frame, multiple subsequent voice frames may be lost continuously. In this case, the subsequent voice frame obtained after packet loss compensation can be used as a new current voice frame to restore the next subsequent voice frame that is sequentially consecutive to it.

[0045] Step S1200: obtaining a speech frame sequence including a current speech frame from a speech stream, respectively extracting global feature information and local feature information of the acoustic feature, and constructing comprehensive feature information;

[0046] The voice stream is generally stored in a memory buffer, and the voice frames in the voice stream can be obtained in the memory buffer. In order to form the input of the vocoder, in one embodiment, after determining that the current voice frame is missing the subsequent voice frame, based on the timestamp of the current voice frame, the current voice frame is used as the last voice frame in the time sequence, and multiple voice frames with continuous time sequence are obtained from the voice stream, and these voice frames constitute a voice frame sequence. It is not difficult to understand that in the voice frame sequence, the current voice frame is the latest voice frame in the time sequence, and the other voice frames are all voice frames whose timestamps are earlier than the timestamp of the current voice frame, and these voice frames are continuous in timestamps.

[0047] For the speech frame sequence, audio preprocessing can be performed on each speech frame therein by means of speech preprocessing means to obtain the acoustic features of each speech frame, thereby constructing the acoustic features of the entire speech frame sequence.

[0048] The acoustic features serve to describe relevant information of relatively stable style features in a speech frame, such as phonemes, rhythms, etc., and may be any one of logarithmic Mel spectrum information, time-frequency spectrum information, and CQT filter information.

[0049] Those skilled in the art understand that the above various acoustic features can be encoded using corresponding algorithms. In the encoding process, the speech signal is first subjected to conventional processing such as pre-emphasis, framing and windowing, and then analyzed in the time domain or frequency domain, that is, speech signal analysis is achieved. The purpose of pre-emphasis is to enhance the high-frequency part of the speech signal and smooth the spectrum; generally, pre-emphasis is achieved through a first-order high-pass filter. Before analyzing the speech signal, it is also necessary to frame it. Usually, the length of each frame of the speech signal is set to 20ms. Considering the frame shift factor, there can be 10ms overlap between adjacent frames. In order to achieve framing, it can be achieved by windowing the speech signal. Different window selections will affect the results of speech signal analysis. More commonly, the window function corresponding to the Hamming window is used to implement the windowing operation.

[0050] In one embodiment, for the time-frequency spectrum information, the speech data of each speech information in the time domain is pre-emphasized, framed, windowed, and short-time Fourier transformed (STFT) to the frequency domain to obtain data corresponding to the spectrogram, thereby forming the time-frequency spectrum information.

[0051] In another embodiment, the logarithmic Mel spectrum may be obtained by filtering the time-frequency spectrum information with a Mel-scale filter bank and then taking the logarithm.

[0052] In another embodiment, for the CQT filtering information, CQT (Constant Q Transform), i.e., constant Q transform, refers to a filter group whose center frequencies are distributed according to an exponential law, whose filtering bandwidths are different, but whose center frequency to bandwidth ratio is a constant Q. It is different from Fourier transform in that the horizontal axis frequency of its spectrum is not linear, but based on log2, and the filter window length can be changed according to the different spectral line frequencies to obtain better performance.

[0053] Any of the above acoustic features can be used as the input of the vocoder of the present application. In order to facilitate the processing of the vocoder, in one embodiment, the acoustic features can be constructed according to a certain preset format. For example, the acoustic features corresponding to each speech frame are organized into a row vector, and for the entire speech frame sequence to be encoded, the row vectors of each speech frame are spliced ​​together longitudinally in time sequence to obtain a two-dimensional matrix as the acoustic feature of the entire speech frame sequence.

[0054] Step S1300: using a conditional network in a preset vocoder to extract comprehensive feature information of acoustic features, including global feature information and local feature information of acoustic features;

[0055] The vocoder is pre-trained to convergence so that it acquires the ability to generate the last speech frame in the speech frame sequence, ie, the subsequent speech frame of the current speech frame, according to the acoustic features of the input speech frame sequence.

[0056] When the acoustic features of the speech frame sequence are input into the vocoder, they enter the conditional network in the vocoder to extract the deep semantic information. Specifically, the acoustic features are divided into two paths and input into the residual network and upsampling network in the conditional network respectively.

[0057] The residual network performs residual convolution operations on the acoustic features input into it, extracts the deep semantic information in the acoustic features at all scales, and thus obtains the global feature information corresponding to the acoustic features. This global feature information is not only used in the conditional network, but will also be referenced in the recurrent network to provide reference information for the recurrent network.

[0058] The upsampling network performs feature sampling operations on the acoustic features at different scales by pre-matching scaling coefficients corresponding to multiple scales, thereby obtaining local feature information corresponding to each scale.

[0059] In the conditional network, the global feature information obtained by the residual network and the global feature information obtained by the upsampling network are concatenated into comprehensive feature information, which includes relatively stable style features such as phonemes and rhythms obtained by extracting deep semantic information from acoustic features at different scales.

[0060] Step S1400: using a recurrent network in the vocoder, taking the global feature information as a reference, and generating the subsequent speech frame according to the comprehensive feature information.

[0061] The recurrent network in the vocoder takes the comprehensive feature information output by the conditional network as input, and uses two or more gated recurrent units (GRU) to extract features from it in order to obtain feature information corresponding to the generation of subsequent speech frames. In this process, the output of each gated recurrent unit will be further spliced ​​with the global feature information output by the residual network in the conditional network as the input of the next processing node. Finally, the classification network at the end of the recurrent network is used, in which the classification mapping is fully connected based on the summary feature information composed of the feature information output by the last gated recurrent unit and the global feature information, and the subsequent speech frame is constructed through the classification mapping result. This subsequent speech frame is the next speech frame obtained by implementing packet loss compensation for the current speech frame.

[0062] According to the above embodiments, after determining that the current speech frame lacks subsequent speech frames, the present application utilizes the acoustic features of the speech frame sequence composed of the current speech frame and its previous multiple speech frames, and extracts the global feature information and local feature information therein respectively with the help of the conditional network of the vocoder to construct comprehensive feature information, and then generates the subsequent speech frames according to the comprehensive feature information with the help of the recurrent network in the vocoder, with reference to the global feature information, because the conditional network extracts the significant features of different scales in the acoustic features in the speech frame sequence to obtain the comprehensive feature information, so that the slowly changing phonemes, rhythms and other information in the speech information can be effectively obtained, and the recurrent network can generate subsequent speech frames according to the comprehensive feature information under the condition of referring to the global feature information, and the subsequent speech frames effectively inherit the phonemes, rhythms and other information of the previous speech frames that are continuous in time, so that the generated subsequent speech frames have a high degree of restoration, which can effectively avoid the existence of repeated sounds, mechanical sounds and other phenomena in the speech after packet loss compensation, and it can be expected that the compensated speech can obtain a higher subjective quality score.

[0063] Based on any of the above embodiments, please refer to Figure 4 , after the step of generating the subsequent speech frame according to the comprehensive feature information, the method further comprises:

[0064] Step S1500: Taking the generated subsequent speech frame as the new current speech frame, continue to iteratively generate new subsequent speech frames until the generated subsequent speech frames reach the maximum compensation quantity, or until all the consecutive missing subsequent speech frames are completed.

[0065] When multiple voice frames are missing continuously in the voice stream, after the vocoder implements packet loss compensation based on the current video frame to obtain the subsequent voice frame, it can also implement packet loss compensation for the next subsequent voice frame by iterating steps S1100 to S1400 of the present application.

[0066] Specifically, the subsequent voice frame recovered by the vocoder is used as a new current voice frame and added to the end of the voice frame sequence. In one embodiment, the first voice frame in the voice frame sequence can also be deleted to maintain the same length of the voice frame sequence. Then, based on the new voice frame sequence, the loop is executed from step S1100 until the next subsequent voice frame is obtained in step S1400. By extension, multiple subsequent voice frames can be continuously determined, thereby completing the packet loss compensation for multiple missing subsequent voice frames.

[0067] In one embodiment, the maximum compensation number of subsequent voice frames can be preset to control the cyclic iteration process of the vocoder. When the vocoder cyclically iterates multiple times and reaches the maximum compensation number, the iteration is terminated, so that based on the original current voice frame as the starting point, a plurality of subsequent voice frames corresponding to the maximum compensation number are generated in a time sequence. Generally, this maximum compensation number can be set to 4 to 8, which can ensure that the subtle differences generated by compensating the subsequent voice frames are not easily perceived by the human ear and maintain excellent sound quality. For example, if it is set to 6, and the time length of the voice frame is 20ms, then a subsequent voice frame of 120ms can be generated accordingly.

[0068] In another embodiment, the total amount of speech frames actually missing continuously in the speech stream may be detected and determined as the maximum compensation amount, so as to dynamically determine the number of iterations to complete all subsequent speech frames that are missing continuously.

[0069] In another embodiment, the maximum compensation quantity can be determined in association with the duration of the speech frame, and the maximum compensation quantity can be determined according to the preset compensation duration. Generally speaking, in order to make it difficult for auditory perception to perceive the subtle changes caused by the supplementary frame, the compensation duration can be set to no more than 180ms. When determining the maximum compensation quantity, 180ms is divided by the duration of the speech frame, such as 20ms, and rounded down to obtain the maximum compensation quantity, such as 180 / 20=9.

[0070] In another embodiment, the vocoder may introduce the duration of the missing voice frames in the voice stream during the control loop iteration process. If the duration is short, such as 40ms, the iteration may be terminated after two subsequent voice frames are generated. If the duration is long, such as 180ms, the iteration may be terminated after the maximum compensation quantity is limited, for example, six subsequent voice frames are generated. For the subsequent missing 60ms voice frames, a silence replacement mark is implanted to keep them silent during playback, thereby reasonably controlling the effective duration of packet loss compensation.

[0071] The various methods used in the prior art are prone to produce mechanical and repeated sounds after compensation exceeds 40ms. Through actual testing, the present application has achieved better sound quality even if 120ms of compensated speech frames are continuously generated, and there will be no mechanical and repeated sounds that affect the sound quality.

[0072] Based on any of the above embodiments, please refer to Figure 5 , after the step of generating the subsequent speech frame according to the comprehensive feature information, specifically after step S1400 or after step S1500 is iterated once or multiple times and finally executes step S1400, including:

[0073] Step S2100, splicing the generated subsequent speech frames to obtain a compensation frame sequence;

[0074] For subsequent speech frames generated by the vocoder, no matter whether it generates a single subsequent speech frame or multiple subsequent speech frames, they can be processed collectively, and these subsequent speech frames are sequentially spliced ​​according to the time stamps to construct a compensation frame sequence.

[0075] Step S2200, adjusting the volume corresponding to the compensation frame sequence so that it does not exceed the volume of the speech frame sequence;

[0076] The volume of each subsequent voice frame obtained after packet loss compensation may be different. In order to unify the volume effect, a preset compressor, specifically a volume compressor, can be called to implement volume compression control on each subsequent voice frame in the compensation frame sequence. Taking the volume in the voice frame sequence as a reference, the excessive volume in the subsequent voice frames is reduced so that the volume of the subsequent voice frames in the compensation frame sequence does not exceed the volume of the voice frames in the voice frame sequence, thereby controlling the volume of the voice frames obtained by packet loss compensation to a reasonable level and maintaining the consistency of sound quality.

[0077] Step S2300: smoothly connect the compensation frame sequence to the voice stream where the voice frame sequence is located.

[0078] By inserting the compensation frame sequence into the voice stream, packet loss compensation for the voice stream can be further realized. In order to maintain the smoothness of the compensation frame sequence obtained by the vocoder after inserting it into the voice stream, the compensation frame sequence can be smoothly inserted into the voice stream by fading in and out.

[0079] In one embodiment, after the compensation frame sequence is connected to the voice stream, it is controlled to start fading out at 20ms, that is, it starts fading out from the second subsequent voice frame until it is completely silent 20ms after the packet loss ends or 120ms after the packet loss starts. In addition, the voice frames in the voice stream after the compensation frame sequence are controlled to fade in within a 20ms time window. The above time settings corresponding to fade-out and fade-in can be flexibly adjusted according to actual needs and are not limited to the above examples.

[0080] It can be understood from the above embodiments that the subsequent voice frames restored by the vocoder through packet loss compensation can be smoothly connected to the original voice stream, so that the original voice stream remains smooth in hearing and obtains good sound quality.

[0081] Based on any of the above embodiments, please refer to Figure 6 , the step of obtaining a speech frame sequence including a current speech frame from a speech stream and extracting acoustic features of the speech frame sequence comprises:

[0082] Step S1210: acquiring a current voice frame and its previous voice frame that is continuous in time sequence from the voice stream according to a preset duration to form a voice frame sequence;

[0083] When it is detected that the voice frame after the current voice frame is missing, it is necessary to determine the voice frame sequence including the current voice frame. To this end, in order to adapt to the unified specifications of the vocoder, a preset time length can be used to obtain each time-series continuous voice frame corresponding to the preset time length from the memory buffer area of ​​the voice stream, and these time-series continuous voice frames are based on the current voice frame as the last voice frame.

[0084] The preset duration can be set as needed. In one embodiment, it can be between 200ms and 400ms, for example, 300ms. Relatively speaking, such a setting can obtain enough speech frames and provide enough audio information to effectively help the vocoder generate subsequent speech frames.

[0085] Step S1220, performing short-time Fourier transform on the speech frame sequence to obtain spectrum information;

[0086] The speech frame sequence already contains sufficient speech frame data, so a short-term Fourier transform (STFT) can be performed on the speech frame sequence to convert it from time domain information to frequency domain information, thereby obtaining spectrum information corresponding to the speech frame sequence.

[0087] In one embodiment, the short-time Fourier transform is performed using the following formula:

[0088]

[0089] Among them, x(t) is the input signal, that is, the speech frame sequence, w(t) is the window function, and the Hamming window function is recommended. The formula indicates that STFT{x(t)}(τ,ω) is the short-time Fourier transform of x(t)w(t-τ).

[0090] Step S1230: Apply a Mel filter bank to convert the frequency spectrum information into a logarithmic Mel spectrum as an acoustic feature of the speech frame sequence.

[0091] The spectral information of the speech frame sequence is linearly scaled. Because the height of the sound heard by the human ear is not linearly proportional to the frequency of the sound, the Mel frequency scale is more in line with the auditory characteristics of the human ear. Therefore, the Mel filter bank is further applied to convert the spectral information into the Mel scale, and then the logarithm is calculated to obtain the corresponding logarithmic Mel spectrum, so as to obtain the contour information in the acoustic feature. When performing logarithmic transformation, the following formula can be applied:

[0092]

[0093] Among them, S represents the spectrum information obtained by Mel scale conversion, Represents the logarithmic Mel spectrum obtained after conversion.

[0094] According to the above embodiments, the speech frame sequence extracted from the speech stream is converted from the time domain to the frequency domain for representation, and the corresponding logarithmic Mel spectrum is obtained by taking the logarithm. The logarithmic Mel spectrum can effectively represent the contour features in the speech frame sequence, and thus can be used as the acoustic features of the speech frame sequence to achieve a preliminary and effective representation of the audio feature information in the speech frame sequence.

[0095] Based on any of the above embodiments, please refer to Figure 7 The method of using a conditional network in a preset vocoder to extract comprehensive feature information of acoustic features includes:

[0096] Step S1310: extracting the acoustic features based on the residual network in the conditional network to obtain the global feature information;

[0097] In this embodiment, SC-WaveRNN can be used as a prototype to obtain the network structure of the vocoder of this application. Compared with the prototype network provided by the original author, Figure 2 As shown, the Speaker Encoder in the prototype network can be omitted. Of course, in another embodiment, the encoder can also be used.

[0098] From the perspective of implementing speech synthesis in this application, the speaker encoder in the prototype network of SC-WaveRNN is not necessary. The speaker encoder is an important contribution of the SC-WaveRNN paper. The author uses PESQ (Perceptual evaluation of speech quality, objective speech quality evaluation) to measure that the speaker encoder has positive benefits in all situations; this application uses the same indicator to measure, and the contribution of the speaker encoder is not obvious for the task of performing packet loss compensation in this application. The reason is that SC-WaveRNN is based on TTS (Text to Speech, from text to speech), the model input contains a complete Mel spectrum, and it is more important for the speaker encoder to map the Mel spectrum to speaker features; for the application scenario of packet loss hiding in this application, the speaker features can only affect the first frame of the compensated speech, and the added speaker encoder includes LSTM (Long Short-Term Memory, long short-term memory network), which has a high computational complexity and the benefits are not obvious. Therefore, those skilled in the art can adopt or not adopt the speaker encoder according to the principles disclosed herein to realize the configuration of the vocoder.

[0099] The acoustic features of the speech frame sequence obtained above, which are used to generate subsequent speech frames, are input into the conditional network of the vocoder, one of which is input into the residual network of the conditional network. The residual network is responsible for performing residual convolution operations on the acoustic features, extracting deep semantic information from the speech frame sequence at the global scale, thereby obtaining corresponding global feature information and realizing global representation of the acoustic features of the speech frame sequence.

[0100] Step S1320: performing multi-scale sampling on the acoustic features based on the upsampling network in the conditional network to obtain the local feature information;

[0101] The acoustic features of the speech frame sequence are input from the second path to the upsampling network in the conditional network. The upsampling network pre-sets multiple scaling scales, for example, three scaling scales. Deep semantic information is extracted from the acoustic features at these different scaling scales, and the information granularity is continuously refined to obtain the corresponding local feature information at different information granularities, thereby realizing the local representation of the acoustic features of the speech frame sequence.

[0102] Step S1330: Based on the concatenation layer in the conditional network, the global feature information and the local feature information are concatenated to obtain the comprehensive feature information.

[0103] Finally, the splicing layer set in the conditional network performs feature splicing on the global feature information and local feature information obtained by the residual network to construct comprehensive feature information. The comprehensive feature information includes both the global information of the acoustic features and the local information of the acoustic features at different finer scales. It can comprehensively and completely characterize the important features of the acoustic features and help guide the recurrent network to generate effective subsequent speech frames.

[0104] According to the above embodiments, it can be understood that the vocoder realizes effective feature representation of the speech frame sequence by integrating the important features of acoustic features under global and local conditions, which is the basis for generating subsequent speech frames. In addition, by adopting a vocoder with an optimal network structure, the working efficiency of the vocoder can be improved and good benefits can be obtained.

[0105] Based on any of the above embodiments, please refer to Figure 8 , the adopting of the recurrent network in the vocoder, taking the global feature information as a reference, and generating the subsequent speech frame according to the comprehensive feature information, comprises:

[0106] Step S1410, after the comprehensive feature information is fully connected, multiple gated recurrent units refer to the global feature information for context sorting and then fully connected and output to obtain predicted feature information;

[0107] First, the comprehensive feature information output from the conditional network is fully connected through the first fully connected layer in the recurrent network to further achieve feature integration.

[0108] Then, the fully connected comprehensive feature information is first input into the first gated recurrent unit for feature extraction to select important features and obtain the first gated feature information. The first gated feature information is further spliced ​​with the global feature information obtained by the residual network and then input into the second gated recurrent unit. The second gated recurrent unit similarly extracts the feature information input therein to obtain the second gated feature information, and the second gated feature information is similarly spliced ​​with the global feature information obtained by the residual network and then output.

[0109] In one embodiment, the output after the second gated feature information is spliced ​​with the global feature information can be used as the prediction feature information. In another embodiment, the feature information obtained by splicing the second gated feature information with the global feature information is further fully connected, and after the full connection, it is further spliced ​​with the global feature information obtained by the residual network to obtain the prediction feature information. In each of the above steps, the global feature information obtained by the residual network is continuously referenced to provide contextual references, which helps to accurately extract important features in the acoustic features and make the subsequent speech frames generated by the recurrent network more effective.

[0110] Step S1420: classify and map the predicted feature information based on a classification network to obtain subsequent speech frames.

[0111] The predicted feature information is input into a preset classification network of each cyclic network, and after classification mapping by the classification network, the probability of each bit position required for constructing a subsequent speech frame is determined, thereby constructing a subsequent speech frame.

[0112] In one embodiment, in the classification network, in the process of constructing subsequent speech frames according to the predicted feature information, the following temperature coefficient-based formula is applied to perform audio sampling:

[0113]

[0114] Where T is the sampling temperature, y i is the predicted label, P i is the probability of the i-th bit of the subsequent speech frame.

[0115] It can be understood from the above embodiments that the recurrent network can effectively generate subsequent speech frames of the current speech frame according to the output of the conditional network, under the guidance of the global features and local features of the acoustic features, to achieve effective packet loss compensation for the speech stream.

[0116] Based on any of the above embodiments, please refer to Fig. 9, before the step of determining the current voice frame in the voice stream that lacks subsequent voice frames, the method comprises:

[0117] Step S0100, using the first type of training samples in the data set to perform the first stage training on the vocoder, and training the vocoder to a convergence state;

[0118] Two types of training samples are prepared, namely, first type of training samples and second type of training samples. The two types of training samples can be stored in the same data set or in different data sets.

[0119] The first type of training samples are mainly used for pre-training of the vocoder, and the second type of training samples are mainly used for fine-tuning of the vocoder. Therefore, the first type of training samples can use materials with appropriately relaxed environmental noise, and the second type of training samples can use materials with clearer foreground speech.

[0120] In one embodiment, the first type of training samples can use a public data set. A public data set selected in the actual test training of this application is a speech data set containing multiple languages ​​and is composed of audio data provided by tens of thousands of contributors, where each speech data can be used as the first type of training samples.

[0121] In one embodiment, the second type of training samples can collect online user data by themselves. The online user data used in the actual test training of this application contains audio data sampling fragments of tens of thousands of online users. After the background noise of the original sampling fragments is eliminated, the pure human voice fragments are intercepted through Voice Activity Detection (VAD), and finally a 15-30s training sample is formed.

[0122] Both the first type of training samples and the second type of training samples may be pre-processed in advance to determine each speech frame therein.

[0123] When implementing the first stage training of the vocoder, first call a single first-category training sample from the data set to perform an iterative training to obtain the acoustic features of the speech frame sequence therein. The length of the speech frame sequence can be in units of 20ms. The acoustic features are input into the vocoder, and an activation function is generated through the conditional network. Comprehensive feature information representing information such as phonemes and rhythm is extracted from the acoustic features, and then input into the recurrent network to predict subsequent speech frames. The next speech frame of the last speech frame in the speech frame sequence is used as a supervisory label to calculate the loss value of the subsequent speech frame relative to the next speech frame. Then, based on the loss value, it is decided whether the vocoder has converged. If not, back propagation is performed on the vocoder based on the loss value, and the weight parameters of its conditional network and recurrent network are updated with gradients. The next first-category training sample is called from the data set to continue iterative training of the vocoder until the vocoder converges.

[0124] It can be seen that during the vocoder pre-training process, the last speech frame in the speech frame sequence sampled in the first type of training samples and its subsequent speech frame in time sequence are used as supervision labels for subsequent speech frames generated based on the speech frame sequence, and are used to calculate the loss values ​​of subsequent speech frames, thereby achieving effective supervision of the vocoder pre-training process.

[0125] Step S0200, using the second type of training samples in the data set to perform the second stage training on the vocoder, and training the vocoder to a convergence state;

[0126] When implementing the second stage training of the vocoder, first call a single second-category training sample from the data set to perform an iterative training to obtain the acoustic features of the speech frame sequence therein. The length of the speech frame sequence can be in units of 20ms. The acoustic features are input into the vocoder, and an activation function is generated through the conditional network. Comprehensive feature information representing information such as phonemes and rhythm is extracted from the acoustic features, and then input into the recurrent network to predict subsequent speech frames. The next speech frame of the last speech frame in the speech frame sequence is used as a supervisory label to calculate the loss value of the subsequent speech frame relative to the next speech frame. Then, based on the loss value, it is decided whether the vocoder has converged. If not, back propagation is performed on the vocoder based on the loss value, and the weight parameters of its conditional network and recurrent network are updated with gradients. The next second-category training sample is called from the data set to continue iterative training of the vocoder until the vocoder converges.

[0127] It can be seen that during the vocoder pre-training process, the last speech frame in the speech frame sequence sampled in the second type of training samples and its subsequent speech frame in time sequence are used as supervision labels for subsequent speech frames generated based on the speech frame sequence, and are used to calculate the loss values ​​of subsequent speech frames, thereby achieving effective supervision of the vocoder pre-training process.

[0128] During the second stage of training, the vocoder is trained at a smaller learning rate than in the first stage, so that the parameterization of the conditional network is more consistent with the data distribution of online users. Since the second type of training samples provided by online users are pure human voice clips, using the second type of training samples to train the vocoder at a smaller learning rate helps further improve the vocoder's ability to generate subsequent speech frames.

[0129] Step S0300, solidifying the weights of the conditional network of the vocoder, using the second type of training samples in the data set to perform the third stage training on the vocoder, training the vocoder to a convergence state, so as to adjust the weights of the recurrent network in the vocoder;

[0130] The third stage of training for the vocoder is actually to implement fine-tuning training on the vocoder after the vocoder has been pre-trained in the first two stages, so as to further adjust the weights of the recurrent network therein to ensure that it can effectively produce subsequent speech frames of a given speech frame sequence.

[0131] Therefore, before executing the third stage of training, the conditional network is considered to have reached the desired requirements and its weights are frozen so that the weights of the conditional network will not be updated by the gradient during the third stage of training. As for the weights of the recurrent network, they are still maintained as learnable weights and will be further modified during the third stage of training. Then, the third stage of training for the vocoder can be started.

[0132] When implementing the third stage training for the vocoder, first call a single second-category training sample from the data set to perform an iterative training to obtain the acoustic features of the speech frame sequence therein. The length of the speech frame sequence can be in units of 20ms. The acoustic features are input into the vocoder, and an activation function is generated through the conditional network. Comprehensive feature information representing information such as phonemes and rhythm is extracted from the acoustic features, and then input into the recurrent network to predict subsequent speech frames. The next speech frame of the last speech frame in the speech frame sequence is used as a supervisory label to calculate the loss value of the subsequent speech frame relative to the next speech frame. Then, based on the loss value, it is decided whether the vocoder has converged. If not, back propagation is performed on the vocoder based on the loss value, and the weight parameters of its conditional network and recurrent network are updated with gradients. The next second-category training sample is called from the data set to continue iterative training of the vocoder until the vocoder converges.

[0133] It can be seen that during the vocoder pre-training process, the last speech frame in the speech frame sequence sampled in the second type of training samples and its subsequent speech frame in time sequence are used as supervision labels for subsequent speech frames generated based on the speech frame sequence, and are used to calculate the loss values ​​of subsequent speech frames, thereby achieving effective supervision of the vocoder pre-training process.

[0134] During the third stage of training, the vocoder is trained at a smaller learning rate than in the first stage, so that the parameterization of the conditional network is more consistent with the data distribution of online users. Since the second type of training samples provided by online users are pure human voice clips, using the second type of training samples to train the vocoder at a smaller learning rate helps further improve the vocoder's ability to generate subsequent speech frames.

[0135] After the third stage of training, the weights of the conditional network remain unchanged, and the weights of the recurrent network are constantly modified, eventually reaching a convergence state, which terminates the entire training process of the vocoder.

[0136] According to the above embodiments, it is not difficult to understand that the acoustic device of the present application undergoes multi-stage training, wherein the first type of training samples are used for pre-training in the first stage, the second type of training samples are used in the second stage to improve the pre-training effect with a smaller learning rate, and the second type of training samples are used in the third stage to improve the ability of the recurrent network to generate subsequent speech frames while solidifying the weights of the conditional network, and finally to achieve comprehensive training, so that the obtained vocoder has the ability to generate effective subsequent speech frames for a given sequence of speech frames. The vocoder obtained according to the above training process has a low parameter amount and a low number of floating-point operations per second, and can realize real-time inference on the mobile terminal, which is particularly suitable for deployment in mobile terminals such as mobile phones and computers.

[0137] In another embodiment, according to the idea of ​​generative adversarial training, the conditional network can be used as a generator, the recurrent network can be used as a discriminator, the conditional network and the recurrent network can be constructed as a generative adversarial model for training, and the first type of training samples and the second type of training samples in the data set can be used to train the generative adversarial model to a convergence state. Similarly, the training of the vocoder can be completed.

[0138] Based on any of the above embodiments, please refer to Fig.10 The third stage of training the vocoder using the second type of training samples in the data set includes:

[0139] Step S0210, replacing the acoustic features of a preset number of subsequent speech frames whose time sequence is continuous with the last speech frame in the speech frame sequence sampled from each second type of training sample with mask representation;

[0140] During the third stage of training, the vocoder can be trained accordingly according to the number of subsequent speech frames that the vocoder is expected to predict and generate for the speech frame sequence, that is, the maximum compensation number. To this end, in each iterative training, a speech frame sequence is obtained for the second type of training samples called, and for the continuous multiple speech frames after the last speech frame of the speech frame sequence, the specific number is determined according to the preset maximum compensation number, and the acoustic features of these speech frames are replaced with mask representations. For example, the maximum compensation number is determined as 6 speech frames according to the duration of 120ms, and the acoustic features of the 6 subsequent speech frames that are consecutive in time sequence after the speech frame sequence are replaced with mask representations. The mask representation method can be, for example, replacing all the feature values ​​of the corresponding acoustic features with values ​​1, 0.5, etc., which can be flexibly set.

[0141] Step S0220: the vocoder iteratively generates subsequent speech frames corresponding to the plurality of subsequent speech frames based on the speech frame sequence of the training sample;

[0142] For the second type of training samples, the vocoder first starts to generate subsequent speech frames based on the acoustic features of the first speech frame sequence of the second type of training samples. After generating a subsequent speech frame, the subsequent speech frame is queued as the last speech frame in the speech frame sequence to obtain a new speech frame sequence, and then continues to iteratively extract the acoustic features of the new speech frame sequence for generating the next subsequent speech frame. This iterative execution is performed until multiple subsequent speech frames corresponding to the maximum compensation number are generated.

[0143] Step S0230: Calculate the loss values ​​of the corresponding subsequent speech frames according to the multiple subsequent speech frames in the second type of training samples, and modify the weights of the recurrent network in the vocoder according to the loss values.

[0144] In the process of the vocoder generating multiple subsequent speech frames for each second-category training sample, for each subsequent speech frame, the vocoder uses the subsequent speech frame whose timestamp in the training sample corresponds to the subsequent speech frame as the supervisory label of the subsequent speech frame, calculates the loss value of the subsequent speech frame, and performs backpropagation on the vocoder based on the loss value, and gradient updates the weights of the recurrent network, while the conditional network does not participate in the gradient update because its weights have been frozen.

[0145] According to the above embodiments, it is not difficult to understand that by replacing the acoustic features of the subsequent speech frame in the second type of training samples with a mask representation during the third stage of training, the recurrent network can be guided to generate a subsequent speech frame corresponding to the subsequent speech frame, and the corresponding multiple subsequent speech frames can be generated according to the preset maximum compensation number through connection iteration, so that the vocoder can learn the ability to continuously compensate for multiple speech frames, thereby improving the efficiency of the vocoder in performing packet loss compensation.

[0146] See also Fig.11 According to one aspect of the present application, a speech stream packet loss compensation device is provided. In one embodiment, the device includes a current frame processing module 1100, a sequence processing module 1200, a feature construction module 1300, and a speech frame generation module 1400, wherein the current frame processing module 1100 is configured to determine a current speech frame in which a subsequent speech frame is missing in the speech stream; the sequence processing module 1200 is configured to obtain a speech frame sequence including the current speech frame from the speech stream and extract acoustic features of the speech frame sequence; the feature construction module 1300 is configured to use a conditional network in a preset vocoder to extract global feature information and local feature information of the acoustic features respectively, and construct them into comprehensive feature information; the speech frame generation module 1400 is configured to use a recurrent network in the vocoder to generate the subsequent speech frame according to the comprehensive feature information with reference to the global feature information.

[0147] On the basis of any of the above embodiments, after the speech frame generation module 1400, it includes: an iterative decision module, which is configured to use the generated subsequent speech frame as the new current speech frame, and continue to iteratively generate new subsequent speech frames until the generated subsequent speech frames reach the maximum compensation quantity, or until all the consecutive missing subsequent speech frames are completed.

[0148] On the basis of any of the above embodiments, after the voice frame generation module 1400, it includes: a frame splicing module, which is configured to splice the generated subsequent voice frames to obtain a compensation frame sequence; a volume control module, which is configured to adjust the volume corresponding to the compensation frame sequence so that it does not exceed the volume of the voice frame sequence; and a compensation access module, which is configured to smoothly access the compensation frame sequence to the voice stream where the voice frame sequence is located.

[0149] Based on any of the above embodiments, the sequence processing module 1200 includes: a sequence acquisition unit, configured to acquire a current speech frame and its previous speech frame that is continuous in time from a speech stream according to a preset duration to form a speech frame sequence; a spectrum conversion unit, configured to perform a short-time Fourier transform on the speech frame sequence to obtain spectrum information; and a spectrum conversion unit, configured to apply a Mel filter group to convert the spectrum information into a logarithmic Mel spectrum as an acoustic feature of the speech frame sequence.

[0150] Based on any of the above embodiments, the feature construction module 1300 includes: a global extraction unit, which is configured to extract the acoustic features based on the residual network in the conditional network to obtain the global feature information; a local extraction unit, which is configured to perform multi-scale sampling on the acoustic features based on the upsampling network in the conditional network to obtain the local feature information; and a feature splicing unit, which is configured to perform feature splicing on the global feature information and the local feature information based on the splicing layer in the conditional network to obtain the comprehensive feature information.

[0151] Based on any of the above embodiments, the speech frame generation module 1400 includes: a prediction execution unit, which is configured to fully connect the comprehensive feature information, and then fully connect and output it after context sorting with reference to the global feature information through multiple gated loop units to obtain predicted feature information; a speech frame generation unit, which is configured to classify and map the predicted feature information based on a classification network to obtain subsequent speech frames.

[0152] On the basis of any of the above embodiments, prior to the current frame processing module 1100, it includes: a first training module, configured to use the first type of training samples in the data set to implement the first stage training on the vocoder, and train the vocoder to a convergence state; a second training module, configured to use the second type of training samples in the data set to implement the second stage training on the vocoder, and train the vocoder to a convergence state; a third training module, configured to solidify the weight of the conditional network of the vocoder, and use the second type of training samples in the data set to implement the third stage training on the vocoder, and train the vocoder to a convergence state to adjust the weight of the recurrent network in the vocoder; wherein the training samples are audio data, including a plurality of time-sequentially continuous speech frames, the time-sequentially continuous speech frames are used to calculate the loss value of the subsequent speech frames corresponding to the time-sequentially continuous speech frames generated by the vocoder, and the second type of training samples are audio data corresponding to pure human voice segments.

[0153] Based on any of the above embodiments, the second training module includes: a mask representation unit, which is configured to replace the acoustic features of a preset number of subsequent speech frames that are consecutive in time sequence from the last speech frame in the speech frame sequence sampled from each second-category training sample with a mask representation; an iterative generation unit, which is configured to iteratively generate subsequent speech frames corresponding to the multiple subsequent speech frames based on the speech frame sequence of the training sample by the vocoder; and a weight correction unit, which is configured to calculate the loss values ​​of the corresponding subsequent speech frames based on the multiple subsequent speech frames in the second-category training sample, and correct the weights of the recurrent network in the vocoder according to the loss values.

[0154] Another embodiment of the present application also provides a voice stream packet loss compensation device. Fig.12 As shown, a schematic diagram of the internal structure of a voice stream packet loss compensation device. The voice stream packet loss compensation device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable non-volatile readable storage medium of the voice stream packet loss compensation device stores an operating system, a database, and computer-readable instructions. The database may store an information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a voice stream packet loss compensation method.

[0155] The processor of the voice stream packet loss compensation device is used to provide computing and control capabilities to support the operation of the entire voice stream packet loss compensation device. The memory of the voice stream packet loss compensation device may store computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor may execute the voice stream packet loss compensation method of the present application. The network interface of the voice stream packet loss compensation device is used to connect and communicate with a terminal.

[0156] Those skilled in the art will understand that Fig.12The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the voice stream packet loss compensation device to which the scheme of the present application is applied. The specific voice stream packet loss compensation device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0157] In this embodiment, the processor is used to execute Fig.11 The memory stores the program code and various data required to execute the above modules or submodules. The network interface is used to realize data transmission between user terminals or servers. The non-volatile readable storage medium in this embodiment stores the program code and data required to execute all modules in the voice stream packet loss compensation device of this application, and the server can call the program code and data of the server to execute the functions of all modules.

[0158] The present application also provides a non-volatile readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the voice stream packet loss compensation method of any embodiment of the present application.

[0159] The present application also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any embodiment of the present application when executed by one or more processors.

[0160] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0161] In summary, the present application can realize packet loss compensation for voice streams and generate missing subsequent voice frames for the voice stream. The generated subsequent voice frames have a high degree of restoration and can effectively avoid the existence of repeated sounds, mechanical sounds, etc. in the voice after packet loss compensation. The compensated voice stream can obtain a higher subjective quality score.

Claims

1. A voice stream packet loss compensation method, characterized in that: include: Determining a current speech frame in a speech stream that lacks subsequent speech frames; Acquire a speech frame sequence including a current speech frame from a speech stream, and extract acoustic features of the speech frame sequence; Using a conditional network in a preset vocoder, respectively extracting global feature information and local feature information of the acoustic feature to construct comprehensive feature information; The subsequent speech frame is generated according to the comprehensive feature information by using a recurrent network in the vocoder and taking the global feature information as a reference.

2. The voice stream packet loss compensation method according to claim 1, characterized in that: After the step of generating the subsequent speech frame according to the comprehensive feature information, the method further comprises: The generated subsequent speech frame is used as a new current speech frame, and new subsequent speech frames are continuously iterated and generated until the generated subsequent speech frames reach a maximum compensation quantity, or until all the consecutively missing subsequent speech frames are completed.

3. The voice stream packet loss compensation method according to claim 1 or 2, characterized in that: After the step of generating the subsequent speech frame according to the comprehensive feature information, the method further comprises: splicing the generated subsequent speech frames to obtain a compensation frame sequence; Adjusting the volume corresponding to the compensation frame sequence so that it does not exceed the volume of the speech frame sequence; The compensation frame sequence is smoothly connected to the voice stream where the voice frame sequence is located.

4. The voice stream packet loss compensation method according to claim 1, characterized in that: The step of obtaining a speech frame sequence including a current speech frame from a speech stream and extracting acoustic features of the speech frame sequence includes: Acquire a current voice frame and its previous voice frame that is continuous in time sequence from the voice stream according to a preset duration to form a voice frame sequence; Performing short-time Fourier transform on the speech frame sequence to obtain spectrum information; A Mel filter bank is applied to convert the frequency spectrum information into a logarithmic Mel spectrum as the acoustic feature of the speech frame sequence.

5. The voice stream packet loss compensation method according to claim 1, characterized in that: The method of using a conditional network in a preset vocoder to extract comprehensive feature information of acoustic features includes: Extracting the acoustic features based on the residual network in the conditional network to obtain the global feature information; Performing multi-scale sampling on the acoustic features based on an upsampling network in the conditional network to obtain the local feature information; The global feature information and the local feature information are feature spliced ​​based on the splicing layer in the conditional network to obtain the comprehensive feature information.

6. The voice stream packet loss compensation method according to claim 1, characterized in that: The method of using a recurrent network in the vocoder, taking the global feature information as a reference, and generating the subsequent speech frame according to the comprehensive feature information comprises: After the comprehensive feature information is fully connected, it is contextually sorted by multiple gated recurrent units with reference to the global feature information and then fully connected and output to obtain predicted feature information; The predicted feature information is classified and mapped based on a classification network to obtain subsequent speech frames.

7. The voice stream packet loss compensation method according to claim 1, characterized in that: Before the step of determining the current voice frame in the voice stream that lacks subsequent voice frames, the method includes: The first stage of training is performed on the vocoder using the first type of training samples in the data set, and the vocoder is trained to a convergence state; The second stage of training is performed on the vocoder using the second type of training samples in the data set, and the vocoder is trained to a convergence state; Solidifying the weight of the conditional network of the vocoder, performing the third stage training on the vocoder using the second type of training samples in the data set, training the vocoder to a convergence state, and adjusting the weight of the recurrent network in the vocoder; Among them, the training samples are audio data, including multiple temporally continuous speech frames, and the temporally consecutive speech frames are used to calculate the loss values ​​of subsequent speech frames generated by the vocoder corresponding to the temporally consecutive speech frames. The second type of training samples are audio data corresponding to pure human voice fragments.

8. The voice stream packet loss compensation method according to claim 7, characterized in that: The third stage of training the vocoder using the second type of training samples in the data set includes: Replacing the acoustic features of a preset number of subsequent speech frames that are sequentially consecutive to the last speech frame in the speech frame sequence sampled from each second type of training sample with mask representations; The vocoder iteratively generates subsequent speech frames corresponding to the plurality of subsequent speech frames based on the speech frame sequence of the training sample; The loss values ​​of the corresponding subsequent speech frames are calculated according to the multiple subsequent speech frames in the second type of training samples, and the weights of the recurrent network in the vocoder are modified according to the loss values.

9. A voice stream packet loss compensation device, characterized in that: include: A current frame processing module configured to determine a current speech frame in a speech stream that lacks subsequent speech frames; A sequence processing module is configured to obtain a speech frame sequence including a current speech frame from a speech stream and extract acoustic features of the speech frame sequence; A feature construction module is configured to use a conditional network in a preset vocoder to extract global feature information and local feature information of the acoustic feature respectively to construct comprehensive feature information; The speech frame generation module is configured to use a recurrent network in the vocoder, take the global feature information as a reference, and generate the subsequent speech frame according to the comprehensive feature information.

10. A voice stream packet loss compensation device, comprising a central processing unit and a memory, characterized in that: The central processor is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 8.

11. A non-volatile readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 8 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

12. A computer program product, characterized in that The method comprises a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Behavior recognition method based on weak supervised learning video segmentation

    CN112861758A

  • Audio packet loss compensation processing method and device and electronic equipment

    CN113035205A