Intelligent Packet Loss Resistant Voice Coding Method and Terminal for VOIP Internet Telephony

Through parameterized encoding and adaptive interleaving packaging based on network state and voice characteristics, the voice quality problem of VOIP network phones in complex network environments is solved, and higher robustness and voice recovery capabilities are achieved, especially in the high packet loss scenarios to effectively protect high-frequency information.

CN120236595BActive Publication Date: 2025-08-01LVSWITCHES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510703456.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

VOIP network phones face network packet loss, jitter and bandwidth fluctuations in complex network environments, resulting in a decline in voice quality, especially the loss of high-frequency information. The existing encoding methods lack real-time response capabilities and speech semantic perception capabilities.

Method used

Parameterized encoding based on real-time network state and speech signal characteristics is adopted, combined with a layered protection encoder and adaptive interleaving packaging strategy, the encoding layer bit allocation and high-frequency compensation are dynamically adjusted, and the linkage modeling of the speech spectrum structure and the dynamic behavior of the network is achieved through adaptive LPC window length control and unsound judgment threshold adjustment.

Benefits of technology

It improves the robustness and recovery ability of voice encoding in complex network environments, alleviates the problems of spectrum leakage and voice blur, and improves the intelligibility and nature of speech, especially in the scenario of high packet loss rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236595B_ABST
    Figure CN120236595B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of voice coding, and particularly to an intelligent packet-loss-resistant voice coding method and terminal for VOIP network telephones, including the following steps: Based on real-time network state parameters and the time-frequency characteristics of voice signals, perform parametric coding on the original voice frames to generate coding feature vectors; input the coding feature vectors into a hierarchical protection encoder, determine the bit allocation of the basic coding layer according to the fundamental frequency trajectory parameters, generate differential parameters of the enhanced coding layer based on the formant bandwidth parameters, and configure a high-frequency compensation coding layer according to the voiceless / voiced flag; combine the energy correlation characteristics of adjacent voice frames to perform adaptive interleaving encapsulation of the coding parameters, wherein the encapsulation interval of the high-frequency compensation coding layer is negatively correlated with the current network packet loss rate. The present invention effectively alleviates the problem of sudden drop in voice intelligibility caused by the loss of voiceless sounds, and improves the recovery ability and subjective listening stability of the coding structure in a harsh network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice coding, and particularly to an intelligent packet loss resistant voice coding method and a terminal for VOIP network phones. Background Art

[0002] With the wide application of VOIP (Voice over IP) network phone technology, voice communication systems increasingly rely on an asynchronous network environment based on packet switching. However, problems such as network packet loss, jitter, and bandwidth fluctuation commonly existing in the public Internet seriously affect the continuous transmission and real-time restoration of voice data packets, resulting in a decline in voice quality, reduced clarity, and even information loss. Especially under weak channel conditions, high-frequency weak information such as voiceless consonants in speech is more vulnerable to packet loss impact, thus significantly damaging the intelligibility and naturalness of speech.

[0003] Traditional voice encoders mostly adopt LPC analysis with a fixed window length and a static bit allocation strategy, lacking the ability to respond in real time to changes in network status, and it is difficult to accurately model and preferentially encode key voice features under dynamic network conditions. At the same time, common interleaving encapsulation methods mostly adopt a linear uniform insertion mode, failing to fully consider the stability and time structure of voice energy distribution, resulting in the possibility of voice backbone loss or collective damage of key frequency band information even in burst packet loss scenarios. In addition, existing high-frequency protection mechanisms often rely on fixed thresholds or redundant replication, bringing unnecessary bandwidth overhead and lacking voice semantic perception ability. Summary of the Invention

[0004] The present invention provides an intelligent packet loss resistant voice coding method and a terminal for VOIP network phones, which is a packet loss resistant voice coding method with differential protection capabilities and an intelligent encapsulation strategy to improve the robustness and voice restoration quality of VOIP network phones in complex communication environments.

[0005] The intelligent packet loss resistant voice coding method for VOIP network phones includes the following steps:

[0006] S1: Based on real-time network state parameters and the time-frequency characteristics of the voice signal, perform parametric coding on the original voice frame to generate a coding feature vector including fundamental frequency trajectory parameters, formant bandwidth parameters, and voiceless / voiced flags;

[0007] S2: Input the coding feature vector output by S1 into a hierarchical protection encoder, determine the basic coding layer bit allocation according to the fundamental frequency trajectory parameters, generate enhanced coding layer differential parameters based on the formant bandwidth parameters, and configure the high-frequency compensation coding layer according to the voiceless / voiced flags;

[0008] S3: Using the hierarchical coding structure output by S2, combined with the energy correlation features of adjacent speech frames, perform adaptive interleaved encapsulation of coding parameters, where the encapsulation interval of the high-frequency compensation coding layer is negatively correlated with the current network packet loss rate.

[0009] Optionally, in S1, a network status parameter set is collected through a sliding window, including the current packet loss rate, network jitter factor, and available bandwidth fluctuation value. At the same time, the short-time spectrum features of the speech frame are extracted, and an improved linear prediction coding analysis with an adaptive window length is used to adjust the analysis window length according to the network jitter factor to generate fundamental frequency trajectory parameters.

[0010] Optionally, in the fundamental frequency trajectory parameters, the analysis window length is shortened to 10 ms in the high-frequency jitter scenario and extended to 30 ms in the stable scenario.

[0011] Optionally, S1 further includes extracting formant bandwidth parameters based on the short-time spectrum features through the cepstral coefficients non-uniformly divided by the Mel scale, and applying bandwidth expansion compensation to the frequency bands above the preset frequency band threshold.

[0012] Optionally, the voiced / unvoiced flag in S1 is generated by combining the zero-crossing rate statistic, fundamental frequency trajectory parameters, and short-time energy gradient of the speech frame, where the decision threshold for unvoiced sound is correspondingly lowered as the current packet loss rate increases.

[0013] Optionally, S2 specifically includes:

[0014] S21: Input the coding feature vector into a hierarchical protection encoder, and calculate the bit allocation weight of the basic coding layer according to the smoothness index of the fundamental frequency trajectory parameters;

[0015] S22: Based on the formant bandwidth parameters, construct a dynamic differential coding model, extract the bandwidth change gradient between the current frame and the previous frame, and generate signed enhanced coding layer differential coding parameters when the bandwidth change gradient exceeds the preset gradient threshold;

[0016] S23: Configure the high-frequency compensation coding layer according to the confidence value of the voiced / unvoiced flag. If the flag is an unvoiced sound frame and the confidence is greater than the confidence threshold, activate the high-frequency compensation coding channel, and the compensation intensity is positively correlated with the current network packet loss rate;

[0017] S24: Perform time-frequency domain interleaving on the coding parameters of the basic coding layer, enhanced coding layer, and high-frequency compensation coding layer.

[0018] Optionally, in S21, when the change rate between adjacent frames of the fundamental frequency trajectory exceeds the change threshold, the number of bits in the basic coding layer increases accordingly (20% - 30%);

[0019] S22 also includes applying a bandwidth expansion factor to the frequency bands above 4000 Hz.

[0020] When performing time-frequency domain interleaving, the high-frequency compensation layer parameters are inserted into the encapsulated frame in a discontinuous manner, and the number of frames between adjacent high-frequency parameters is negatively correlated with the network jitter factor.

[0021] Optionally, S3 specifically includes:

[0022] S31, extract the basic coding layer coding parameters, enhanced coding layer differential coding parameters, and high-frequency compensation coding parameters in the hierarchical coding structure, and generate a time-domain energy correlation matrix in combination with the energy change gradient within a three-frame sliding window, where the row vector of the time-domain energy correlation matrix represents the energy correlation of the same frequency band parameters in adjacent frames;

[0023] S32, select the starting position of the encapsulation of the high-frequency compensation parameters according to the sparsity index of the time-domain energy correlation matrix, and when the energy gradient exceeds the energy mutation threshold, disperse the high-frequency parameters into the sequence of energy-stable frames;

[0024] S33, construct an anti-packet-loss interleaving template, and perform non-uniform encapsulation on the high-frequency compensation coding layer parameters, and the number of frames between encapsulations meets the preset interval condition.

[0025] Optionally, S3 further includes performing block interleaving coding on the basic coding layer and enhanced coding layer parameters, where the block size is positively correlated with the main diagonal strength of the time-domain energy correlation matrix, and performing time-slot cross-multiplexing on the interleaved data packet and the high-frequency compensation parameter packet.

[0026] A terminal, the terminal includes a memory and a processor, where:

[0027] The memory is used to store executable program code;

[0028] The processor is used to call the executable program code in the memory to implement the above voice coding method.

[0029] Advantages of the present invention:

[0030] In the present invention, by introducing an adaptive LPC window length control mechanism based on the network jitter factor and a voicing decision threshold adjustment mechanism based on the dynamic packet loss rate in the parametric coding stage, a linkage modeling of the voice spectrum structure and the network dynamic behavior is realized, effectively alleviating the spectrum leakage and voicing ambiguity problems generated by a fixed window length in a high-jitter environment; at the same time, combined with Mel scale division and bandwidth expansion compensation, the extraction accuracy of high-frequency formants is enhanced, providing a structurally stable and change-sensitive basic feature for subsequent coding, and improving the robustness and discrimination of anti-packet-loss preprocessing.

[0031] In the present invention, through a hierarchical protection encoder design, by combining the fundamental frequency trajectory change rate to control the bit weights of the base layer, the formant bandwidth difference to trigger the enhanced layer modeling, and the voiceless confidence-guided high-frequency compensation activation mechanism, a dynamic tilt scheduling of coding resources between speech mutation segments (such as plosives) and vulnerable high-frequency segments (such as voiceless consonants) is achieved; especially in the scenario of increasing network packet loss rate, the high-frequency compensation intensity adaptively increases with the packet loss rate, effectively alleviating the sudden drop in speech intelligibility caused by the loss of voiceless sounds, and enhancing the recovery ability and subjective listening stability of the coding structure under poor network conditions.

[0032] In the present invention, in the adaptive interleaved encapsulation stage, by introducing a three-frame sliding energy correlation matrix, accurate identification of the temporal stationarity of speech segments and high-energy mutation points is achieved, and based on the gradient trigger mechanism, high-frequency compensation parameters are scattered and inserted into the energy-stable frame segments; combined with the dynamic encapsulation interval control algorithm based on the packet loss rate, transmission redundancy can be saved in low-packet-loss scenarios, and strong protection can be achieved in high-packet-loss scenarios. Brief Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. Detailed Embodiments

[0035] The following will describe the present invention in detail with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specifically describing the embodiments, and are not intended to specifically limit the present invention.

[0036] It should be pointed out that in the specification, it is mentioned that "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. indicate that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. In addition, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0037] Generally, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or property in the singular sense, or can be used to describe a combination of features, structures, or properties in the plural sense. Additionally, the term "based on" can be understood to not necessarily be intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0038] As Figure 1 shown, the intelligent packet loss resistant voice coding method for VOIP network phones includes the following steps:

[0039] S1: Based on real-time network state parameters and the time-frequency characteristics of the voice signal, perform parametric coding on the original voice frame to generate a coding feature vector including fundamental frequency trajectory parameters, formant bandwidth parameters, and voiceless / voiced identification;

[0040] S2: Input the coding feature vector output by S1 into a hierarchical protection encoder, determine the bit allocation for the basic coding layer according to the fundamental frequency trajectory parameters, generate enhanced coding layer differential parameters based on the formant bandwidth parameters, and configure the high-frequency compensation coding layer according to the voiceless / voiced identification;

[0041] S3: Utilize the hierarchical coding structure output by S2, combine the energy correlation characteristics of adjacent voice frames, and perform adaptive interleaving encapsulation of the coding parameters, wherein the encapsulation interval of the high-frequency compensation coding layer is negatively correlated with the current network packet loss rate.

[0042] S1 specifically includes:

[0043] S11: Collect the network state parameter set through a sliding window, including:

[0044] Current packet loss rate ;

[0045] Network jitter factor ;

[0046] Available bandwidth fluctuation value ;

[0047] At the same time, extract the short-time spectrum characteristics of the voice frame , where is the frequency, is the frame time stamp.

[0048] Network jitter factor , defined as the root mean square value of the delay variation:

[0049] , represents the th measured end-to-end delay value, is the mean of the delay values, representing all the measured delay values;

[0050] Short-time spectrum extraction of speech frames: , is the original speech signal, is the analysis window function (Hamming window), represents the Fourier transform, represents the spectrum amplitude of the nth frame.

[0051] S12, using improved linear predictive coding (LPC) analysis with an adaptive window length, dynamically adjusts the analysis window length according to the network jitter factor , and calculates the fundamental frequency trajectory parameters . The window length adjustment rule is as follows:

[0052] ; where is the network jitter threshold, set according to the training data, with a default value of approximately 25 ms;

[0053] Linear predictive coding (LPC) analysis obtains the prediction coefficients , and solves the prediction error signal:

[0054] ; represents the order (e.g., 10 - 16 orders), represents the amplitude value of the speech signal at the nth sampling point, representing the nth sample of the time-domain discrete speech signal sequence in the current frame, is the LPC prediction value at the nth sampling point, is the LPC prediction error;

[0055] Fundamental frequency trajectory estimation (extracted using the autocorrelation function), expressed as:

[0056] ;

[0057] where represents the autocorrelation function when the lag is m, is the audio sampling rate, is the fundamental frequency estimate of the current frame, , are the minimum and maximum lags allowed for fundamental frequency detection, respectively;

[0058] Prediction coefficients ​is obtained as follows:

[0059] a. First, extract a windowed time-domain signal from the current speech frame, and then calculate the autocorrelation sequence of this signal, which represents the correlation of the signal itself at different time lags;

[0060] b. Construct a symmetric positive definite matrix with autocorrelation values as elements, and use it as the coefficient matrix of the linear equations. The solution of the equations is a set of linear prediction coefficients, which can predict the linear relationship between the current sampling point and its previous x sampling points in the way of minimum mean square error;

[0061] c. To efficiently solve this set of equations, the Levinson-Durbin recursive algorithm is adopted, which significantly reduces the computational complexity while maintaining numerical stability. The finally obtained prediction coefficients , describe the short-time spectral envelope characteristics of this frame of speech and can be used for subsequent parametric coding processing.

[0062] S13. Based on the short-time spectrum obtained in S11 , perform Mel-scale cepstral analysis to generate formant bandwidth parameters ;

[0063] Definition of Mel filter bank (non-uniform): , represents the frequency corresponding Mel-scale value;

[0064] Calculation of energy after Mel-frequency filtering: ; represents the th Mel filter's frequency response, the th Mel filter's output energy (the th frame);

[0065] Take the logarithm and then perform discrete cosine transform to obtain cepstral coefficients :

[0066] , represents the th frame's th cepstral coefficient, represents the number of Mel filters;

[0067] Formant bandwidth parameter is obtained according to cepstral difference: ;

[0068] Apply bandwidth expansion compensation to the high-frequency band:

[0069] ; where, is the th formant bandwidth (obtained by Mel cepstrum analysis), is the high-frequency bandwidth expansion compensation amount, set to 100 - 200 Hz, depending on the frequency band energy distribution, is the th formant corresponding frequency, represents the value of the bandwidth compensation of the th formant, is the preset frequency band threshold.

[0070] S14, combined with the zero-crossing rate , fundamental frequency trajectory and short-time energy gradient , to generate a dynamic voiceless / voiced flag , where the decision threshold for voicelessness is adaptively adjusted according to the packet loss rate:

[0071] ; where is the basic noise decision threshold, initially set to about -40 dB, is the current packet loss rate (percentage), represents the floor operation.

[0072] Zero-crossing rate is calculated as: ; represents the indicator function, which is 1 when the condition holds and 0 otherwise, represents the total number of sampling points included in the current speech frame.

[0073] Short-time energy and energy gradient are calculated as:

[0074] ; where is the short-time energy of the th frame (sum of signal powers within the frame), is the difference in short-time energy between adjacent frames;

[0075] Final voiceless / voiced flag is defined as:

[0076] ;

[0077] where represents voiceless, represents voiced.

[0078] S2 specifically includes:

[0079] S21, adjustment of the bit allocation weight in the basic coding layer:

[0080] S211, Fundamental frequency trajectory change rate calculation: ; where is the fundamental frequency value of the current frame, is the inter-frame time interval (10 ms);

[0081] S212, Bit allocation gain factor:

[0082] ; where , is the change threshold;

[0083] S213, Base layer bit number allocation: ; where, is the initial bit number of the base layer, is the base layer bit number after dynamic adjustment, is the bit allocation weight factor of the base coding layer, used to dynamically adjust the bit number of the base layer according to the change degree of the fundamental frequency trajectory, bit increase ratio, with a value range of 0.2 to 0.3 (indicating an increase of 20% - 30%).

[0084] S22, Differential modeling and high - frequency extension of the enhancement coding layer:

[0085] S221, Bandwidth change gradient Calculation: ;

[0086] S222, Conditions for generating differential coding parameters: ; where, , represents the preset gradient threshold;

[0087] S223, High - frequency extension compensation (for frequencies above 4000 Hz):

[0088] ; where, , represents the bandwidth extension factor. Here, the extension compensation and the bandwidth extension compensation for the high - frequency band in S13 belong to two stages and work together. The compensation in S13 is to perform high - frequency enhancement during cepstrum extraction to make the formant bandwidth parameters capture high - frequency details more accurately; the compensation in S22 here is to further amplify the coding weight for the detected high - frequency part during coding, for a stronger robust expression of high - frequency changes.

[0089] S23, Dynamic configuration of the high - frequency compensation coding layer:

[0090] S231, Confidence determination and channel activation: ; where, represents the voiceless / voiced sound identifier (0 for voiceless sound), Indicates the confidence value for voiceless sound determination, and 0.7 is the confidence threshold;

[0091] S232, Compensation intensity calculation: ; where is the basic compensation intensity, is the packet loss rate of the current frame.

[0092] S24, Time-frequency interleaved encapsulation of layer parameters:

[0093] S241, Calculation of the insertion interval of high-frequency compensation parameters: ; where is the network jitter factor (normalized), is the insertion frame interval between high-frequency parameters;

[0094] S242, Interleaving strategy:

[0095] Base layer parameters: Packed in chronological order;

[0096] Enhanced layer parameters: Arranged in reverse order of frequency bands;

[0097] High-frequency layer parameters: Inserted one per frame to form a discontinuous encapsulation layout.

[0098] S3 specifically includes:

[0099] S31, Generation of energy correlation matrix:

[0100] S311, Calculation of the energy gradient of a three-frame sliding window:

[0101] ; where is the energy (cepstral energy) of the frame in the frequency band, is the energy change gradient of the frequency band in the sliding window;

[0102] S312, Construction of the time-domain energy correlation matrix: ; where represents the energy correlation when the adjacent frames of the frequency band lag by , and each row vector

[0103] describes the energy stability of a certain frequency band in the sliding window.

[0104] S32, Dynamically selecting the encapsulation position based on the energy correlation matrix: ; where is the maximum frame lag number, taking 3, Frequency band The larger the energy sparsity within the sliding window, the more severe the fluctuation;

[0105] S322, high frequency parameter insertion strategy determination: If , then the high frequency compensation parameters are scattered and inserted into the subsequent frames, It is the energy mutation threshold, which is 8dB / frame.

[0106] S33, encapsulation interval control based on packet loss rate:

[0107] S331, dynamic interval encapsulation function: ;in, is the number of frame intervals for high frequency compensation parameters, is the current network packet loss rate (expressed as a ratio, for example, 20% is recorded as 0.2), Indicates rounding down.

[0108] S34, block interleaving coding and time slot cross multiplexing:

[0109] S341, dynamic adjustment of interleaving block size: ;in, is the minimum interleaved block size (16×16 for the base layer), is the main diagonal average of the energy correlation matrix, To adjust the coefficient, the degree to which the strength of the control block changes with stability;

[0110] S342, multiplexing structure arrangement: suppose the high frequency compensation package is , the base layer package is , then the packaging timing is: , that is, insert high-frequency packets at the 1 / 3 and 2 / 3 positions to form uniform interleaving in the time domain.

[0111] The present invention also relates to a terminal comprising a memory and a processor, wherein: the memory is used to store executable program code; the processor is used to call the executable program code in the memory to implement the above-mentioned speech encoding method.

[0112] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.

[0113] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. The intelligent packet loss resistant voice coding method for VOIP network phone is characterized in that, Including the following steps: S1: Based on real-time network state parameters and the time-frequency characteristics of the voice signal, perform parametric coding on the original voice frame to generate a coded feature vector including fundamental frequency trajectory parameters, formant bandwidth parameters, and voiceless / voiced flags; S2: Input the coded feature vector output by S1 into a hierarchical protection encoder, determine the bit allocation of the basic coding layer according to the fundamental frequency trajectory parameters, generate differential parameters of the enhanced coding layer based on the formant bandwidth parameters, and configure the high-frequency compensation coding layer according to the voiceless / voiced flags; S3: Utilize the hierarchical coding structure output by S2, combine the energy correlation characteristics of adjacent voice frames, and perform adaptive interleaving encapsulation of the coding parameters, where the encapsulation interval of the high-frequency compensation coding layer is negatively correlated with the current network packet loss rate; The specific steps of S2 include: S21: Input the coded feature vector into a hierarchical protection encoder, and calculate the bit allocation weight of the basic coding layer according to the smoothness index of the fundamental frequency trajectory parameters; S22: Based on the formant bandwidth parameters, construct a dynamic differential coding model, extract the bandwidth change gradient between the current frame and the previous frame, and generate signed differential coding parameters of the enhanced coding layer when the bandwidth change gradient exceeds a preset gradient threshold; S23: Configure the high-frequency compensation coding layer according to the confidence value of the voiceless / voiced flag. If the flag indicates a voiceless frame and the confidence is greater than the confidence threshold, activate the high-frequency compensation coding channel, and the compensation intensity is positively correlated with the current network packet loss rate; S24: Perform time-frequency domain interleaving on the coding parameters of the basic coding layer, the enhanced coding layer, and the high-frequency compensation coding layer.

2. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 1, characterized in that, In S1, a set of network state parameters is collected through a sliding window, including the current packet loss rate, network jitter factor, and available bandwidth fluctuation value. At the same time, the short-time spectrum characteristics of the voice frame are extracted, and an improved linear prediction coding analysis with an adaptive window length is adopted. The analysis window length is adjusted according to the network jitter factor to generate fundamental frequency trajectory parameters.

3. The intelligent packet loss resistant voice coding method for VOIP network phone according to claim 2, characterized in that In the fundamental frequency trajectory parameters, the analysis window length is shortened to 10 ms in a high-frequency jitter scenario and extended to 30 ms in a stable scenario.

4. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 2, characterized in that, S1 further includes extracting formant bandwidth parameters based on the short-time spectrum characteristics through Mel-scale non-uniformly divided cepstral coefficients, and applying bandwidth expansion compensation to the frequency bands above a preset frequency band threshold.

5. The intelligent packet loss resistant voice coding method for VOIP network phone according to claim 2, characterized in that The voiceless / voiced flag in S1 is generated by combining the zero-crossing rate statistic, fundamental frequency trajectory parameters, and short-time energy gradient of the voice frame. Among them, the determination threshold for voiceless is adjusted downward correspondingly with the increase of the current packet loss rate.

6. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 1, characterized in that In S21, when the change rate between adjacent frames of the fundamental frequency trajectory exceeds the change threshold, the number of bits in the basic coding layer increases accordingly; S22 also includes applying a bandwidth expansion factor to the frequency bands above 4000 Hz; When performing time-frequency domain interleaving, the parameters of the high-frequency compensation layer are inserted into the encapsulated frame in a non-continuous manner, and the number of frames between adjacent high-frequency parameters is negatively correlated with the network jitter factor.

7. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 1, characterized in that, The specific steps of S3 include: S31. Extract the base coding layer coding parameters, enhanced coding layer differential coding parameters, and high-frequency compensation coding parameters in the hierarchical coding structure, and generate a time-domain energy correlation matrix in combination with the energy change gradient within a three-frame sliding window, where the row vectors of the time-domain energy correlation matrix represent the energy correlation of the same frequency band parameters in adjacent frames; S32. Select the starting position of the encapsulation of the high-frequency compensation parameters according to the sparsity index of the time-domain energy correlation matrix, and when the energy gradient exceeds the energy mutation threshold, disperse the high-frequency parameters into the sequence of energy-stable frames; S33. Construct an anti-packet-loss interleaving template, and perform non-uniform encapsulation on the high-frequency compensation coding layer parameters, where the number of frames between encapsulations meets the preset interval condition.

8. The intelligent packet loss resistant voice coding method for VOIP network phone according to claim 7, characterized in that The S3 further includes performing block interleaving coding on the base coding layer and enhanced coding layer parameters, where the block size is positively correlated with the main diagonal strength of the time-domain energy correlation matrix, and performing time-slot cross-multiplexing on the interleaved data packets and the high-frequency compensation parameter packets.

9. A terminal, the terminal comprising a memory and a processor, characterized in that: The memory is used for storing executable program code; The processor is used for calling the executable program code in the memory to implement the voice coding method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Packet loss concealment for audio codec

    CN103688306A

  • Flexible frequency and time partitioning in perceptual transform coding of audio

    US20080312759A1