Intelligent anti-packet-loss voice coding method and terminal for VOIP (Voice Over Internet Protocol)

Through the intelligent anti-packet loss voice encoding method, combined with real-time network status and voice signal characteristics, a layered protection encoder and adaptive interleaving packaging technology are used to solve the problem of voice quality degradation in complex environments of VOIP network phones, achieving higher robustness and voice restoration quality.

CN120236595AActive Publication Date: 2025-07-01LVSWITCHES INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510703456.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

VOIP network phones suffer from network packet loss, jitter and bandwidth fluctuations in complex communication environments, especially under weak channel conditions, high-frequency and weak information are easily damaged, affecting the intelligibility and naturalness of voice.

Method used

The intelligent anti-packet loss voice encoding method is adopted to generate coded feature vectors through the linkage modeling of real-time network state parameters and the time-frequency characteristics of voice signals, and to achieve differentiated protection and intelligent packaging through layered protection encoder and adaptive interleaving packaging technology, and dynamically adjust the coding resource allocation.

Benefits of technology

It improves the robustness and voice restoration quality of VOIP network phones in complex communication environments, effectively alleviates the impact of packet loss on voice, and improves the anti-packet loss ability and voice intelligibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236595A_ABST
    Figure CN120236595A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice coding, in particular to an intelligent anti-packet-loss voice coding method and terminal for a VOIP (Voice Over Internet Protocol), and the method comprises the following steps: carrying out the parametric coding of an original voice frame based on a real-time network state parameter and a voice signal time-frequency feature, and generating a coding feature vector; inputting the coding feature vector into a hierarchical protection encoder, determining the bit allocation of a basic coding layer according to the fundamental frequency track parameter, generating an enhanced coding layer differential parameter based on the formant bandwidth parameter, and configuring a high-frequency compensation coding layer according to unvoiced and voiced sound identifiers; and in combination with the energy correlation characteristics of the adjacent voice frames, adaptive interleaving packaging of coding parameters is executed, and the packaging interval of the high-frequency compensation coding layer is in negative correlation with the current network packet loss rate. According to the method, the problem of sudden reduction of speech intelligibility caused by loss of unvoiced sound is effectively relieved, and the recovery capability and subjective hearing stability of a coding structure in a severe network are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech coding, and particularly to an intelligent packet loss resistant speech coding method and terminal for VOIP network telephones. Background Art

[0002] With the wide application of VOIP (Voice over IP) network telephone technology, voice communication systems increasingly rely on asynchronous network environments based on packet switching. However, problems such as network packet loss, jitter, and bandwidth fluctuations commonly existing in the public Internet seriously affect the continuous transmission and real-time restoration of voice data packets, resulting in degraded voice quality, reduced clarity, and even information loss. Especially under weak channel conditions, high-frequency weak information such as voiceless consonants in speech is more vulnerable to packet loss impact, thereby significantly damaging the intelligibility and naturalness of speech.

[0003] Traditional speech encoders mostly adopt LPC analysis with a fixed window length and static bit allocation strategies, lacking the ability to respond in real time to changes in network status, and it is difficult to accurately model and preferentially encode key speech features under dynamic network conditions. At the same time, common interleaving encapsulation methods mostly adopt a linear uniform insertion mode, failing to fully consider the stability of voice energy distribution and time structure, resulting in the phenomenon that the voice backbone may still be lost or key frequency band information may be collectively damaged in the case of burst packet loss. In addition, existing high-frequency protection mechanisms often rely on fixed thresholds or redundant replication, bringing unnecessary bandwidth overhead and lacking voice semantic perception ability. Summary of the Invention

[0004] The present invention provides an intelligent packet loss resistant speech coding method and terminal for VOIP network telephones, a packet loss resistant speech coding method with differential protection capabilities and intelligent encapsulation strategies, to improve the robustness and voice restoration quality of VOIP network telephones in complex communication environments.

[0005] The intelligent packet loss resistant speech coding method for VOIP network telephones includes the following steps: S1: Based on real-time network state parameters and the time-frequency characteristics of the voice signal, perform parametric coding on the original voice frame to generate a coding feature vector including fundamental frequency trajectory parameters, formant bandwidth parameters, and voiceless / voiced flags. S2: Input the coding feature vector output by S1 into a hierarchical protection encoder, determine the basic coding layer bit allocation according to the fundamental frequency trajectory parameters, generate enhanced coding layer differential parameters based on the formant bandwidth parameters, and configure the high-frequency compensation coding layer according to the voiceless / voiced flags. S3: Utilize the hierarchical coding structure output by S2, combined with the energy correlation characteristics of adjacent voice frames, to perform adaptive interleaving encapsulation of the coding parameters, wherein the encapsulation interval of the high-frequency compensation coding layer is negatively correlated with the current network packet loss rate.

[0006] Optionally, in S1, a network state parameter set is collected through a sliding window, including the current packet loss rate, network jitter factor, and available bandwidth fluctuation value. Meanwhile, the short-time spectrum features of the speech frames are extracted, and an improved linear prediction coding analysis with an adaptive window length is adopted. The analysis window length is adjusted according to the network jitter factor to generate fundamental frequency trajectory parameters.

[0007] Optionally, in the fundamental frequency trajectory parameters, the analysis window length is shortened to 10 ms in the high-frequency jitter scenario and extended to 30 ms in the stable scenario.

[0008] Optionally, S1 further includes extracting formant bandwidth parameters based on the short-time spectrum features through the cepstral coefficients non-uniformly divided by the Mel scale, and applying bandwidth expansion compensation to the frequency bands above the preset frequency band threshold.

[0009] Optionally, the voiceless / voiced flag in S1 is generated by combining the zero-crossing rate statistic, fundamental frequency trajectory parameters, and short-time energy gradient of the speech frames. Among them, the decision threshold for voiceless sound is down-regulated correspondingly with the increase of the current packet loss rate.

[0010] Optionally, S2 specifically includes: S21, input the encoded feature vector into a hierarchical protection encoder, and calculate the bit allocation weight of the basic encoding layer according to the smoothness index of the fundamental frequency trajectory parameters; S22, construct a dynamic differential coding model based on the formant bandwidth parameters, extract the bandwidth change gradient between the current frame and the previous frame, and generate signed enhanced encoding layer differential coding parameters when the bandwidth change gradient exceeds the preset gradient threshold; S23, configure the high-frequency compensation encoding layer according to the confidence value of the voiceless / voiced flag. If the flag is a voiceless sound frame and the confidence is greater than the confidence threshold, activate the high-frequency compensation encoding channel, and the compensation intensity is positively correlated with the current network packet loss rate; S24, perform time-frequency domain interleaving on the encoding parameters of the basic encoding layer, enhanced encoding layer, and high-frequency compensation encoding layer.

[0011] Optionally, in S21, when the change rate of adjacent frames of the fundamental frequency trajectory exceeds the change threshold, the number of bits in the basic encoding layer increases accordingly (20%-30%); The S22 also includes applying a bandwidth expansion factor to the frequency bands above 4000 Hz; When performing time-frequency domain interleaving, the parameters of the high-frequency compensation layer are inserted into the encapsulated frame in a discontinuous manner, and the number of frames between adjacent high-frequency parameters is negatively correlated with the network jitter factor.

[0012] Optionally, S3 specifically includes: S31. Extract the coding parameters of the basic coding layer, the differential coding parameters of the enhancement coding layer, and the high-frequency compensation coding parameters in the hierarchical coding structure, and generate a time-domain energy correlation matrix in combination with the energy change gradient within a three-frame sliding window, where the row vector of the time-domain energy correlation matrix represents the energy correlation of the same frequency band parameters in adjacent frames; S32. Select the starting position of the encapsulation of the high-frequency compensation parameters according to the sparsity index of the time-domain energy correlation matrix, and when the energy gradient exceeds the energy mutation threshold, disperse the high-frequency parameters into the energy-stable frame sequence; S33. Construct an anti-packet-loss interleaving template, and perform non-uniform encapsulation on the parameters of the high-frequency compensation coding layer, where the number of frames between encapsulations meets the preset interval condition.

[0013] Optionally, S3 further includes performing block interleaving coding on the parameters of the basic coding layer and the enhancement coding layer, where the block size is positively correlated with the main diagonal strength of the time-domain energy correlation matrix, and performing time-slot cross-multiplexing on the interleaved data packet and the high-frequency compensation parameter packet.

[0014] A terminal, the terminal includes a memory and a processor, where: The memory is used to store executable program code; The processor is used to call the executable program code in the memory to implement the above-mentioned voice coding method.

[0015] Advantages of the present invention: In the present invention, by introducing an adaptive LPC window length control mechanism based on the network jitter factor and a voiceless determination threshold adjustment mechanism based on the dynamic packet loss rate in the parametric coding stage, a linkage modeling of the speech spectrum structure and the network dynamic behavior is realized, effectively alleviating the spectrum leakage and voiceless ambiguity problems generated by the fixed window length in a high-jitter environment; at the same time, combined with Mel scale division and bandwidth expansion compensation, the extraction accuracy of high-frequency formants is enhanced, providing a structurally stable and change-sensitive basic feature for subsequent coding, and improving the robustness and discrimination of the anti-packet-loss preprocessing.

[0016] In the present invention, through a hierarchical protection encoder design, combining the fundamental frequency trajectory change rate to control the basic layer bit weight, the formant bandwidth difference to trigger the enhancement layer modeling, and the voiceless confidence-guided high-frequency compensation activation mechanism, a dynamic tilt scheduling of coding resources between the speech mutation segment (such as plosive sounds) and the vulnerable high-frequency segment (such as voiceless consonants) is realized; especially in the scenario of increasing network packet loss rate, the high-frequency compensation intensity adapts and increases with the packet loss rate, effectively alleviating the sudden drop in speech intelligibility caused by the loss of voiceless sounds, and enhancing the recovery ability and subjective listening stability of the coding structure in a harsh network.

[0017] In the adaptive interleaving encapsulation stage of the present invention, by introducing a three-frame sliding energy correlation matrix, the accurate recognition of the temporal stationarity of voice segments and high-energy mutation points is realized, and the high-frequency compensation parameters are scattered and inserted into the energy-stable frame segments based on the gradient trigger mechanism; combined with the dynamic encapsulation interval control algorithm based on the packet loss rate, it can save transmission redundancy in low-packet-loss scenarios and achieve strong protection in high-packet-loss scenarios. Brief Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a schematic flowchart of the method of the embodiment of the present invention. Detailed Embodiments

[0020] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.

[0021] It should be noted that in the specification, when referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc., it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific feature, structure or characteristic. In addition, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such a feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0022] Generally, the terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. In addition, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but instead, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.

[0023] As Figure 1 shown, the intelligent packet-loss-resistant voice coding method for VOIP network telephones includes the following steps: S1: Perform parametric encoding on the original speech frame based on real-time network state parameters and the time-frequency characteristics of the speech signal to generate a coded feature vector including fundamental frequency trajectory parameters, formant bandwidth parameters, and voiced / unvoiced flags. S2: Input the coded feature vector output by S1 into a hierarchical protection encoder. Determine the bit allocation for the basic coding layer according to the fundamental frequency trajectory parameters, generate differential parameters for the enhanced coding layer based on the formant bandwidth parameters, and configure the high-frequency compensation coding layer according to the voiced / unvoiced flags. S3: Use the hierarchical coding structure output by S2 and combine the energy correlation characteristics of adjacent speech frames to perform adaptive interleaving encapsulation of the coding parameters. Among them, the encapsulation interval of the high-frequency compensation coding layer is negatively correlated with the current network packet loss rate.

[0024] S1 specifically includes: S11. Collect the network state parameter set through a sliding window, including: The current packet loss rate ; The network jitter factor ; The available bandwidth fluctuation value ; At the same time, extract the short-time spectrum characteristics of the speech frame , where is the frequency, is the frame time stamp.

[0025] The network jitter factor , defined as the root mean square value of the delay variation: , represents the th measured end-to-end delay value, is the mean value of the delay values, represents all measured delay values; Short-time spectrum extraction of the speech frame: , is the original speech signal, is the analysis window function (Hamming window), represents the Fourier transform, represents the th frame's spectral amplitude.

[0026] S12. Adopt an improved linear predictive coding (LPC) analysis with an adaptive window length, and dynamically adjust the analysis window length according to the network jitter factor , and calculate the fundamental frequency trajectory parameter . The window length adjustment rule is as follows: ; Among them, is the network jitter threshold, which is set according to the training data and has a default value of approximately 25 ms; The linear prediction coding (LPC) analysis obtains the prediction coefficients , and solve the prediction error signal: ; denotes the order (such as 10 to 16 orders), denotes the amplitude value of the speech signal at the -th sampling point, representing the -th sample in the time-domain discrete speech signal sequence of the current frame, is the LPC prediction value at the -th sampling point, is the LPC prediction error; The fundamental frequency trajectory estimation (extracted using the autocorrelation function) is expressed as: ; wherein, denotes the autocorrelation function when the lag is , is the audio sampling rate, is the fundamental frequency estimation value of the current frame, , are respectively the minimum and maximum lags allowed for fundamental frequency detection; The prediction coefficients are obtained as follows: a. First, extract a windowed time-domain signal from the current speech frame, and then calculate the autocorrelation sequence of the signal, which represents the correlation of the signal with itself at different time lags; b. Construct a symmetric positive definite matrix with the autocorrelation values as elements, and use it as the coefficient matrix of the linear equations. The solution of the equations is a set of linear prediction coefficients, which can predict the linear relationship between the current sampling point and its previous x sampling points in the way of minimum mean square error; c. To efficiently solve the equations, the Levinson-Durbin recursive algorithm is used, which significantly reduces the computational complexity while maintaining numerical stability. The finally obtained prediction coefficients describe the short-time spectral envelope characteristics of this frame of speech and can be used for subsequent parametric coding processing.

[0027] S13. Based on the short-time spectrum obtained in S11 , perform Mel-scale cepstral analysis to generate formant bandwidth parameters ; Definition of Mel filter bank (non-uniform): , denotes the Mel-scale value corresponding to the frequency ; Mel-frequency filtered energy calculation: ; represents the frequency response of the -th Mel filter, the output energy of the -th Mel filter (for the -th frame); After taking the logarithm and performing discrete cosine transform, cepstral coefficients are obtained , represents the -th cepstral coefficient of the -th frame, represents the number of Mel filters; Formant bandwidth parameter is obtained according to the cepstral difference: ; Bandwidth expansion compensation is applied to the high-frequency band: ; where, is the -th formant bandwidth (obtained by Mel cepstrum analysis), is the high-frequency bandwidth expansion compensation amount, set to 100 - 200 Hz, depending on the energy distribution of the frequency band, is the frequency corresponding to the -th formant, represents the value of the bandwidth compensation of the -th formant, is the preset frequency band threshold.

[0028] S14, combined with the zero-crossing rate , fundamental frequency trajectory and short-time energy gradient , to generate a dynamic voiceless / voiced flag , where the decision threshold for voiceless sound is adaptively adjusted according to the packet loss rate: ; where, is the basic noise decision threshold, initially set to about -40 dB, is the current packet loss rate (percentage), represents the floor operation.

[0029] Zero-crossing rate is calculated as: ; represents the indicator function, which is 1 when the condition holds and 0 otherwise, represents the total number of sampling points included in the current speech frame.

[0030] Short-time energy and the energy gradient are calculated as follows: ; where is the short-time energy of the th frame (the sum of the signal power within the frame), is the difference in short-time energy between adjacent frames; The final unvoiced / voiced flag is defined as: ; where represents unvoiced sound, represents voiced sound.

[0031] S2 specifically includes: S21, adjustment of the bit allocation weight in the base coding layer: S211, calculation of the fundamental frequency trajectory change rate: ; where is the fundamental frequency value of the current frame, is the inter-frame time interval (10 ms); S212, bit allocation gain factor: ; where , is the change threshold; S213, bit allocation for the base layer: ; where is the initial number of bits in the base layer, is the number of bits in the base layer after dynamic adjustment, is the bit allocation weight factor for the base coding layer, used to dynamically adjust the number of bits in the base layer according to the degree of change in the fundamental frequency trajectory, the bit increase ratio, with a value range of 0.2 to 0.3 (indicating a 20% - 30% increase).

[0032] S22, differential modeling and high-frequency extension in the enhancement coding layer: S221, bandwidth change gradient is calculated as: ; S222, conditions for generating differential coding parameters: ; where , represents the preset gradient threshold; S223, high-frequency extension compensation (for frequencies above 4000 Hz): ; where , representing the bandwidth expansion factor, the expansion compensation here and the bandwidth expansion compensation applied to the high-frequency band in S13 belong to two stages and work together. The compensation in S13 is to perform high-frequency enhancement during cepstrum extraction, so that the formant bandwidth parameter can more accurately capture high-frequency details; the compensation in S22 here is to further amplify the coding weight for the detected changing high-frequency part during coding, for a stronger robust expression of high-frequency changes.

[0033] S23, dynamic configuration of the high-frequency compensation coding layer: S231, confidence determination and channel activation: ; where represents the voiceless / voiced flag (0 for voiceless), represents the confidence value of voiceless determination, and 0.7 is the confidence threshold; S232, compensation intensity calculation: ; where is the basic compensation intensity, is the packet loss rate of the current frame.

[0034] S24, time-frequency interleaved encapsulation of hierarchical parameters: S241, calculation of the insertion interval of high-frequency compensation parameters: ; where is the network jitter factor (normalized), is the insertion frame interval between high-frequency parameters; S242, interleaving strategy: Basic layer parameters: Packed in chronological order; Enhancement layer parameters: Arranged in reverse order by frequency band; High-frequency layer parameters: Inserted one per frame to form a discontinuous encapsulation layout.

[0035] S3 specifically includes: S31, generation of the energy correlation matrix: S311, calculation of the energy gradient of the three-frame sliding window: ; where is the energy (cepstrum energy) of the th frame in the frequency band, is the energy change gradient of the frequency band in the sliding window; S312, construction of the time-domain energy correlation matrix: ; where represents the energy correlation when the adjacent frames lag by , and each row vector describes the energy stability of a certain frequency band in the sliding window.

[0036] S32, Dynamic selection of packaging location based on energy correlation matrix: S321, energy sparsity index calculation: ;in, is the maximum frame lag number, set to 3, For frequency band The larger the energy sparseness within the sliding window, the more severe the fluctuation; S322, high frequency parameter insertion strategy determination: If , then the high-frequency compensation parameters are scattered and inserted into the subsequent frames. It is the energy mutation threshold, which is 8dB / frame.

[0037] S33, encapsulation interval control based on packet loss rate: S331, dynamic interval encapsulation function: ;in, is the number of frame intervals for high frequency compensation parameters, is the current network packet loss rate (expressed as a ratio, for example, 20% is recorded as 0.2), Indicates rounding down.

[0038] S34, block interleaving coding and time slot cross multiplexing: S341, dynamic adjustment of interleaving block size: ;in, is the minimum interleaved block size (16×16 for the base layer), is the main diagonal average of the energy correlation matrix, To adjust the coefficient, the control block strength changes with stability; S342, multiplexing structure arrangement: assuming that the high frequency compensation package is , the base layer package is , then the packaging timing is: , that is, insert high-frequency packets at the 1 / 3 and 2 / 3 positions to form uniform interleaving in the time domain.

[0039] The present invention also relates to a terminal, which comprises a memory and a processor, wherein: the memory is used to store executable program codes; the processor is used to call the executable program codes in the memory to implement the above-mentioned speech encoding method.

[0040] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion with the essence of the present invention.

[0041] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements can be made without departing from the principle of the present invention, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. The intelligent packet loss resistant voice coding method for VOIP network phone, characterized in that, Including the following steps: S1: Based on real-time network state parameters and the time-frequency characteristics of the voice signal, perform parametric encoding on the original voice frame to generate an encoded feature vector including fundamental frequency trajectory parameters, formant bandwidth parameters, and voiceless / voiced flags; S2: Input the encoded feature vector output by S1 into a hierarchical protection encoder, determine the bit allocation of the basic encoding layer according to the fundamental frequency trajectory parameters, generate differential parameters of the enhanced encoding layer based on the formant bandwidth parameters, and configure the high-frequency compensation encoding layer according to the voiceless / voiced flags; S3: Utilize the hierarchical encoding structure output by S2, combine the energy correlation characteristics of adjacent voice frames, and perform adaptive interleaving encapsulation of the encoding parameters, where the encapsulation interval of the high-frequency compensation encoding layer is negatively correlated with the current network packet loss rate.

2. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 1, characterized in that In S1, collect a set of network state parameters through a sliding window, including the current packet loss rate, network jitter factor, and available bandwidth fluctuation value. At the same time, extract the short-time spectrum characteristics of the voice frame, use improved linear prediction coding analysis with an adaptive window length, and adjust the analysis window length according to the network jitter factor to generate fundamental frequency trajectory parameters.

3. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 2, characterized in that In the fundamental frequency trajectory parameters, the analysis window length is shortened to 10 ms in a high-frequency jitter scenario and extended to 30 ms in a stable scenario.

4. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 2, characterized in that, S1 further includes extracting formant bandwidth parameters based on the short-time spectrum characteristics through Mel-scale non-uniformly divided cepstral coefficients, and applying bandwidth expansion compensation to the frequency band above a preset frequency band threshold.

5. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 2, wherein The voiceless / voiced flag in S1 is generated by jointly considering the zero-crossing rate statistic, fundamental frequency trajectory parameters, and short-time energy gradient of the voice frame, where the determination threshold for voiceless sounds is down-regulated corresponding to the increase in the current packet loss rate.

6. The intelligent packet loss resistant voice coding method for VOIP network phone according to claim 1, characterized in that, S2 specifically includes: S21: Input the encoded feature vector into a hierarchical protection encoder, and calculate the bit allocation weight of the basic encoding layer according to the smoothness index of the fundamental frequency trajectory parameters; S22: Based on the formant bandwidth parameters, construct a dynamic differential coding model, extract the bandwidth change gradient between the current frame and the previous frame, and generate signed differential coding parameters of the enhanced encoding layer when the bandwidth change gradient exceeds a preset gradient threshold; S23: Configure the high-frequency compensation encoding layer according to the confidence value of the voiceless / voiced flag. If the flag indicates a voiceless frame and the confidence is greater than the confidence threshold, activate the high-frequency compensation encoding channel, and the compensation intensity is positively correlated with the current network packet loss rate; S24: Perform time-frequency domain interleaving on the encoding parameters of the basic encoding layer, enhanced encoding layer, and high-frequency compensation encoding layer.

7. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 6, characterized in that, In S21, when the change rate of adjacent frames of the fundamental frequency trajectory exceeds the change threshold, the number of bits in the basic encoding layer increases accordingly; S22 further includes applying a bandwidth expansion factor to the frequency band above 4000 Hz; When performing time-frequency domain interleaving, the parameters of the high-frequency compensation layer are inserted into the encapsulated frame in a discontinuous manner, and the number of frames between adjacent high-frequency parameters is negatively correlated with the network jitter factor.

8. The intelligent packet loss resistant voice coding method for VOIP network phone according to claim 1, characterized in that, S3 specifically includes: S31. Extract the coding parameters of the base coding layer, the differential coding parameters of the enhancement coding layer, and the high-frequency compensation coding parameters in the hierarchical coding structure, and generate a time-domain energy correlation matrix in combination with the energy change gradient within a three-frame sliding window, where the row vector of the time-domain energy correlation matrix represents the energy correlation of the same frequency band parameters in adjacent frames; S32. Select the starting position of the encapsulation of the high-frequency compensation parameters according to the sparsity index of the time-domain energy correlation matrix, and when the energy gradient exceeds the energy mutation threshold, disperse and insert the high-frequency parameters into the sequence of energy-stable frames; S33. Construct an anti-packet-loss interleaving template, and perform non-uniform encapsulation on the parameters of the high-frequency compensation coding layer, where the number of frames between encapsulations meets the preset interval condition.

9. The intelligent packet loss resistant voice coding method for VOIP network telephone according to claim 8, characterized in that, The S3 further includes performing block interleaving coding on the parameters of the base coding layer and the enhancement coding layer, where the block size is positively correlated with the main diagonal strength of the time-domain energy correlation matrix, and performing time-slot cross-multiplexing on the interleaved data packet and the high-frequency compensation parameter packet.

10. A terminal, the terminal includes a memory and a processor, characterized in that: The memory is used to store executable program code; The processor is used to call the executable program code in the memory to implement the voice coding method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Packet loss concealment for audio codec

    CN103688306A

  • High-band encoding method and device, and high-band decoding method and device

    CN111105806A

  • Low-rate voice coding and decoding method based on sine harmonic model

    CN118230741A

  • Voice coding stimulation method based on multi-peak extraction

    CN1604188A

  • Flexible frequency and time partitioning in perceptual transform coding of audio

    US20080312759A1