A narrowband speech compression method and system suitable for use in a PDT system

By using a narrowband speech compression method for the PDT system, the problem of inaccurate speech parameter extraction was solved, and clear speech output was achieved in noisy environments and with high bit error rates. This method enables low-complexity, high-quality speech compression that is adaptable to various dialects.

CN120853588BActive Publication Date: 2025-12-30XINRUIDI (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511352873.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-30
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing PDT systems suffer from several issues with speech compression technology, including inaccurate parameter extraction leading to speech distortion in synthesized speech, severe speech distortion in noisy environments, and poor speech quality under high error rates.

Method used

The narrowband speech compression method is adopted, including framing and encoding of the input speech signal, quantization, non-uniform weight forward error correction coding, and synthesis filtering. The speech quality is improved by acquiring speech activation, multi-dimensional line spectrum frequency, spectral amplitude, energy, voiced/unvoiced tones and fundamental frequency parameters from the speech signal parameters and performing non-uniform weight forward error correction coding.

Benefits of technology

It ensures speech quality in noisy environments and with high bit error rates, adapts to various dialects, and achieves low-complexity, high-quality, low-rate speech compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853588B_ABST
    Figure CN120853588B_ABST
Patent Text Reader

Abstract

The application discloses a narrowband speech compression method and system suitable for a PDT system and belongs to the technical field of digital speech processing. The method comprises the following steps: acquiring speech signal parameters corresponding to each frame; performing quantization processing on the speech signal parameters corresponding to each frame to acquire speech coding quantization bits corresponding to each frame; performing non-equal-weight forward error correction coding on part of the speech coding quantization bits corresponding to each frame to acquire forward error correction coding bits corresponding to each frame; performing forward error correction decoding on the forward error correction coding bits corresponding to each frame to acquire speech decoding quantization bits corresponding to each frame; performing inverse quantization on the speech decoding quantization bits corresponding to each frame to synthesize mixed excitation signals corresponding to each frame; and sequentially performing synthesis filtering and post-filtering processing on the mixed excitation signals corresponding to each frame to synthesize compressed narrowband speech signals. The application can ensure that the synthesized speech is clear, the speech parameters are accurate in a noise environment, and the speech quality is good under a high bit error rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital speech processing technology, and in particular relates to a narrowband speech compression method and system suitable for PDT systems. Background Technology

[0002] This invention belongs to the field of digital speech processing technology, and in particular relates to a narrowband speech compression method and system suitable for PDT systems.

[0003] Table 1. Two-bit mapping to symbols and frequency offset

[0004]

[0005] However, the speech compression technology used in the existing PDT system has problems such as inaccurate parameter extraction leading to broken speech in synthesized speech, severe speech distortion in noisy environments, and poor speech quality or inaudibility under high error rates. Summary of the Invention

[0006] One of the objectives of this invention is to provide a narrowband speech compression method suitable for PDT systems. This narrowband speech compression method can ensure clear synthesized speech, accurate speech parameters in noisy environments, good speech quality under high bit error rates, and can adapt to various dialects, low complexity, high quality, and strong noise and bit error resistance low-rate speech compression.

[0007] The second objective of this invention is to provide a narrowband voice compression system suitable for PDT systems.

[0008] To achieve one of the above objectives, the present invention employs the following technical solution:

[0009] A narrowband speech compression method suitable for PDT systems, the narrowband speech compression method comprising:

[0010] Step S1: Perform frame segmentation and encoding processing on the input speech signal in sequence to obtain the speech signal parameters corresponding to each frame;

[0011] Step S2: Quantize the speech signal parameters corresponding to each frame to obtain the speech coding quantization bits corresponding to each frame;

[0012] Step S3: Perform non-equal weight forward error correction coding on a portion of the quantized bits of the speech coding corresponding to each frame to obtain the forward error correction coding bits corresponding to each frame.

[0013] Step S4: Perform forward error correction decoding on the forward error correction coding bits corresponding to each frame to obtain the speech decoding quantization bits corresponding to each frame.

[0014] Step S5: Dequantize the speech decoding quantization bits corresponding to each frame to synthesize the hybrid excitation signal corresponding to each frame;

[0015] Step S6: Perform synthesis filtering and post-filtering processing on the hybrid excitation signal corresponding to each frame in sequence to synthesize the compressed narrowband speech signal.

[0016] Furthermore, in step S1, the specific process of encoding includes:

[0017] Step S11: Obtain the speech activation parameters from the speech signal parameters corresponding to each frame;

[0018] Step S12: Perform linear prediction analysis and spectral frequency conversion on the input speech signal of each frame and multiple sampling points in sequence to obtain the residual signal and multidimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame;

[0019] Step S13: Perform frequency domain transformation and spectral envelope extraction on the residual signal corresponding to each frame in sequence to obtain the spectral amplitude parameter in the speech signal parameters corresponding to each frame.

[0020] Step S14: Calculate the sum of squares of each sampling point of the residual signal corresponding to each frame to obtain the energy parameter in the speech signal parameters corresponding to each frame.

[0021] Step S15: Perform different bandpass filtering on the residual signal corresponding to each frame to determine the voiced / unvoiced parameters of each sub-band in the speech signal parameters corresponding to each frame.

[0022] Step S16: Using the input speech signal of each frame and multiple sampling points, determine the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0023] Furthermore, in step S16, the specific process of determining the fundamental frequency parameter in the speech signal parameters corresponding to each frame includes:

[0024] Step S1601: Perform second-order inverse filtering, low-pass filtering, and time-domain pitch period traversal on the input speech signal of each frame and multiple sampling points in sequence to obtain the time-domain signal corresponding to each selectable value of the time-domain pitch period in each frame.

[0025] Step S1602: Perform autocorrelation processing on the time-domain signal corresponding to each selectable value of the time-domain pitch period in each frame to obtain the autocorrelation value and the corresponding energy value of the current frame for each selectable value of the time-domain pitch period in each frame.

[0026] Step S1603: Sort the autocorrelation values ​​corresponding to all selectable time-domain pitch period values ​​in each frame from largest to smallest, and then sort the previous... Each autocorrelation value corresponds to a selectable time-domain pitch period, which can be used as the corresponding frame. One candidate value for the time-domain pitch period;

[0027] Step S1604: Using the energy value of this frame, normalize the autocorrelation value corresponding to each temporal pitch period candidate value of the corresponding frame;

[0028] Step S1605: Perform Discrete Fourier Transform on the time-domain signal corresponding to each frame after sequentially passing through second-order inverse filtering and low-pass filtering to obtain the digital frequency point corresponding to each time-domain pitch period candidate value of each frame.

[0029] Step S1606: Take the digital frequency point corresponding to each time-domain pitch period candidate value of each frame as the current digital frequency point, and take the maximum energy corresponding to the current digital frequency point and the left and right digital frequency points of the current digital frequency point as the harmonic energy corresponding to each time-domain pitch period candidate value of each frame.

[0030] Step S1607: Using the harmonic energy corresponding to each time-domain pitch period candidate value in each frame, calculate the maximum value of the sum of the three adjacent harmonic energies of each time-domain pitch period candidate value in each frame;

[0031] Step S1608: Select the largest value from the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame as a weighting factor to normalize the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame.

[0032] Step S1609: Using multiple temporal pitch period candidate values ​​of each frame and their corresponding weighting factors and normalized autocorrelation values, perform dynamic programming to determine the cost function of the trajectory path corresponding to each temporal pitch period candidate value of each frame.

[0033] Step S1610: Select the trajectory path with the smallest cost function from the cost functions of the trajectory paths corresponding to all temporal pitch period candidate values ​​in each frame as the temporal pitch period change trajectory of the corresponding frame, so as to obtain the temporal pitch period candidate values ​​passed by the temporal pitch period change trajectory of each frame.

[0034] Step S1611: Take the candidate values ​​of the time-domain pitch period through which the time-domain pitch period change trajectory of each frame passes as the time-domain pitch period of the corresponding frame, so as to calculate the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0035] Furthermore, in step S2, the specific implementation process of quantization includes:

[0036] Step S21: Represent the speech activation parameters in the speech signal parameters corresponding to each frame using 1 bit;

[0037] Step S22: Perform three-level residual vector quantization on the multidimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame, and represent them using three levels of bit numbers: 8, 8, and 7.

[0038] Step S23: Perform 3-bit vector quantization on the spectral amplitude parameter and the voiced / unvoiced parameters of each subband in the speech signal parameters corresponding to each frame;

[0039] Step S24: After taking the logarithmic value of the energy parameter in the corresponding speech signal parameters of each frame, perform 6-bit non-uniform scalar quantization in the range of 10~77dB.

[0040] Step S25: Perform 7-bit non-uniform scalar quantization on the fundamental frequency parameter in the speech signal parameters corresponding to each frame within the range of 54~444Hz.

[0041] Furthermore, in step S3, the specific process of non-equal weight forward error correction coding includes:

[0042] Step S31: Using 1 bit, perform parity check on bits 1 to 9 of the speech coding quantization bits corresponding to each frame to output 1 bit;

[0043] Step S32: Perform (15,11) BCH encoding on bits 14 to 24 of the speech coding quantization bits corresponding to each frame to output 14 bits;

[0044] Step S33: Encode the 2nd to 13th bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits;

[0045] Step S34: Encode the 25th, 26th, 27th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, and 41st bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits.

[0046] To achieve the second objective mentioned above, the present invention employs the following technical solution:

[0047] A narrowband voice compression system suitable for PDT systems, the narrowband voice compression system comprising:

[0048] The encoding processing module is configured to perform framing and encoding processing on the input speech signal to obtain the speech signal parameters corresponding to each frame.

[0049] The quantization processing module is configured to: perform quantization processing on the speech signal parameters corresponding to each frame to obtain the speech coding quantization bits corresponding to each frame;

[0050] The non-equal weight forward error correction coding module is configured to perform non-equal weight forward error correction coding on a portion of the bits in the speech coding quantization bits corresponding to each frame to obtain the forward error correction coding bits corresponding to each frame.

[0051] The non-equal weight forward error correction decoding module is configured to perform forward error correction decoding on the forward error correction coding bits corresponding to each frame to obtain the speech decoding quantization bits corresponding to each frame.

[0052] The dequantization module is configured to dequantize the speech decoding quantization bits corresponding to each frame in order to synthesize the hybrid excitation signal corresponding to each frame.

[0053] The post-filtering module is configured to sequentially perform synthesis filtering and post-filtering processing on the hybrid excitation signal corresponding to each frame to synthesize a compressed narrowband speech signal.

[0054] Furthermore, the encoding processing module includes:

[0055] The acquisition submodule is configured to: acquire the speech activation parameters from the speech signal parameters corresponding to each frame;

[0056] The spectrum frequency conversion submodule is configured to perform linear prediction analysis and spectrum frequency conversion on the input speech signal of each frame and multiple sampling points in sequence, so as to obtain the residual signal and multi-dimensional line spectrum frequency parameters in the speech signal parameters of each frame.

[0057] The extraction submodule is configured to sequentially perform frequency domain transformation and spectral envelope extraction on the residual signal corresponding to each frame in order to obtain the spectral amplitude parameter in the speech signal parameters corresponding to each frame.

[0058] The calculation submodule is configured to: calculate the sum of squares of each sampling point of the residual signal corresponding to each frame, so as to obtain the energy parameter in the speech signal parameter corresponding to each frame;

[0059] The bandpass filtering submodule is configured to perform different bandpass filtering on the residual signal corresponding to each frame in order to determine the voiced / unvoiced parameters of each subband in the speech signal parameters corresponding to each frame.

[0060] The determination submodule is configured to: use the input speech signal of each frame and multiple sampling points to determine the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0061] Furthermore, the sub-modules are identified as including:

[0062] The filtering subunit is configured to sequentially perform second-order inverse filtering, low-pass filtering, and time-domain pitch period traversal on the input speech signal of each frame and multiple sampling points to obtain the time-domain signal corresponding to each selectable value of the time-domain pitch period in each frame.

[0063] The autocorrelation processing subunit is configured to perform autocorrelation processing on the temporal signal corresponding to each selectable value of the temporal pitch period in each frame, so as to obtain the autocorrelation value and the corresponding energy value of the current frame for each selectable value of the temporal pitch period in each frame.

[0064] The sorting subunit is configured to: sort the autocorrelation values ​​corresponding to all selectable temporal pitch period values ​​in each frame from largest to smallest, and then sort the previous... Each autocorrelation value corresponds to a selectable time-domain pitch period, which can be used as the corresponding frame. One candidate value for the time-domain pitch period;

[0065] The first normalization subunit is configured to: use the energy value of this frame to normalize the autocorrelation value corresponding to each temporal pitch period candidate value of the corresponding frame;

[0066] The Discrete Fourier Transform subunit is configured to perform a Discrete Fourier Transform on the time-domain signal corresponding to each frame after sequentially passing through a second-order inverse filter and a low-pass filter, in order to obtain the digital frequency point corresponding to each candidate value of the time-domain pitch period in each frame.

[0067] In this process, the digital frequency point corresponding to each candidate value of the time-domain pitch period in each frame is taken as the current digital frequency point, and the maximum energy corresponding to the current digital frequency point and the left and right digital frequency points of the current digital frequency point is taken as the harmonic energy corresponding to each candidate value of the time-domain pitch period in each frame.

[0068] The first calculation subunit is configured to: use the harmonic energy corresponding to each time-domain pitch period candidate value in each frame to calculate the maximum value of the sum of the three adjacent harmonic energies of each time-domain pitch period candidate value in each frame;

[0069] The second normalization subunit is configured to: select the largest value from the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame as a weighting factor, so as to normalize the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame.

[0070] The dynamic programming subunit is configured to: use multiple temporal pitch period candidate values ​​of each frame and their corresponding weighting factors and normalized autocorrelation values ​​to perform dynamic programming in order to determine the cost function of the trajectory path corresponding to each temporal pitch period candidate value of each frame;

[0071] The selected sub-unit is configured to: select the trajectory path with the smallest cost function from the cost functions of the trajectory paths corresponding to all temporal pitch period candidate values ​​in each frame as the temporal pitch period change trajectory of the corresponding frame, so as to obtain the temporal pitch period candidate values ​​passed by the temporal pitch period change trajectory of each frame.

[0072] The second calculation subunit is configured to: take the candidate values ​​of the time-domain pitch period through which the trajectory of the time-domain pitch period change of each frame passes as the time-domain pitch period of the corresponding frame, so as to calculate the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0073] Furthermore, the quantization processing module includes:

[0074] The bit representation submodule is configured to represent the speech activation parameters in the speech signal parameters corresponding to each frame with 1 bit.

[0075] The three-level residual vector quantization submodule is configured to perform three-level residual vector quantization on the multi-dimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame, and to represent them using three levels of bit numbers: 8, 8, and 7.

[0076] The 3-bit vector quantization submodule is configured to perform 3-bit vector quantization on the spectral amplitude parameter and the voiced / unvoiced parameters of each subband in the speech signal parameters corresponding to each frame.

[0077] The first non-uniform scalar quantization submodule is configured to: take the logarithm of the energy parameter in the speech signal parameters corresponding to each frame, and then perform 6-bit non-uniform scalar quantization in the range of 10~77dB.

[0078] The second non-uniform scalar quantization submodule is configured to perform 7-bit non-uniform scalar quantization on the fundamental frequency parameters in the corresponding speech signal parameters of each frame within the range of 54~444Hz.

[0079] Furthermore, the non-equal weight forward error correction coding module includes:

[0080] The parity check submodule is configured to use 1 bit to perform parity check on bits 1 to 9 of the speech coding quantization bits corresponding to each frame, so as to output 1 bit.

[0081] The BCH encoding submodule is configured to perform (15,11) BCH encoding on bits 14 to 24 of the speech coding quantization bits corresponding to each frame to output 14 bits.

[0082] The first systematic Gray code encoding submodule is configured to encode the 2nd to 13th bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits.

[0083] The second systematic Gray code encoding submodule is configured to encode the 25th, 26th, 27th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, and 41st bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits.

[0084] In summary, the solution proposed in this invention has the following technical effects:

[0085] This invention encodes and quantizes the input speech signal to select important high-order bits for forward error correction coding, while neglecting unimportant low-order bits. By using non-equal-weight forward error correction coding on a subset of the quantized bits in each frame, the accuracy of the important bits in each frame is ensured, thereby guaranteeing clear speech synthesized from all frames. This results in accurate speech parameters even in noisy environments, good speech quality under high error rates, and low-rate speech compression that is adaptable to various Chinese dialects, with low complexity, high quality, and strong noise and error resistance. Attached Figure Description

[0086] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0087] Figure 1 This is a flowchart of a narrowband voice compression method for a PDT system according to an embodiment of the present invention. Detailed Implementation

[0088] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0089] This embodiment presents a narrowband voice compression method suitable for PDT systems, referencing... Figure 1 The narrowband speech compression method includes:

[0090] Step S1: Perform frame segmentation and encoding processing on the input speech signal in sequence to obtain the speech signal parameters corresponding to each frame.

[0091] This embodiment divides the input voice signal into frames. For example, a voice signal input from 50 to 4000 Hz and sampled at 8 kHz is processed into frames of 20 ms.

[0092] Because LSF parameters have better robustness compared to LPC, this embodiment employs a hybrid excitation linear prediction model to perform linear prediction coding (LPC) on the input speech signal to obtain the margin signal. The LPC coefficients are then converted into easily quantifiable multidimensional (e.g., 10-dimensional) line spectrum frequency (LSF) parameters. Simultaneously, parameters characterizing the margin signal, including the gain, unvoiced / voiced (U / V) parameters for multiple subbands (e.g., 5 subbands), the fundamental frequency parameter, and the prototype parameter, are extracted. The specific encoding process includes:

[0093] Step S11: Obtain the speech activation parameters from the speech signal parameters corresponding to each frame;

[0094] Step S12: Perform linear prediction analysis and spectral frequency conversion on the input speech signal of each frame and multiple sampling points in sequence to obtain the residual signal and multidimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame;

[0095] For example, linear prediction analysis is performed on a 20ms frame with 160 sampling points of input speech signal to obtain 10-dimensional linear prediction coefficients, which are then converted into 10-dimensional line spectrum frequency (LSF) parameters.

[0096] Step S13: Perform frequency domain transformation and spectral envelope extraction on the residual signal corresponding to each frame in sequence to obtain the spectral amplitude parameter in the speech signal parameters corresponding to each frame.

[0097] After linear predictive analysis, the residual signal is transformed in the frequency domain, and the relative values ​​of the first 10-dimensional spectral envelope are extracted to obtain the spectral amplitude parameter in the speech signal parameters.

[0098] Step S14: Calculate the sum of squares of each sampling point of the residual signal corresponding to each frame to obtain the energy parameter in the speech signal parameter corresponding to each frame.

[0099] Step S15: Perform different bandpass filtering on the residual signal corresponding to each frame to determine the voiced / unvoiced parameters of each sub-band in the speech signal parameters corresponding to each frame.

[0100] For example, a 20ms frame with 160 sampling points of input speech signal is processed through five bandpass filters at frequencies of 50~500Hz, 501~1000Hz, 1001~2000Hz, 2001~3000Hz, and 3001~4000Hz to determine the voiced / unvoiced parameters.

[0101] Step S16: Using the input speech signal of each frame and multiple sampling points, determine the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0102] The 20ms frame of input speech with 160 sample points is processed by an 800Hz low-pass filter and a second-order inverse filter. Then, the fundamental frequency parameters of the frame are analyzed and smoothed by combining time and frequency domains.

[0103] The specific process for determining the fundamental frequency parameter in the speech signal parameters corresponding to each frame in this embodiment includes:

[0104] Step S1601: Perform second-order inverse filtering, low-pass filtering, and time-domain pitch period traversal on the input speech signal of each frame and multiple sampling points in sequence to obtain the time-domain signal corresponding to each selectable value of the time-domain pitch period in each frame.

[0105] In this embodiment, the cutoff frequency of the low-pass filter can be 800Hz. After second-order inverse filtering and low-pass filtering, the selectable value range of the time-domain pitch period is traversed from 18 to 148.

[0106] Step S1602: Perform autocorrelation processing on the time-domain signal corresponding to each selectable value of the time-domain pitch period in each frame to obtain the autocorrelation value and the corresponding energy value of the current frame for each selectable value of the time-domain pitch period in each frame.

[0107] Perform autocorrelation processing according to the following formula:

[0108]

[0109] in, Selectable values ​​for the time-domain pitch period The autocorrelation value at time; This indicates the 800Hz low-pass filter result. A time-domain signal, when At that time, , For a fixed threshold value, The total number of autocorrelation sampling points; when At that time, , For a fixed threshold value, This represents the total number of autocorrelation sampling points.

[0110] Step S1603: Sort the autocorrelation values ​​corresponding to all selectable time-domain pitch period values ​​in each frame from largest to smallest, and then sort the previous... Each autocorrelation value corresponds to a selectable time-domain pitch period, which can be used as the corresponding frame. One time-domain pitch period candidate value.

[0111] The five largest values ​​of the correlation function are selected as candidate values ​​for the time-domain pitch period, denoted as follows: The corresponding autocorrelation value is denoted as .

[0112] Step S1604: Using the energy value of this frame, normalize the autocorrelation value corresponding to each temporal pitch period candidate value of the corresponding frame.

[0113] Utilize the energy of this frame (Selectable values ​​for the time domain pitch period) The autocorrelation value is 0. Normalization is performed on each candidate value of the time-domain pitch period to obtain the corresponding normalized autocorrelation value. .

[0114] Step S1605: Perform Discrete Fourier Transform on the time-domain signal corresponding to each frame after sequentially passing through second-order inverse filtering and low-pass filtering to obtain the digital frequency point corresponding to each candidate value of the time-domain pitch period of each frame.

[0115] To obtain the maximum sum of the energies of the three adjacent harmonics for each candidate time-domain pitch period, a Discrete Fourier Transform (DFT) is performed on the 800Hz low-pass filtered time-domain signal. The number of points in the Fourier Transform is... Five candidate values ​​for the time-domain pitch period. The corresponding digital frequency points are roughly distributed in ,in, This indicates rounding.

[0116] Step S1606: Take the digital frequency point corresponding to each candidate value of the time-domain pitch period in each frame as the current digital frequency point, and take the maximum energy corresponding to the current digital frequency point and the left and right digital frequency points of the current digital frequency point as the harmonic energy corresponding to each candidate value of the time-domain pitch period in each frame.

[0117] Because the calculation of harmonic energy in the high-frequency range may have some deviation, the maximum energy of the three frequency points (two left and right frequency points plus the current frequency point) is taken as the maximum energy. The energy of the second harmonic is:

[0118]

[0119] in, For the first The corresponding candidate values ​​of each time-domain pitch period Subharmonic energy; , and The first The digital frequency points, left digital frequency points, and right digital frequency points corresponding to each candidate value of the time-domain pitch period Secondary energy.

[0120] Step S1607: Using the harmonic energy corresponding to each time-domain pitch period candidate value in each frame, calculate the maximum value of the sum of the three adjacent harmonic energies of each time-domain pitch period candidate value in each frame;

[0121] For each time-domain pitch period candidate value , To avoid interference from high-frequency noise, the maximum number of harmonics to be calculated is limited to [number missing]. The maximum value of the sum of its three adjacent harmonic energies for:

[0122]

[0123] in, , and Indicates the first The corresponding candidate values ​​of each time-domain pitch period Subharmonic energy, Subharmonic energy and Subharmonic energy.

[0124] Step S1608: Select the largest value from the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame as the normalization factor, so as to normalize the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame.

[0125] Select the normalization factor according to the following formula:

[0126] ;

[0127] in, This is the normalization factor.

[0128] Normalize using the following formula:

[0129] ;

[0130] in, For each frame corresponding to the first The weighting factor corresponding to each candidate value of the time-domain pitch period.

[0131] Step S1609: Using multiple temporal pitch period candidate values ​​of each frame and their corresponding weighting factors and normalized autocorrelation values, perform dynamic programming to determine the cost function of the trajectory path corresponding to each temporal pitch period candidate value of each frame.

[0132] This embodiment utilizes the first The first frame corresponding to Weighting factor corresponding to each candidate value of the time-domain pitch period Combined with the first The temporal pitch period of the frame, the first The candidate values ​​of each temporal pitch period corresponding to the frame and the first The candidate values ​​of each temporal pitch period corresponding to the frame are dynamically planned.

[0133] Calculate the cost function for each trajectory path using the following formula:

[0134] ;

[0135] ;

[0136] in, For the first The first frame The cost function of the trajectory path corresponding to the candidate value of the time-domain pitch period (i.e., from the ) Frame to the The first frame corresponding to From the candidate value of the first time-domain pitch period to the first The first frame corresponding to (the entire path cost function for each candidate time-domain pitch period). Indicates the first The temporal pitch period of the frame; Indicates the first The first frame corresponding to One candidate value for the time-domain pitch period; Indicates the first The first frame corresponding to One candidate value for the time-domain pitch period; For the first Frame to the The first frame corresponding to The path cost of candidate values ​​for each time-domain pitch period; For the first The first frame corresponding to The point cost of a candidate value for a time-domain pitch period; For the first The first frame corresponding to The candidate values ​​of the time-domain pitch period are up to the first... The first frame corresponding to The path cost of candidate values ​​for each time-domain pitch period; For the first The first frame corresponding to The point cost of a candidate value for a time-domain pitch period; For the first The first frame corresponding to Weighting factors corresponding to each candidate value of the time-domain pitch period; For the first The first frame corresponding to Normalized autocorrelation values ​​corresponding to candidate values ​​of time-domain pitch period; It is a fixed value, usually taken as 0.8.

[0137] Step S1610: Select the trajectory path with the smallest cost function from the cost functions of the trajectory paths corresponding to all temporal pitch period candidate values ​​in each frame as the temporal pitch period change trajectory of the corresponding frame, so as to obtain the temporal pitch period candidate values ​​passed by the temporal pitch period change trajectory of each frame.

[0138] After obtaining the cost function of each trajectory path, the path with the smallest cost function is selected as the final time-domain pitch period variation trajectory.

[0139] Step S1611: Take the candidate values ​​of the time-domain pitch period through which the time-domain pitch period change trajectory of each frame passes as the time-domain pitch period of the corresponding frame, so as to calculate the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0140] The candidate values ​​of the temporal pitch period of the current frame through which the trajectory of the temporal pitch period variation passes are the temporal pitch period of the corresponding frame. Then, using the formula fundamental frequency = sampling rate / pitch period, the final fundamental frequency parameter is obtained.

[0141] Step S2: Quantize the speech signal parameters corresponding to each frame to obtain the speech coding quantization bits corresponding to each frame.

[0142] The specific implementation process of quantization in this embodiment includes:

[0143] Step S21: Represent the speech activation parameters in the speech signal parameters corresponding to each frame using 1 bit;

[0144] In this embodiment, for input voice, a bit with a default value of 1 is used to represent that the input is a voice signal.

[0145] Step S22: Perform three-level residual vector quantization on the multidimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame, and represent them using three levels of bit numbers: 8, 8, and 7.

[0146] In this embodiment, the 10-dimensional line spectrum frequency parameters are quantized using a three-level residual vector quantizer (RVQ), with the number of bits for each level being 8, 8, and 7.

[0147] Step S23: Perform 3-bit vector quantization on the spectral amplitude parameter and the voiced / unvoiced parameters of each subband in the speech signal parameters corresponding to each frame.

[0148] The 10-dimensional spectral amplitude parameter and the voiced / unvoiced parameters of the five subbands are all quantized using 3-bit vector quantization.

[0149] Step S24: After taking the logarithmic value of the energy parameter in the corresponding speech signal parameters of each frame, perform 6-bit non-uniform scalar quantization in the range of 10~77dB.

[0150] The energy parameters are logarithmically calculated and limited to the range of 10~77dB. 6-bit non-uniform scalar quantization is used to facilitate the selection of important high bits for forward error correction coding, while unimportant low bits are not forward error correction coded.

[0151] Step S25: Perform 7-bit non-uniform scalar quantization on the fundamental frequency parameter in the speech signal parameters corresponding to each frame within the range of 54~444Hz.

[0152] The fundamental frequency parameter is limited to the range of 54~444Hz, and 7-bit non-uniform scalar quantization is used to facilitate the selection of important high bits for forward error correction coding, while unimportant low bits are not subject to forward error correction coding.

[0153] Step S3: Perform non-equal weight forward error correction coding on a portion of the quantized bits of the speech coding corresponding to each frame to obtain the forward error correction coding bits corresponding to each frame.

[0154] In this embodiment, some bits include bits 1-9, bits 14-24, bits 2-13, and bits 25, 26, 27, 31, 32, 33, 34, 35, 36, 37, 38, and 41. Bits 1-9 of the speech encoding output are parity checked using 1 bit. Based on the standard (15,11) BCH (Bose Chaudhuri Hocquenghem) code, 4 bits are used for error correction coding of bits 14-24 (11 bits) of the speech encoding output. Based on the standard (24,12) Systematic Gray Code, bits 2-13 (12 bits) are error-corrected to generate a 12-bit parity bit. Based on the standard (24,12) Systematic Gray Code, bits 25, 26, 27, 31, 32, 33, 34, 35, 36, 37, 38, and 41 (12 bits) are error-corrected to generate a 12-bit parity bit. The two sets of Gray codes result in a total of 48 bits of output. For each 20ms frame of speech encoding, 43 bits are output. 35 bits are then subjected to non-uniform weight forward error correction coding by adding 29 parity bits as described above. The remaining 8 bits are output directly without forward error correction coding. After non-uniform weight forward error correction coding, 72 bits are output per 20ms frame.

[0155] In 4FSK modulation, when two bits 01 and 11 are mapped to symbols +3 and -3, and 00 and 10 are mapped to symbols +1 and -1, under 01 error, it is easy to be mistaken for 00 but not for 10 and 11. Under 00 error, the probability of being mistaken for 01 and 10 is about the same. Therefore, the error probability of high bits is less than that of low bits.

[0156] The specific process of non-equal weight forward error correction coding in this embodiment includes:

[0157] Step S31: Using 1 bit, perform parity check on bits 1 to 9 of the speech coding quantization bits corresponding to each frame to output 1 bit;

[0158] Step S32: Perform (15,11) BCH encoding on bits 14 to 24 of the speech coding quantization bits corresponding to each frame to output 14 bits;

[0159] Step S33: Encode the 2nd to 13th bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits;

[0160] Step S34: Encode the 25th, 26th, 27th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, and 41st bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits.

[0161] Step S4: Perform forward error correction decoding on the forward error correction coding bits corresponding to each frame to obtain the speech decoding quantization bits corresponding to each frame.

[0162] Step S5: Dequantize the speech decoding quantization bits corresponding to each frame to synthesize the hybrid excitation signal corresponding to each frame.

[0163] The excitation signal is synthesized using the parameters obtained from the inverse quantization. For each subband, if it is a voiced tone, the energy is superimposed using the fundamental frequency parameter to synthesize a periodic excitation signal; if it is an unvoiced tone, the energy is superimposed using a random noise generator to synthesize a noise excitation signal.

[0164] Step S6: Perform synthesis filtering and post-filtering processing on the hybrid excitation signal corresponding to each frame in sequence to synthesize the compressed narrowband speech signal.

[0165] After the excitation signals of the five sub-bands are superimposed, they are synthesized into the final speech signal through a synthesis filter and a post-filter. The coefficients of the synthesis filter are the LPC coefficients.

[0166] This embodiment uses encoding and quantization of the input speech signal parameters to select important high-order bits for forward error correction coding, while unimportant low-order bits are not. By using non-equal weight forward error correction coding of some bits in the speech coding quantization bits corresponding to each frame, the accuracy of the important bits corresponding to each frame is ensured, thereby ensuring clear speech synthesized from all frames, accurate speech parameters in noisy environments, good speech quality under high bit error rate, and adaptability to various Chinese dialects, low complexity, high quality, and strong noise and bit error resistance low-rate speech compression.

[0167] The technical solutions of the above embodiments can be implemented through the technical solutions given in the following embodiments:

[0168] Another embodiment provides a narrowband voice compression system suitable for PDT systems, the narrowband voice compression system comprising:

[0169] The encoding processing module is configured to perform framing and encoding processing on the input speech signal to obtain the speech signal parameters corresponding to each frame.

[0170] The quantization processing module is configured to: perform quantization processing on the speech signal parameters corresponding to each frame to obtain the speech coding quantization bits corresponding to each frame;

[0171] The non-equal weight forward error correction coding module is configured to perform non-equal weight forward error correction coding on a portion of the bits in the speech coding quantization bits corresponding to each frame to obtain the forward error correction coding bits corresponding to each frame.

[0172] Among them, some bits include bits 1-9, bits 14-24, bits 2-13, and bits 25, 26, 27, 31, 32, 33, 34, 35, 36, 37, 38, and 41;

[0173] The non-equal weight forward error correction decoding module is configured to perform forward error correction decoding on the forward error correction coding bits corresponding to each frame to obtain the speech decoding quantization bits corresponding to each frame.

[0174] The dequantization module is configured to dequantize the speech decoding quantization bits corresponding to each frame in order to synthesize the hybrid excitation signal corresponding to each frame.

[0175] The post-filtering module is configured to sequentially perform synthesis filtering and post-filtering processing on the hybrid excitation signal corresponding to each frame to synthesize a compressed narrowband speech signal.

[0176] Furthermore, the encoding processing module includes:

[0177] The acquisition submodule is configured to: acquire the speech activation parameters from the speech signal parameters corresponding to each frame;

[0178] The spectrum frequency conversion submodule is configured to perform linear prediction analysis and spectrum frequency conversion on the input speech signal of each frame and multiple sampling points in sequence, so as to obtain the residual signal and multi-dimensional line spectrum frequency parameters in the speech signal parameters of each frame.

[0179] The extraction submodule is configured to sequentially perform frequency domain transformation and spectral envelope extraction on the residual signal corresponding to each frame in order to obtain the spectral amplitude parameter in the speech signal parameters corresponding to each frame.

[0180] The calculation submodule is configured to: calculate the sum of squares of each sampling point of the residual signal corresponding to each frame, so as to obtain the energy parameter in the speech signal parameter corresponding to each frame;

[0181] The bandpass filtering submodule is configured to perform different bandpass filtering on the residual signal corresponding to each frame in order to determine the voiced / unvoiced parameters of each subband in the speech signal parameters corresponding to each frame.

[0182] The determination submodule is configured to: use the input speech signal of each frame and multiple sampling points to determine the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0183] Furthermore, the sub-modules are identified as including:

[0184] The filtering subunit is configured to sequentially perform second-order inverse filtering, low-pass filtering, and time-domain pitch period traversal on the input speech signal of each frame and multiple sampling points to obtain the time-domain signal corresponding to each selectable value of the time-domain pitch period in each frame.

[0185] The autocorrelation processing subunit is configured to perform autocorrelation processing on the temporal signal corresponding to each selectable value of the temporal pitch period in each frame, so as to obtain the autocorrelation value and the corresponding energy value of the current frame for each selectable value of the temporal pitch period in each frame.

[0186] The sorting subunit is configured to: sort the autocorrelation values ​​corresponding to all selectable temporal pitch period values ​​in each frame from largest to smallest, and then sort the previous... Each autocorrelation value corresponds to a selectable time-domain pitch period, which can be used as the corresponding frame. One candidate value for the time-domain pitch period;

[0187] The first normalization subunit is configured to: use the energy value of this frame to normalize the autocorrelation value corresponding to each temporal pitch period candidate value of the corresponding frame;

[0188] The Discrete Fourier Transform subunit is configured to perform a Discrete Fourier Transform on the time-domain signal corresponding to each frame after sequentially passing through a second-order inverse filter and a low-pass filter, in order to obtain the digital frequency point corresponding to each candidate value of the time-domain pitch period in each frame.

[0189] In this process, the digital frequency point corresponding to each candidate value of the time-domain pitch period in each frame is taken as the current digital frequency point, and the maximum energy corresponding to the current digital frequency point and the left and right digital frequency points of the current digital frequency point is taken as the harmonic energy corresponding to each candidate value of the time-domain pitch period in each frame.

[0190] The first calculation subunit is configured to: use the harmonic energy corresponding to each time-domain pitch period candidate value in each frame to calculate the maximum value of the sum of the three adjacent harmonic energies of each time-domain pitch period candidate value in each frame;

[0191] The second normalization subunit is configured to: select the largest value from the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame as a weighting factor, so as to normalize the maximum sum of the three adjacent harmonic energies of each temporal pitch period candidate value in each frame.

[0192] The dynamic programming subunit is configured to: use multiple temporal pitch period candidate values ​​of each frame and their corresponding weighting factors and normalized autocorrelation values ​​to perform dynamic programming in order to determine the cost function of the trajectory path corresponding to each temporal pitch period candidate value of each frame;

[0193] The selected sub-unit is configured to: select the trajectory path with the smallest cost function from the cost functions of the trajectory paths corresponding to all temporal pitch period candidate values ​​in each frame as the temporal pitch period change trajectory of the corresponding frame, so as to obtain the temporal pitch period candidate values ​​passed by the temporal pitch period change trajectory of each frame.

[0194] The second calculation subunit is configured to: take the candidate values ​​of the time-domain pitch period through which the trajectory of the time-domain pitch period change of each frame passes as the time-domain pitch period of the corresponding frame, so as to calculate the fundamental frequency parameter in the speech signal parameters corresponding to each frame.

[0195] Furthermore, the quantization processing module includes:

[0196] The bit representation submodule is configured to represent the speech activation parameters in the speech signal parameters corresponding to each frame with 1 bit.

[0197] The three-level residual vector quantization submodule is configured to perform three-level residual vector quantization on the multi-dimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame, and to represent them using three levels of bit numbers: 8, 8, and 7.

[0198] The 3-bit vector quantization submodule is configured to perform 3-bit vector quantization on the spectral amplitude parameter and the voiced / unvoiced parameters of each subband in the speech signal parameters corresponding to each frame.

[0199] The first non-uniform scalar quantization submodule is configured to: take the logarithm of the energy parameter in the speech signal parameters corresponding to each frame, and then perform 6-bit non-uniform scalar quantization in the range of 10~77dB.

[0200] The second non-uniform scalar quantization submodule is configured to perform 7-bit non-uniform scalar quantization on the fundamental frequency parameters in the corresponding speech signal parameters of each frame within the range of 54~444Hz.

[0201] Furthermore, the non-equal weight forward error correction coding module includes:

[0202] The parity check submodule is configured to use 1 bit to perform parity check on bits 1 to 9 of the speech coding quantization bits corresponding to each frame, so as to output 1 bit.

[0203] The BCH encoding submodule is configured to perform (15,11) BCH encoding on bits 14 to 24 of the speech coding quantization bits corresponding to each frame to output 14 bits.

[0204] The first systematic Gray code encoding submodule is configured to encode the 2nd to 13th bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits.

[0205] The second systematic Gray code encoding submodule is configured to encode the 25th, 26th, 27th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, and 41st bits of the speech coding quantization bits corresponding to each frame using (24,12) systematic Gray code to output 24 bits.

[0206] The principles, formulas, and parameter definitions involved in the above embodiments are all applicable and will not be repeated here.

[0207] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A narrowband speech compression method suitable for use in a PDT system, characterized by, The narrowband speech compression method comprises: Step S1, sequentially performing frame division and encoding processing on the input speech signal to obtain speech signal parameters corresponding to each frame; Step S2, performing quantization processing on the speech signal parameters corresponding to each frame to obtain speech encoding quantization bits corresponding to each frame; Step S3, performing non-equal-weight forward error correction coding on part of the speech encoding quantization bits corresponding to each frame to obtain forward error correction coding bits corresponding to each frame; Step S4, performing forward error correction decoding on the forward error correction coding bits corresponding to each frame to obtain speech decoding quantization bits corresponding to each frame; Step S5, performing inverse quantization on the speech decoding quantization bits corresponding to each frame to synthesize mixed excitation signals corresponding to each frame; Step S6, sequentially performing synthesis filtering and post-filtering processing on the mixed excitation signals corresponding to each frame to synthesize compressed narrowband speech signals.

2. The narrowband speech compression method of claim 1, wherein, In the step S1, the specific process of the encoding processing comprises: Step S11, obtaining speech activation parameters in the speech signal parameters corresponding to each frame; Step S12, sequentially performing linear prediction analysis and spectral frequency conversion on the input speech signal of each frame and multiple sampling points to obtain residual signals corresponding to each frame and multi-dimensional line spectrum frequency parameters in the speech signal parameters; Step S13, sequentially performing frequency domain transformation and spectral envelope extraction on the residual signals corresponding to each frame to obtain spectral amplitude parameters in the speech signal parameters corresponding to each frame; Step S14, calculating the square sum of each sampling point of the residual signals corresponding to each frame to obtain energy parameters in the speech signal parameters corresponding to each frame; Step S15, performing different band-pass filtering on the residual signals corresponding to each frame to determine voicing parameters of each sub-band in the speech signal parameters corresponding to each frame; Step S16, determining a fundamental frequency parameter in the speech signal parameters corresponding to each frame by using the input speech signal of each frame and multiple sampling points.

3. The narrowband speech compression method of claim 2, wherein, In the step S16, the specific process of determining the fundamental frequency parameter in the speech signal parameters corresponding to each frame comprises: Step S1601, sequentially performing second-order inverse filtering, low-pass filtering and time-domain pitch period traversal on the input speech signal of each frame and multiple sampling points to obtain time-domain signals corresponding to each time-domain pitch period candidate value in each frame; Step S1602, performing autocorrelation processing on the time-domain signals corresponding to each time-domain pitch period candidate value in each frame to obtain autocorrelation values corresponding to each time-domain pitch period candidate value in each frame and corresponding frame energy values; Step S1603, for all time-domain pitch period optional values in each frame, the autocorrelation values corresponding to the time-domain pitch period optional values are sorted from large to small, and the time-domain pitch period optional values corresponding to the first autocorrelation values are taken as the first time-domain pitch period candidate values of the corresponding frame. Step S1604, normalizing the autocorrelation values corresponding to each time-domain pitch period candidate value of the corresponding frame by using the frame energy values; Step S1605, performing discrete Fourier transform on the time-domain signals corresponding to each frame after sequentially passing through the second-order inverse filtering and the low-pass filtering to obtain digital frequency points corresponding to each time-domain pitch period candidate value of each frame; Step S1606, taking the digital frequency points corresponding to each time-domain pitch period candidate value of each frame as a current digital frequency point, and taking the maximum energy corresponding to the current digital frequency point and the left and right digital frequency points of the current digital frequency point as the harmonic energy corresponding to each time-domain pitch period candidate value of each frame. Step S1607, using the harmonic energy corresponding to each time-domain pitch period candidate value of each frame, calculating the maximum value of the sum of the adjacent three harmonic energies of each time-domain pitch period candidate value of each frame; Step S1608, selecting the maximum value from the maximum values of the sum of the adjacent three harmonic energies of each time-domain pitch period candidate value of each frame as a weighting factor, to normalize the maximum value of the sum of the adjacent three harmonic energies of each time-domain pitch period candidate value of each frame; Step S1609, using the multiple time-domain pitch period candidate values of each frame, the corresponding weighting factors and the normalized autocorrelation values, performing dynamic programming to determine the cost function of the trajectory path corresponding to each time-domain pitch period candidate value of each frame; Step S1610, selecting the trajectory path with the minimum cost function from the cost functions of the trajectory paths corresponding to all time-domain pitch period candidate values of each frame as the time-domain pitch period change trajectory of the corresponding frame, to obtain the time-domain pitch period candidate values passed through by the time-domain pitch period change trajectory of each frame; Step S1611, taking the time-domain pitch period candidate values passed through by the time-domain pitch period change trajectory of each frame as the time-domain pitch period of the corresponding frame, to calculate the pitch parameter in the speech signal parameter corresponding to each frame.

4. The narrow-band speech compression method according to any one of claims 1 to 3, characterized in that, In the step S2, the specific implementation process of the quantization processing includes: Step S21, performing 1-bit representation on the speech activation parameter in the speech signal parameter corresponding to each frame; Step S22, performing three-level residual vector quantization on the multiple-dimensional line spectrum frequency parameter in the speech signal parameter corresponding to each frame, and using 8, 8 and 7 three-level bit numbers for representation; Step S23, performing 3-bit vector quantization on the spectral amplitude parameter and the voiced and unvoiced parameters of each subband in the speech signal parameter corresponding to each frame; Step S24, taking the logarithmic value of the energy parameter in the speech signal parameter corresponding to each frame, and performing 6-bit non-uniform scalar quantization in the range of 10~77 dB; Step S25, performing 7-bit non-uniform scalar quantization on the pitch parameter in the speech signal parameter corresponding to each frame in the range of 54~444 Hz.

5. The narrowband speech compression method of claim 4, wherein, In the step S3, the specific process of the non-equal weight forward error correction coding includes: Step S31, using 1 bit, performing parity check on the 1st~9th bits in the speech coding quantization bits corresponding to each frame to output 1 bit; Step S32, performing (15, 11) BCH coding on the 14th~24th in the speech coding quantization bits corresponding to each frame to output 14 bits; Step S33, performing (24, 12) systematic Golay code coding on the 2nd~13th in the speech coding quantization bits corresponding to each frame to output 24 bits; Step S34, performing (24, 12) systematic Golay code coding on the 25th, 26th, 27th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th and 41st bits in the speech coding quantization bits corresponding to each frame to output 24 bits.

6. A narrowband speech compression system suitable for use in a PDT system, characterized by The narrowband speech compression system includes: An encoding processing module configured to perform frame division and encoding processing on an input speech signal to obtain speech signal parameters corresponding to each frame; The quantization processing module is configured to perform quantization processing on the speech signal parameter corresponding to each frame to obtain speech coding quantization bits corresponding to each frame; The non-equal weight forward error correction coding module is configured to perform non-equal weight forward error correction coding on part of the speech coding quantization bits corresponding to each frame to obtain forward error correction coding bits corresponding to each frame; The non-equal weight forward error correction decoding module is configured to perform forward error correction decoding on the forward error correction coding bits corresponding to each frame to obtain speech decoding quantization bits corresponding to each frame; The inverse quantization module is configured to perform inverse quantization on the speech decoding quantization bits corresponding to each frame to synthesize a mixed excitation signal corresponding to each frame; The post-filtering module is configured to sequentially perform synthesis filtering and post-filtering processing on the mixed excitation signal corresponding to each frame to synthesize a compressed narrowband speech signal.

7. The narrow-band speech compression system of claim 6, wherein, The encoding processing module includes: The obtaining submodule is configured to obtain a speech activation parameter in the speech signal parameter corresponding to each frame; The spectral frequency conversion submodule is configured to sequentially perform linear prediction analysis and spectral frequency conversion on the input speech signal of each frame and multiple sampling points to obtain a residual signal corresponding to each frame and a multi-dimensional line spectrum frequency parameter in the speech signal parameter; The extraction submodule is configured to sequentially perform frequency domain transformation and spectral envelope extraction on the residual signal corresponding to each frame to obtain a spectral amplitude parameter in the speech signal parameter corresponding to each frame; The calculation submodule is configured to calculate the square sum of each sampling point of the residual signal corresponding to each frame to obtain an energy parameter in the speech signal parameter corresponding to each frame; The band-pass filtering submodule is configured to perform different band-pass filtering on the residual signal corresponding to each frame to determine a voiced and unvoiced parameter of each subband in the speech signal parameter corresponding to each frame; The determination submodule is configured to determine a fundamental frequency parameter in the speech signal parameter corresponding to each frame by using the input speech signal of each frame and multiple sampling points.

8. The narrow-band speech compression system of claim 7, wherein, The determination submodule includes: The filtering subunit is configured to sequentially perform second-order inverse filtering, low-pass filtering and time-domain pitch period traversal on the input speech signal of each frame and multiple sampling points to obtain a time-domain signal corresponding to each time-domain pitch period candidate value in each frame; The autocorrelation processing subunit is configured to perform autocorrelation processing on the time-domain signal corresponding to each time-domain pitch period candidate value in each frame to obtain an autocorrelation value corresponding to each time-domain pitch period candidate value in each frame and a corresponding frame energy value; The sorting subunit is configured to sort the autocorrelation values corresponding to all the time-domain pitch period optional values in each frame from large to small, and take the time-domain pitch period optional value corresponding to each of the first K autocorrelation values as the K time-domain pitch period candidate values of the corresponding frame. ​​ The first normalization subunit is configured to normalize the autocorrelation value corresponding to each time-domain pitch period candidate value of the corresponding frame by using the frame energy value; The discrete Fourier transform subunit is configured to perform discrete Fourier transform on the time-domain signal corresponding to each frame after sequentially passing through the second-order inverse filtering and the low-pass filtering to obtain a digital frequency point corresponding to each time-domain pitch period candidate value of each frame; In which, the digital frequency point corresponding to each time-domain pitch period candidate value of each frame is taken as a current digital frequency point, and the maximum energy corresponding to the current digital frequency point and the left and right digital frequency points of the current digital frequency point is taken as the harmonic energy corresponding to each time-domain pitch period candidate value of each frame. The first calculation subunit is configured to calculate the maximum value of the sum of the adjacent three harmonic energies of each time-domain pitch period candidate value of each frame by using the harmonic energy corresponding to each time-domain pitch period candidate value of each frame; The second normalization subunit is configured to select the maximum value from the maximum values of the sum of the adjacent three harmonic energies of each time-domain pitch period candidate value of each frame as a weighting factor to normalize the maximum value of the sum of the adjacent three harmonic energies of each time-domain pitch period candidate value of each frame; The dynamic programming subunit is configured to perform dynamic programming by using the multiple time-domain pitch period candidate values, the corresponding weighting factors and the normalized autocorrelation values of each frame to determine a cost function of a track path corresponding to each time-domain pitch period candidate value of each frame; The selection subunit is configured to select a track path with the minimum cost function from the cost functions of the track paths corresponding to all the time-domain pitch period candidate values of each frame as a time-domain pitch period variation track of the corresponding frame to obtain the time-domain pitch period candidate values passed through by the time-domain pitch period variation track of each frame; The second calculation subunit is configured to take the time-domain pitch period candidate values passed through by the time-domain pitch period variation track of each frame as the time-domain pitch period of the corresponding frame to calculate a fundamental frequency parameter in the speech signal parameters corresponding to each frame.

9. The narrow-band speech compression system according to any one of claims 6 to 8, characterized in that The quantization processing module comprises: The bit representation sub-module is configured to perform 1-bit representation on the voice activity parameter in the speech signal parameters corresponding to each frame; The three-level residual vector quantization sub-module is configured to perform three-level residual vector quantization on the multiple-dimensional line spectrum frequency parameters in the speech signal parameters corresponding to each frame and adopt 8, 8 and 7 three-level bit numbers for representation; The 3-bit vector quantization sub-module is configured to perform 3-bit vector quantization on the spectral amplitude parameters and the voiced and unvoiced parameters of each subband in the speech signal parameters corresponding to each frame; The first non-uniform scalar quantization sub-module is configured to take the logarithmic value of the energy parameter in the speech signal parameters corresponding to each frame and perform 6-bit non-uniform scalar quantization in the range of 10-77 dB; The second non-uniform scalar quantization sub-module is configured to perform 7-bit non-uniform scalar quantization on the fundamental frequency parameter in the speech signal parameters corresponding to each frame in the range of 54-444 Hz.

10. The narrowband speech compression system of claim 9, wherein, The non-equal weight forward error correction encoding module comprises: The parity check sub-module is configured to perform parity check on the 1st-9th bits in the speech coding quantization bits corresponding to each frame by using 1 bit to output 1 bit; The BCH encoding sub-module is configured to perform (15, 11) BCH encoding on the 14th-24th bits in the speech coding quantization bits corresponding to each frame to output 14 bits; The first systematic Golay code encoding sub-module is configured to perform (24, 12) systematic Golay code encoding on the 2nd-13th bits in the speech coding quantization bits corresponding to each frame to output 24 bits; The second system Golay code encoding submodule is configured to perform (24, 12) system Golay code encoding on the 25th, 26th, 27th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th and 41st bits in the speech coding quantization bits corresponding to each frame to output 24 bits.

Citation Information

Patent Citations

  • 1.2kb / s low-rate speech encoding and decoding method based on mixed excitation linear prediction MELP

    CN105118513A

  • Beyond-the-horizon ultra-short wave radio station based on BP neural network and communication method thereof

    CN118413250A