Sub-band unvoiced and voiced sound parameter quantization method and system based on double-resolution codebook
By constructing a subband voiced/unvoiced parameter quantization method based on a dual-resolution codebook, and utilizing maximum likelihood search and likelihood matrix, the problem of insufficient expressive power of speech coding at low bit rates is solved, and more natural and clear speech synthesis is achieved.
Patent Information
- Application Number
- CN202511391256.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-12-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech coding technologies, under low bitrate conditions, suffer from limited expressive power in subband voiced/unvoiced parameter quantization, leading to incorrect excitation selection, unnatural speech synthesis, and distorted sound quality. They also fail to effectively exploit the temporal correlation of speech sequences.
A dual-resolution codebook approach is adopted to construct a first-level codebook and a second-level codebook. The maximum likelihood search strategy is used, combined with the voiced/unvoiced tone pattern of the previous frame subband and the likelihood matrix, to improve the precision of parameter quantization and coding accuracy.
Without increasing the bitrate, the naturalness and clarity of speech synthesis are improved, and the ability to express speech features and temporal continuity are enhanced.
Smart Images

Figure CN121054007A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech coding technology, and in particular relates to a method and system for quantizing subband voiced / unvoiced tone parameters based on a dual-resolution codebook. Background Technology
[0002] Voice coding technology plays a crucial role in modern communication systems, voice recording and playback devices, and numerous consumer electronics products. To achieve high-quality voice transmission under limited bandwidth conditions, a series of voice compression coding standards have been introduced. These standards exhibit a trend of continuously improving compression efficiency, decreasing coding rate, and increasing synthesized voice quality. Especially in the field of low-to-medium bit-rate voice coding, they have become core technologies in scenarios with high requirements for transmission efficiency and voice quality, such as wireless communication, secure communication, and underwater acoustic communication.
[0003] In the structural design of this type of encoder, the speech signal is typically divided into multiple sub-bands, and periodic and aperiodic excitation models are performed for each sub-band. To this end, the speech coding system needs to determine the voiced or unvoiced state of each sub-band and generate a set of features called Bandpass Voicing Coefficients (BPVCs). These parameters directly determine the type of speech excitation signal and are one of the key factors in the quality of synthesized speech. To effectively compress these parameters, the industry commonly uses codebook quantization: typical sub-band voiced / unvoiced patterns are extracted from the speech training set, and a codebook is constructed using clustering algorithms; subsequently, during the encoding stage, the optimal codeword is selected for index transmission using criteria such as minimum weighted mean square error (WMSE), and the corresponding parameters are restored at the decoding end. While this process has a significant advantage in reducing the number of transmitted bits, it also introduces limitations on the expressive power of the parameters.
[0004] As speech coding technology evolves towards higher fidelity and stronger perceptual optimization, traditional subband voiced / unvoiced parameter quantization schemes face numerous challenges. For example, under low bitrate conditions, limited by codebook capacity, systems often retain only a few typical voiced / unvoiced modes for encoding. This results in some complex or transitional phoneme states in real speech not being fully expressed, leading to problems such as excitation misselection, unstable speech waveforms, and sound quality distortion, reducing the naturalness and intelligibility of speech synthesis. Furthermore, existing methods rely heavily on current frame information in codebook construction and parameter matching, failing to effectively mine the temporal correlation of speech sequences, thus limiting their ability to model speech continuity and semantic integrity in real-world communication scenarios.
[0005] Therefore, given the increasing emphasis on clarity, naturalness, and low latency in current speech coding systems, how to further improve the quantization accuracy of subband voiced / unvoiced parameters and enhance the model's ability to express speech features without significantly increasing the bit rate has become one of the core issues that urgently need to be addressed in this field. Summary of the Invention
[0006] This invention provides a method and system for subband voiced / unvoiced parameter quantization based on a dual-resolution codebook, addressing the limitations in expressive power and excitation selection errors caused by subband voiced / unvoiced parameter quantization in existing speech coding techniques. This invention aims to fully utilize the correlation of subband voiced / unvoiced parameters by introducing a dual-resolution codebook structure based on maximum likelihood search. While retaining the traditional WMSE training codebook as the primary codebook, a secondary codebook is constructed according to the occurrence probability of subband voiced / unvoiced modes. Based on this, the subband voiced / unvoiced modes of the previous frame are used as observation information, combined with all codewords in the secondary codebook of the current frame, to construct a conditional probability model. The subband voiced / unvoiced modes are then selected based on the maximum likelihood criterion, improving the coding accuracy and perceptual quality of subband voiced / unvoiced parameters.
[0007] The technical solution provided by this invention is as follows: A subband voiced / unvoiced tone parameter quantization method based on dual-resolution codebooks includes: A first-level codebook for quantizing voiced and unvoiced sounds is constructed based on the speech training set. This first-level codebook contains several first-level codewords. The subband voiced / unvoiced parameters of each speech frame in the speech training set are quantized using the first-level codebook, and the subband voiced / unvoiced modes of the same codeword are used as the cell vector of that codeword. Statistical analysis was performed on the frequency of the cell vectors corresponding to each first-level codeword. The top N cell vectors with the highest frequency were selected and their corresponding second-level codewords were set to construct a locally high-resolution second-level codebook. The conditional transition probabilities of the first-level codewords in the previous frame and the second-level codewords in the next frame are statistically obtained in adjacent speech frames, and a likelihood matrix / table is constructed. The conditional transition probabilities reflect the probability that the next frame is quantized into each of the second-level codewords when the previous frame is quantized into a certain level codeword in adjacent speech frames. The encoding end uses the first-level codebook to quantize the sub-band voiced / unvoiced parameters of the target speech to be quantized, obtains the corresponding first-level codewords for each speech frame, and transmits the codeword index to the decoding end. The decoding end obtains the corresponding first-level codeword based on the received codeword index. For adjacent speech frames in the target speech, it searches the likelihood matrix / table based on the first-level codeword of the previous frame. Among the second-level codewords corresponding to the first-level codeword of the subsequent frame, the second-level codeword with the highest conditional transition probability is used as the reconstructed value of the voiced / unvoiced tone parameter of the subsequent frame subband.
[0008] Furthermore, the conditional transition probability is statistically represented by the frequency of occurrence of the codeword; the higher the frequency of occurrence, the greater the corresponding conditional transition probability.
[0009] Furthermore, if the preceding frame in an adjacent speech frame is the starting frame, it is assumed that the first-level codeword of the preceding frame of the starting frame is the same as the first-level codeword of the starting frame.
[0010] Furthermore, the construction of the first-level codebook includes: reading speech from the speech training set frame by frame, performing subband voiced / unvoiced analysis on the speech frames to form a subband voiced / unvoiced parameter set; and using vector clustering technology to perform codebook clustering on the subband voiced / unvoiced parameter set to generate a first-level codebook containing several codewords.
[0011] Furthermore, the subband voiced / unvoiced parameter quantization of the target speech using the first-level codebook includes: performing subband voiced / unvoiced analysis on each speech frame in the target speech, obtaining the corresponding subband voiced / unvoiced parameter vectors, calculating the error between the parameter vector and the cluster center vector corresponding to each first-level codeword in the first-level codebook, and selecting the codeword corresponding to the cluster center vector with the smallest error as the first-level codeword of the subband voiced / unvoiced speech frame.
[0012] Preferably, the error is calculated using the least weighted mean square error method.
[0013] Furthermore, each speech frame in the speech training set is 25ms long and corresponds to 200 sampling points; each speech frame is divided into 5 sub-bands with frequency ranges of 0~500Hz, 500~1000Hz, 1000~2000Hz, 2000~3000Hz, and 3000~4000Hz.
[0014] A subband voiced / unvoiced tone parameter quantization system based on the above method includes a parameter extraction module, a codebook generation module, a likelihood modeling module, a quantization encoding module, and a decoding reconstruction module; The parameter extraction module is used to divide the speech into frames and subbands, and to extract the subband voiced / unvoiced parameter vectors of the speech frames; the codebook generation module is used to construct a first-level codebook for voiced / unvoiced parameter quantization based on the speech training set, and to construct a corresponding high-resolution second-level codebook through statistical analysis; the likelihood modeling module is used to construct a likelihood matrix / table through statistical analysis; the quantization encoding module is used to quantize the subband voiced / unvoiced parameters of the target speech based on the first-level codebook, and output the codeword index corresponding to the first-level codeword of each speech frame of the target speech; the decoding and reconstruction module obtains the corresponding first-level codeword based on the received codeword index, and determines the second-level codeword corresponding to each speech frame in the target speech by searching the likelihood matrix / table, and uses it as the reconstructed value of the subband voiced / unvoiced parameters.
[0015] An electronic device includes a memory, a processor, and a computer program stored in the memory and running on the memory. When the processor executes the program, it implements the steps of the subband voiced / unvoiced tone parameter quantization method as described above.
[0016] A computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the subband voiced / unvoiced tone parameter quantization method as described above.
[0017] This invention fully utilizes the inter-frame correlation of subband voiced / unvoiced parameters to further construct a high-resolution secondary codebook with context-aware capabilities based on the primary codebook. It also introduces a maximum likelihood search strategy, improving the representation accuracy of subband voiced / unvoiced parameters without significantly increasing additional bit rate overhead. By using the previous frame's subband voiced / unvoiced pattern as observation information and combining it with the conditional probabilities in the likelihood matrix to achieve optimal selection of secondary codewords, it effectively compensates for the shortcomings of traditional methods in modeling temporal evolution features, enhances the encoder's adaptability to dynamic speech changes, and thus significantly improves the naturalness, clarity, and intelligibility of speech synthesis. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0019] Figure 1 This is a schematic diagram of the subband voiced / unvoiced tone parameter quantization framework based on a dual-resolution codebook provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of the present invention.
[0021] In existing technologies, the quantization methods for subband voiced / unvoiced sounds mainly include the following steps: 1. Codebook generation process 1.1 Prepare a speech training set, read the speech frame by frame, perform sub-band voiced / unvoiced analysis on the speech frames, and form a sub-band voiced / unvoiced parameter set; 1.2 Using vector clustering technology, codebook training is performed on the subband voiced / unvoiced tone parameter set in step 1.1 to generate a subband voiced / unvoiced tone codebook C with a codebook size of M.
[0022] 2. Parameter quantization process 2.1 Prepare the test speech, read the speech frame by frame, perform sub-band voiced / unvoiced analysis on the speech frame, and obtain the sub-band voiced / unvoiced parameters of the current frame; 2.2. Compare the subband voiced / unvoiced parameters of the current frame in step 2.1 with the weighted mean square error of all codewords in the subband voiced / unvoiced codebook C, select the codeword with the smallest error, and output the corresponding codeword index as the encoding result. 2.3 The decoding end uses the codeword index obtained from the encoding result to extract the corresponding codeword from the subband voiced / unvoiced codebook C, and uses it as the reconstruction value of the subband voiced / unvoiced parameters of the current frame.
[0023] This low-bit-rate speech coding method uses only a single static codebook in the subband voiced / unvoiced sound modeling process and uses the minimum weighted mean square error as the quantization criterion. Although it has a certain compression efficiency, it is limited by the codebook capacity and has limited ability to express complex speech patterns. It is very easy to cause problems such as excitation selection error, unstable sound quality, and unnatural speech. In particular, it is difficult to guarantee the perceptual continuity and clarity of speech in real communication scenarios.
[0024] Therefore, this embodiment provides a sub-band voiced / unvoiced parameter quantization method based on a dual-resolution codebook, which improves the quantization accuracy of sub-band voiced / unvoiced parameters in speech coding and is suitable for voiced / unvoiced excitation modeling in low-bit-rate speech coding systems. Figure 1 As shown, the main technical contents of this method include: 1. Codebook generation process 1.1 Prepare a speech training set. Read the speech frame by frame and perform sub-band voiced / unvoiced analysis on each frame to form a sub-band voiced / unvoiced parameter set. The speech training set uses an 8kHz sampling rate and a 16-bit quantization format. Each subframe is 25ms long, corresponding to 200 sampling points. Each subframe is divided into 5 sub-bands with frequency ranges of 0~500Hz, 500~1000Hz, 1000~2000Hz, 2000~3000Hz, and 3000~4000Hz, and the voiced / unvoiced parameters of each sub-band are extracted sequentially. The encoder uses a two-frame joint quantization structure, with each pair of subframes as an analysis unit, extracting its sub-band voiced / unvoiced parameters to form a 10-dimensional feature vector (composed of the concatenation of the voiced / unvoiced results from the two sub-frames). The above 10-dimensional sub-band voiced / unvoiced parameters are extracted pair by pair to form the sub-band voiced / unvoiced parameter set. 1.2. Using vector clustering technology, codebook training is performed on the subband voiced / unvoiced parameter set from step 1.1 to generate a first-level codebook C1 for subband voiced / unvoiced sounds, with a codebook size of M. The LBG algorithm is used to train the codebook on this set, and iterative clustering is performed under the weighted mean square error criterion to finally generate the first-level codebook C1, where each codeword corresponds to a 10-dimensional joint pattern of subband voiced / unvoiced sounds.
[0025] 1.3. Using the first-level codebook C1 generated in step 1.2, the sub-band voiced / unvoiced consonant parameter set is quantized sequentially. If a certain sub-band voiced / unvoiced consonant pattern is quantized into codeword V in the first-level codebook C1, then the pattern is regarded as the cell vector of codeword V (the same codeword usually corresponds to multiple sub-band voiced / unvoiced consonant patterns; in other words, multiple sub-band voiced / unvoiced consonant patterns may be quantized and encoded into the same codeword). For each codeword in the first-level codebook C1, the N cell vectors with the highest occurrence probability (frequency) are selected from its cell and the corresponding codewords are set to construct the second-level codebook C2. If there are fewer than N cell vectors, all cell vectors are used to construct the second-level codebook C2 of the codeword. 1.4. Treat each vector in the second-level codebook C2 in step 1.3 as a model parameter, take the voiced / unvoiced mode of the previous speech frame (corresponding to the codeword in the first-level codebook C1) as the observation data of the current speech frame, calculate the conditional probability of M observation vectors, and store them as a likelihood matrix R.
[0026] In simple terms, it means that if the previous speech frame is quantized into each first-level codeword, then the probability of the current speech frame being encoded into each second-level codeword during second-level quantization is obtained. The probability can be represented by the frequency of occurrence.
[0027] For example: Suppose that the first-level codebook C1 constructed from the speech training set contains three first-level codewords V1~V3. The second-level codewords corresponding to the V1 cell cavity vector are V101 and V102. The second-level codewords corresponding to the V2 cell cavity vector are V201 and V202. By analyzing the quantization encoding of the speech training set, we can see that: When the first-level codeword of the previous frame is V1 in two adjacent speech frames, the number of times the first-level codeword of the next frame is V1 and the second-level codewords are V101 and V102 are 30 and 70 respectively (i.e., the occurrence probabilities are 0.30 and 0.70 respectively), and the number of times the first-level codeword of the next frame is V2 and the second-level codewords are V201 and V202 are 65 and 35 respectively (i.e., the occurrence probabilities are 0.65 and 0.35 respectively). When the first-level codeword of the previous frame is V2 in two adjacent speech frames, the number of times the first-level codeword of the next frame is V1 and the second-level codewords are V101 and V102 are 85 and 15 respectively (i.e., the occurrence probabilities are 0.85 and 0.15 respectively), and the number of times the first-level codeword of the next frame is V2 and the second-level codewords are V201 and V202 are 40 and 60 respectively (i.e., the occurrence probabilities are 0.40 and 0.60 respectively). The above statistics are recorded as a likelihood matrix / table for storage. Rows in the likelihood matrix / table correspond to first-level codewords, columns to second-level codewords, and the matrix / table elements are the corresponding probability values (or frequency values). For example, the above statistics can be recorded as the following likelihood table:
[0028] Its corresponding likelihood matrix can be expressed as:
[0029] 2. Parameter quantization process 2.1 Read the speech to be quantized frame by frame, perform sub-band voiced / unvoiced analysis on the speech frame, and obtain the sub-band voiced / unvoiced parameters of the current frame; 2.2 Compare the subband voiced / unvoiced parameters of the current frame in step 2.2 with the weighted mean square error of all codewords in the first-level codebook C1, select the codeword with the smallest error, and output the index corresponding to the codeword as the encoding result; 2.3 The decoding end finds the corresponding codeword from the first-level codebook C1 according to the codeword index transmitted in step 2.2, and at the same time obtains the first-level codeword corresponding to the previous frame. If the current frame is the initial frame, the first-level codeword of the previous frame is regarded as the same as the first-level codeword of the current frame. 2.4. Based on the likelihood matrix R, find the second-level codeword with the highest likelihood probability in the first-level codeword encoding case of the previous frame among the second-level codewords corresponding to the first-level codewords in the current frame. Use this second-level codeword as the reconstruction value of the subband voiced / unvoiced parameters in the current frame, which will be used for subsequent excitation modeling and speech synthesis.
[0030] Based on the above method, this embodiment also provides a subband voiced / unvoiced parameter quantization system based on a dual-resolution codebook, to meet the actual needs of accurate modeling of subband voiced / unvoiced parameters in low-to-medium bit-rate speech coding systems, improve speech synthesis quality, enhance speech feature representation capabilities, and optimize perceptual consistency between consecutive frames. Figure 1 As shown, the system can be divided into five core units: parameter extraction module, codebook generation module, likelihood modeling module, real-time quantization and encoding module, and decoding and reconstruction module. These units work together to form a closed-loop process.
[0031] The parameter extraction module is first responsible for extracting voiced and unvoiced features from the original speech training data. The system divides the input speech signal into frames and further divides each frame into sub-bands. Through joint extraction of multiple frames, a high-dimensional parameter vector reflecting the voiced and unvoiced features is formed, which serves as the basis for subsequent codebook generation and probabilistic modeling.
[0032] The codebook generation module takes the set of voiced / unvoiced parameters output by the parameter extraction module as input. First, it performs vector clustering using a weighted mean square error criterion to generate a first-level codebook, which covers typical voiced / unvoiced joint patterns. Based on this, the system also constructs a second-level codebook for each first-level codeword. By analyzing the cell vector frequency of the first-level codewords, highly representative sub-patterns are selected to construct a local high-resolution codebook, thus compensating for the shortcomings of traditional quantization methods in representing local voiced / unvoiced variations.
[0033] The likelihood modeling module takes temporal correlation as its starting point, statistically analyzing the codeword transition relationships between adjacent speech frames in the training set to construct a conditional probability matrix between the first-level codewords of the previous frame and the second-level codewords of the next frame. This likelihood matrix is obtained through long-term offline statistical analysis and can serve as a decision-making basis for the real-time encoding stage. Its design logic implicitly models the temporal continuity of the speech signal, which helps to enhance the auditory smoothness and naturalness of the encoding results.
[0034] The real-time quantization coding module is mainly responsible for frame-level analysis and parameter encoding of the speech to be quantized. It compares the subband voiced / unvoiced parameters of the current frame with the weighted mean square error of all codewords in the first-level codebook, selects the codeword with the smallest error as the optimal match, and outputs the index of the codeword as the encoding result. The module realizes the process of extracting and compressing speech feature parameters from analog signals.
[0035] The decoding and reconstruction module is responsible for reconstructing the corresponding subband voiced / unvoiced parameters at the receiving end based on the index information. The receiving end extracts the corresponding codeword content using the transmitted first-level codeword index and selects the optimal second-level codeword from the likelihood matrix based on the previous frame's subband voiced / unvoiced pattern to complete parameter reconstruction, ensuring contextual consistency and auditory coherence. The finally reconstructed subband voiced / unvoiced parameters will then drive the excitation generator and speech synthesis module, outputting natural and fluent synthesized speech.
[0036] The above system can execute the subband voiced / unvoiced tone parameter quantization method based on dual-resolution codebook described in the embodiments. It has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in this part, please refer to the subband voiced / unvoiced tone parameter quantization method based on dual-resolution codebook provided in the embodiments.
[0037] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0038] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A subband voiced / unvoiced tone parameter quantization method based on a dual-resolution codebook, characterized in that, include: A first-level codebook for quantizing voiced and unvoiced sounds is constructed based on the speech training set. This first-level codebook contains several first-level codewords. The subband voiced / unvoiced parameters of each speech frame in the speech training set are quantized using the first-level codebook, and the subband voiced / unvoiced modes of the same codeword are used as the cell vector of that codeword. Statistical analysis was performed on the frequency of the cell vectors corresponding to each first-level codeword. The top N cell vectors with the highest frequency were selected and their corresponding second-level codewords were set to construct a locally high-resolution second-level codebook. The conditional transition probabilities of the first-level codewords in the previous frame and the second-level codewords in the next frame are statistically obtained in adjacent speech frames, and a likelihood matrix / table is constructed. The conditional transition probabilities reflect the probability that the next frame is quantized into each of the second-level codewords when the previous frame is quantized into a certain level codeword in adjacent speech frames. The encoding end uses the first-level codebook to quantize the sub-band voiced / unvoiced parameters of the target speech to be quantized, obtains the corresponding first-level codewords for each speech frame, and transmits the codeword index to the decoding end. The decoding end obtains the corresponding first-level codeword based on the received codeword index. For adjacent speech frames in the target speech, it searches the likelihood matrix / table based on the first-level codeword of the previous frame. Among the second-level codewords corresponding to the first-level codeword of the subsequent frame, the second-level codeword with the highest conditional transition probability is used as the reconstructed value of the voiced / unvoiced tone parameter of the subsequent frame subband.
2. The sub-band voiced / unvoiced tone parameter quantization method as described in claim 1, characterized in that, The conditional transition probability is statistically represented by the frequency of occurrence of the codeword; the higher the frequency of occurrence, the greater the corresponding conditional transition probability.
3. The sub-band voiced / unvoiced tone parameter quantization method as described in claim 1, characterized in that, If the preceding frame in an adjacent speech frame is the starting frame, then it is assumed that the first-level codeword of the preceding frame is the same as the first-level codeword of the starting frame.
4. The sub-band voiced / unvoiced tone parameter quantization method as described in claim 1, characterized in that, The construction of the first-level codebook includes: reading speech from the speech training set frame by frame, performing subband voiced / unvoiced analysis on the speech frames to form a subband voiced / unvoiced parameter set; and using vector clustering technology to perform codebook clustering on the subband voiced / unvoiced parameter set to generate a first-level codebook containing several codewords.
5. The sub-band voiced / unvoiced tone parameter quantization method as described in claim 4, characterized in that, The subband voiced / unvoiced parameter quantization of target speech using a first-level codebook includes: performing subband voiced / unvoiced analysis on each speech frame in the target speech, obtaining the corresponding subband voiced / unvoiced parameter vectors, calculating the error between the parameter vector and the cluster center vector corresponding to each first-level codeword in the first-level codebook, and selecting the codeword corresponding to the cluster center vector with the smallest error as the first-level codeword of the subband voiced / unvoiced speech frame.
6. The sub-band voiced / unvoiced tone parameter quantization method as described in claim 5, characterized in that, The error is calculated using the least weighted mean square error method.
7. The sub-band voiced / unvoiced tone parameter quantization method as described in claim 1, characterized in that, The speech training set has a speech frame length of 25ms and corresponds to 200 sampling points; each speech frame is divided into 5 sub-bands with frequency ranges of 0~500Hz, 500~1000Hz, 1000~2000Hz, 2000~3000Hz, and 3000~4000Hz.
8. A sub-band voiced / unvoiced tone parameter quantization system based on the method of any one of claims 1 to 7, characterized in that, It includes a parameter extraction module, a codebook generation module, a likelihood modeling module, a quantization and encoding module, and a decoding and reconstruction module; The parameter extraction module is used to divide the speech into frames and subbands, and to extract the subband voiced / unvoiced parameter vectors of the speech frames; the codebook generation module is used to construct a first-level codebook for voiced / unvoiced parameter quantization based on the speech training set, and to construct a corresponding high-resolution second-level codebook through statistical analysis; the likelihood modeling module is used to construct a likelihood matrix / table through statistical analysis; the quantization encoding module is used to quantize the subband voiced / unvoiced parameters of the target speech based on the first-level codebook, and output the codeword index corresponding to the first-level codeword of each speech frame of the target speech; the decoding and reconstruction module obtains the corresponding first-level codeword based on the received codeword index, and determines the second-level codeword corresponding to each speech frame in the target speech by searching the likelihood matrix / table, and uses it as the reconstructed value of the subband voiced / unvoiced parameters.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running thereon, characterized in that, When the processor executes the program, it implements the steps of the subband voiced / unvoiced tone parameter quantization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the subband voiced / unvoiced tone parameter quantization method as described in any one of claims 1 to 7.