Voice reconstruction method and device, storage medium and program product

Through the multi-frame rate granular feature extraction and information density screening mechanism, the frame rate of the neural speech codec is dynamically adjusted, and the redundant encoding problem caused by fixed frame rates is solved, achieving efficient and low-latency speech reconstruction.

CN120356481APending Publication Date: 2025-07-22AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400460.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

Smart Images

  • Figure CN120356481A_ABST
    Figure CN120356481A_ABST
Patent Text Reader

Abstract

The invention discloses a voice reconstruction method and device, a storage medium and a program product, and relates to the technical field of audio processing, and the method comprises the steps: obtaining to-be-reconstructed original voice data, and extracting the granularity voice frame features of the original voice data at each frame rate; calculating the voice information density of each granularity voice frame feature in each time window; screening granularity frame feature fragments according with a preset granularity proportion from each granularity voice frame feature according to the voice information density; the granularity proportion predefines the time proportion of the fragment duration of each granularity frame feature fragment in a unit time period; and performing voice reconstruction based on each granularity frame feature fragment to obtain corresponding reconstructed voice data. Therefore, through voice information density analysis, dynamic frame segment distribution of a variable frame rate is realized, high-sound-quality reconstruction is ensured, the coding complexity and delay are remarkably reduced, and the real-time performance and the self-adaptive capability of the whole system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technologies, and particularly to a voice reconstruction method, device, storage medium, and program product. Background Art

[0002] In recent years, neural networks have been widely applied in the field of audio and speech coding and decoding, thereby enabling the reconstruction of high-fidelity audio. These models typically adopt an end-to-end training method and mainly consist of three components: an encoder that compresses the input signal into a compact representation, a quantization module that discretizes these representations, and a decoder that is responsible for reconstructing the audio from the quantized vectors. Relying on these successful experiences, the discrete token generation modeling method has been successfully extended to the speech field, giving rise to speech generation models such as AudioLM and VALL-E.

[0003] However, existing neural speech codecs typically generate token sequences at a relatively high temporal rate. For example, a system like EnCodec generates up to 75 frames per second, while BPE's text tokenization typically generates only 3 to 5 tokens per second. This significant difference in token length results in low autoregressive prediction efficiency, thereby causing performance degradation and increased latency in real-time applications. To address this issue, in addition to attempts to directly improve the efficiency of downstream speech generation models, some recent studies have begun to explore low-frame-rate neural audio codecs to fundamentally tackle this challenge.

[0004] Despite the related research, most neural speech codecs still operate as constant bitrate (CBR) systems, which may lead to redundant coding, especially in segments such as silence, which require a lower coding rate compared to dynamic regions. This temporal redundancy undermines the goal of reducing the total bitrate consumption and sequence length. Although previous work has preliminarily introduced a variable bitrate (VBR) strategy based on information density, it still belongs to the constant frame rate (CFR) framework and fails to shorten the overall sequence length.

[0005] In response to the above problems, the industry has not yet proposed a better solution. Summary of the Invention

[0006] The present application provides a voice reconstruction method, electronic device, storage medium, and program product, which are used to at least solve the problems of redundant coding and overly long sequence length caused by a fixed frame rate in traditional neural speech codecs.

[0007] In a first aspect, an embodiment of the present application provides a voice reconstruction method, including: obtaining original voice data to be reconstructed, and extracting granular voice frame features of the original voice data at each frame rate; calculating the voice information density of each granular voice frame feature within each time window; screening granular frame feature segments that meet a preset granularity ratio from each granular voice frame feature; the granularity ratio predefines the time ratio of the segment duration of each granular frame feature segment in a unit period; performing voice reconstruction based on each granular frame feature segment to obtain corresponding reconstructed voice data.

[0008] In a second aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the voice reconstruction method according to any embodiment of the present application.

[0009] In a third aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the voice reconstruction method according to any embodiment of the present application are implemented.

[0010] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the training method of the multi-modal speaker anti-spoofing model according to any embodiment of the present application are implemented.

[0011] The beneficial effects of the embodiments of the present application are as follows:

[0012] By introducing multi-frame rate granular feature extraction and an information density-based dynamic selection mechanism, accurate discrimination and resource allocation for different segments of audio during the encoding process are achieved. It can effectively identify low-information-density regions and screen frame segments according to the predefined granularity ratio. Thus, low-frame-rate segments are selected in silent or low-information areas to greatly reduce redundant encoding, and high-frame-rate segments are selected in high-information-density regions to capture sufficient voice details, realizing variable frame-rate dynamic frame segment allocation. While ensuring high-quality voice reconstruction, the encoding complexity and latency are significantly reduced, and the real-time performance and adaptive ability of the overall system are improved. Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 Shows a flowchart of an example of a voice reconstruction method according to an embodiment of the present application;

[0015] Figure 2 Shows an operation flowchart of an example of screening granularity frame feature segments according to an embodiment of the present application;

[0016] Figure 3 Shows an operation flowchart of an example of voice reconstruction based on each of the granularity frame feature segments according to an embodiment of the present application;

[0017] Figure 4 Shows a schematic diagram of the working principle of an example of a time-flexible coding framework according to an embodiment of the present application;

[0018] Figure 5 Shows a schematic diagram of the simulation effect of an example of VFR allocation of time-flexible coding at different granularity ratios according to an embodiment of the present application;

[0019] Figure 6 Shows a schematic diagram of the simulation effect of an example of the performance of DAC+TFC in the VFR mode;

[0020] Figure 7 Is a schematic structural diagram of an embodiment of an electronic device of the present application. Detailed implementation manners

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that in the current related technologies, most neural voice codecs implement bitrate adjustment at a fixed frame rate (Constant Frame Rate, CFR) through an intra-frame mechanism (such as codebook discard). However, voice segments inherently have time-varying information density (such as silent intervals and voiced segments), which makes CFR not optimal in terms of bitrate and token sequence length, thus hindering the efficiency of real-time applications.

[0023] Specifically, the following main types of related technologies exist:

[0024] VRVQ (Variable Residual Vector Quantization): Introduce variable bitrate (VBR) into the framework of Residual Vector Quantization (RVQ), and reduce the overall bitrate by dynamically allocating the number of quantizers for different frames.

[0025] DQ-VAE (Dynamic Quantization VAE): An encoding and decoding method in the field of image compression. It adjusts the quantization granularity according to the signal complexity of the image region in the variational autoencoder to achieve a variable bitrate in the spatial dimension.

[0026] Other CFR (Constant Frame Rate) codecs (such as DAC): Allocate quantization resources through a fixed frame rate, and control the bitrate only by adjusting the number of quantizers or the codebook selection.

[0027] However, none of these related technologies have achieved variable frame rate (VFR) in the field of speech, resulting in the following problems: unable to dynamically adjust the frame rate according to the information density, using the same frame rate for silent segments and complex speech segments, resulting in redundant encoding; the sequence length is not reduced, and downstream generation tasks (such as speech synthesis, dialogue systems) still need to process long sequences, resulting in computational latency and memory consumption.

[0028] Specifically, in the current related technologies, the technologies for implementing variable bitrate for neural network-based codecs focus on adjusting the bitrate at the same frame rate (such as the dynamic allocation of the RVQ codebook in VRVQ), optimizing the bitrate allocation (such as reducing the number of codebooks without sacrificing the reconstruction performance by optimizing the network architecture or quantization strategy), or reducing the fixed frame rate (such as from 75Hz to 37.5Hz, etc.). They all focus on "single-frame optimization" and do not consider the dynamic management of time resolution.

[0029] Currently, no research has introduced a variable frame rate (VFR) strategy in the field of neural speech encoding and decoding.

[0030] Figure 1 The flowchart of an example of the speech reconstruction method according to an embodiment of the present application is shown.

[0031] As Figure 1 shown, in step S110, the original speech data to be reconstructed is obtained, and the granular speech frame features of the original speech data at each frame rate are extracted.

[0032] Here, the sources of the original speech data can be diverse, such as real-time recordings, audio files, voice stream data, etc., and digital audio processing tools can be used to preprocess the speech data, such as denoising, normalization, and removing DC bias. In addition, various frame rate converters can be employed to process the original speech data to achieve the acquisition of speech frame features of the original speech data at different frame rates. In some embodiments, a feature extractor is used to extract the speech frame features of the speech data at the normal frame rate, and the granularity speech frame features at multiple levels of coarse-grained frame rates are extracted by single or multiple downsampling methods.

[0033] In some examples of the embodiments of the present application, feature extraction is performed on the original speech data based on multiple cascaded downsampling layers in a convolutional neural network to determine the granularity speech frame features of the original speech data corresponding to multiple downsampled frame rates respectively.

[0034] Exemplarily, a deep network including multiple convolutional layers is constructed, and each layer is followed by a downsampling operation (for example: convolution + pooling or a design with a convolution stride greater than 1). Each downsampling layer reduces the temporal resolution by adjusting the convolution kernel size and stride, thereby generating feature maps with different downsampling rates. The feature maps generated by each downsampling layer represent the "granularity" speech frame features corresponding to that layer. Thus, through multi-level downsampling, the system can capture the local details and global structure of the speech signal at different time scales.

[0035] In step S120, the speech information density of each granularity speech frame feature within each time window is calculated.

[0036] Here, the speech information density can be measured according to indicators such as the degree of change of the time-frequency characteristics of the speech signal, energy fluctuation, spectral entropy, etc. to identify and distinguish the segments rich in effective information in the speech data. Exemplarily, the speech information density can be determined according to the energy amplitude of the granularity speech frame feature within the corresponding time window. It is relatively easy to implement, has low computational overhead, and has high real-time performance, but it cannot distinguish complex speech from noise, which may lead to incorrect frame rate allocation.

[0037] As an additional or alternative embodiment, the speech information density is determined according to the Shannon entropy of the granularity speech frame feature within the corresponding time window. Using the Shannon entropy as a metric can accurately reflect the uncertainty and diversity of the speech signal within each time window, thereby finely distinguishing the high-information-density regions and low-information-density regions. More details and supporting explanations will be elaborated in combination with other examples below.

[0038] In some embodiments, the continuous speech signals corresponding to each granularity speech frame feature are divided according to a fixed or sliding window, such that each window contains several frames. Exemplarily, the granularity speech frame feature includes the speech frame sequences of the original speech data at fine granularity, medium granularity, and coarse granularity. For speech frame sequences of various different granularities, the speech information density is calculated respectively under the time window of the corresponding granularity.

[0039] In some examples of the embodiments of the present application, the time window corresponding to the low frame rate is larger than the time window corresponding to the high frame rate. The high frame rate short window has high sensitivity to local detail changes and can capture instantaneous speech dynamics; while the low frame rate long window integrates long-term information, is efficient in processing and occupies less resources. In addition, by adopting multi-scale window division, the calculation of the speech information density of each granularity feature can be processed in parallel or hierarchically, reducing the calculation burden of each module. At the same time, since the number of low frame rate features itself is small and is statistically calculated within a longer window, it is more conducive to quickly obtaining a reliable entropy value, further shortening the subsequent processing delay.

[0040] In step S130, according to the speech information density, granularity frame feature segments that meet a preset granularity ratio are screened from each granularity speech frame feature, and the granularity ratio predefines the time ratio of the segment duration of each granularity frame feature segment in the unit time period.

[0041] In some embodiments, the user can, according to application requirements and experimental tuning, predefine a granularity ratio parameter to clearly specify the duration ratio that each granularity feature segment should occupy in each time window. In addition, the sum of the time ratios corresponding to each type of granularity frame feature segment can be 1 to avoid missing coding. For example, within a unit time (such as 1 second), the time ratios between the high, medium, and low information density frames are 0.3:0.5:0.2 respectively to complete the frame coverage of all time spans within the unit time.

[0042] As described above, the segment ratio extracted from different granularity speech frame features is constrained by the granularity ratio, and in addition, the standard of the segments extracted from different granularity speech frame features is constrained by the speech information density. Exemplarily, each frame in the granularity speech frame feature is sorted according to the information density, and frames are selected in descending order until the total duration of the selected frames reaches the number of frames required by the corresponding granularity ratio. Thus, low frame rate feature segments are selected and allocated to the low information density region, and high frame rate feature segments are selected and allocated to the high information density region, realizing variable frame rate adaptive resource allocation, which can greatly reduce irrelevant or redundant data and reduce the overall token sequence length.

[0043] In step S140, speech reconstruction is performed based on each granularity frame feature segment to obtain the corresponding reconstructed speech data.

[0044] In some embodiments, each granularity frame feature segment jointly constitutes the feature of the total speech time of the corresponding original speech data, which includes the fine-grained speech frame features corresponding to the high information density region and the coarse-grained speech frame features corresponding to the low information density region. Since coarse-grained frame feature segments with a corresponding low frame rate can be used for the low information density region (e.g., silent segments), the sequence length is reduced and the processing efficiency is improved. In addition, for the high information density region (e.g., speech active segments), fine-grained frame feature segments with a corresponding high frame rate can be used, which can be reconstructed based on key information frames, not only retaining the dynamic features of the speech but also avoiding noise caused by redundant data, ensuring that the reconstructed audio has high clarity and naturalness. In some embodiments, for features of different granularities, a hierarchical reconstruction strategy can be adopted, that is, first generate a rough audio waveform, and then use fine-grained information for local refinement and enhancement.

[0045] Through the embodiments of the present application, through multi-frame rate feature extraction and frame screening based on information density, the system can adaptively perform dynamic frame rate allocation between the speech active region and the silent region, significantly reducing redundant data, thereby reducing the overall bit rate consumption and sequence length, and improving the reconstruction quality under the same bit rate configuration.

[0046] Figure 2 The operation flowchart showing an example of screening granularity frame feature segments according to the embodiments of the present application is shown.

[0047] As Figure 2 shown, in step S210, the first speech information density sequence and the second speech information density sequence corresponding to the first granularity speech frame feature and the second granularity speech frame feature are obtained respectively. The second frame rate corresponding to the second granularity speech frame feature is higher than the first frame rate corresponding to the first granularity speech frame feature.

[0048] Specifically, for each granularity, within their respective preset time windows, according to the statistical distribution of the frame features within the window, calculate its Shannon entropy and / or energy amplitude to obtain the first speech information density sequence and the second speech information density sequence. Since the frame rate of the second granularity is higher, its information density sequence will provide more frequent and fine-grained dynamic change information, while the first sequence reflects the relatively stable global information density.

[0049] In step S220, perform a first sorting on each speech information density in the first speech information density sequence, and screen the first granularity frame feature segments according to the result of the first sorting and the granularity ratio.

[0050] Specifically, sort the speech information densities of each window (corresponding to the first granularity frame feature) in the first speech information density sequence. For example, a descending order strategy or an ascending order strategy can be adopted. When the descending order strategy is used, the window with the highest information density is ranked at the front to ensure capturing the most significant global changes in the speech signal. Then, according to the granularity ratio for the segment duration constrained by the first granularity, select several windows ranked at the front in the sorting result as the focus of the coarse-grained frame coding.

[0051] Regarding the processing details of the first screened granularity frame feature segment, it can generate a binary mask corresponding to the first frame rate according to the result of the first sorting and the granularity ratio. Furthermore, perform mask processing on the first granularity speech frame feature according to the binary mask to obtain the corresponding first granularity frame feature segment.

[0052] Specifically, use the sorting result to determine a dynamic threshold or directly select the top N frames (meeting the granularity ratio requirements) to generate a binary vector with the same length as the first granularity speech frame sequence. Exemplarily, for each frame, if the Shannon entropy value of the frame reaches or exceeds the threshold (or is within the top N), the corresponding position is marked as "1", otherwise marked as "0". In addition, to avoid overly scattered selections in the mask, a sliding window can be introduced to ensure that the selected frames have a certain continuity in time and guarantee the coherence of the signal during speech reconstruction.

[0053] Regarding the details of the specific mask processing, in some embodiments, multiply the generated binary mask with the first granularity speech frame feature frame by frame, or use the mask to perform zero-padding or directly discard the frames that do not meet the conditions. In this way, after the mask processing, the remaining frame features constitute the screened first granularity frame feature segments, and these segments cover the time periods with the highest global information density, providing key speech information for subsequent reconstruction.

[0054] Here, the low-information density frames are filtered out through the binary mask, effectively reducing redundant data, reducing the length of the token sequence that the subsequent autoregressive model or decoder needs to process, thereby reducing the computational burden and improving the real-time processing ability. In addition, the generation and application process of the binary mask has a small amount of calculation and simple operations, which is convenient to be embedded in an end-to-end system to achieve fast feature screening and dynamic resource allocation, and helps to reduce latency in real-time speech applications.

[0055] In step S230, for the subsequence corresponding to the remaining time period in the second speech information density sequence, perform a second sorting on the respective speech information densities in the subsequence, and screen the second granularity frame feature segment according to the result of the second sorting and the granularity ratio. The remaining time period is the time period obtained by subtracting the segment duration of the first granularity frame feature segment from the total speech time period.

[0056] Specifically, since the coarse-grained frame feature segments occupy a part of the total speech duration, their durations are deducted from the total duration, and the remaining part is the target area for subsequent finer-grained processing. Then, according to the granularity ratio, the sorted local information density windows are screened for the segment durations constrained by the second granularity, and segments with high information content in the remaining time period are selected. In this way, fine coding of the remaining area is achieved, making up for the local dynamic details that may be missed in the first-round global screening.

[0057] In the embodiments of the present application, through multiple rounds of density analysis at different scales, the coarse-grained density analysis first screens to cover the global high-density areas, and then the fine-grained screening analysis conducts a detailed analysis of the remaining time period to ensure that the overall speech information is fully expressed at both the coarse and fine levels. Through hierarchical sorting and screening, an adaptive balance between global and local information is achieved, and the coding resources are accurately allocated to the speech segments that truly need to be strengthened for processing, further reducing redundant coding and improving the system robustness.

[0058] Figure 3 The operation flowchart of an example of speech reconstruction based on frame feature segments of each granularity according to the embodiments of the present application is shown.

[0059] As Figure 3 shown, in step S310, the frame feature segments of each granularity are subjected to residual vector quantization processing to determine the corresponding quantization vectors.

[0060] In residual vector quantization (RVQ), first, a primary quantizer is used to roughly quantize the original features, and then the residuals (i.e., the differences between the original features and the outputs of the primary quantizer) are input into subsequent quantizers, and the iteration continues until the preset quantization accuracy is met or the set number of layers is reached. By using the layer-by-layer refinement mechanism of residual quantization, the quantization error is effectively reduced, ensuring that each quantization vector can accurately represent the key information in the original speech feature segments.

[0061] In addition, since the feature segments of each granularity are independent of each other, a parallel computing strategy can be adopted to improve the efficiency of the entire quantization processing. In addition, the hierarchical design allows multiple time resolutions to be represented by one codebook index frame, thereby reducing the coding bit rate of the low-entropy regions with a smaller frame rate.

[0062] In step S320, the quantization vectors of each granularity are subjected to frame rate alignment processing according to the base frame rate, and the quantized vectors after frame rate alignment processing are concatenated to obtain the corresponding fused latent representation.

[0063] Here, the base frame rate is the frame rate of the original speech data to ensure that the finally reconstructed audio is consistent with the original signal in the time domain. For quantization vectors from different granularities, due to their different downsampling rates, methods such as interpolation, repetition, or time normalization are required to expand the low-frame-rate quantization vectors to the time domain scale consistent with the base frame rate. For example, the low-frame-rate quantization vectors can be filled with missing frames by linear interpolation or periodic replication to align with the high-frame-rate vectors.

[0064] In step S330, speech reconstruction is performed based on the fused latent representation to obtain the corresponding reconstructed speech data.

[0065] Here, a speech decoder can be used to reconstruct the reconstructed speech data from the fused latent representation. The decoder can adopt a convolutional neural network, a recurrent neural network, or a Transformer structure, etc.

[0066] Through the embodiments of the present application, from residual vector quantization to frame rate alignment splicing, and then to speech reconstruction based on the fused latent representation, it ensures the efficient capture and expression of global and local information, providing comprehensive and accurate feature support for speech reconstruction. Through precise quantization and alignment processing, the system can maintain a high-fidelity reconstruction effect while compressing the data volume, reducing the computational complexity of the downstream generation module and meeting the requirements of real-time speech generation. In addition, the fusion of multi-granularity and cross-time-domain feature representations enables the system to maintain a stable reconstruction quality in the face of different noise environments and speech dynamic changes, further enhancing the robustness and adaptability of the overall system.

[0067] Regarding the implementation details of step S330, in some embodiments, it can be to perform average pooling processing on the fused latent representation to obtain the corresponding pooled feature vector. In this way, through the average pooling operation, global statistics and dimensionality reduction of the features at each moment are performed to extract the representative pooled feature vector. Then, transposed convolution and upsampling are performed on the pooled feature vector to obtain the corresponding upsampled feature sequence. In some cases, using the transposed convolution layer, the pooled feature vector is expanded into a continuous upsampled feature sequence. A multi-level step-by-step upsampling strategy may be adopted during the upsampling process to achieve smoother and finer-grained temporal recovery. Furthermore, based on the fusion of the upsampled feature sequence and the quantization vector matching the upsampled frame rate, the corresponding reconstructed speech data is obtained.

[0068] Here, through average pooling and transposed convolution upsampling, an upsampled feature sequence matching the original speech frame rate has been obtained, which contains global statistical information and restored fine-grained temporal features. Then, the quantization vectors corresponding to the original quantization at the corresponding frame rate are fused. Each layer refines the output of its previous layer by aggregating the reconstructed features with the earliest quantization features, avoiding the error propagation that may be caused by solely relying on the previous reconstruction result. In addition, the system can utilize both global statistical information and local details, enabling the reconstructed speech to retain the overall structure of the original speech while finely restoring the instantaneous dynamic changes, improving the accuracy and clarity of speech reconstruction.

[0069] Next, the specific details of a Temporally Flexible Coding (TFC) method for a neural speech codec provided by an embodiment of the present application will be elaborated with reference to examples.

[0070] Figure 4 The schematic diagram of the working principle of an example of the temporally flexible coding framework according to an embodiment of the present application is shown.

[0071] As Figure 4 shown, in temporally flexible coding, the information density in the time domain of the speech signal is estimated through non-parametric entropy, and the frame rate is dynamically allocated to achieve variable frame rate compression.

[0072] Specifically, in the Encoder, the input speech signal extracts multi-level features via a convolutional neural network, obtaining fine-grained, medium-grained, and coarse-grained features respectively.

[0073] In Entropy Values & Generate Binary Masks, for each granularity feature, the entropy value is calculated within the corresponding time window to measure the speech information density of each frame.

[0074] Furthermore, according to the preset granularity ratio, the high-information-density region and the low-information-density region are determined through the entropy value, and the corresponding binary masks are generated respectively for dynamic frame rate allocation:

[0075] - High-entropy value region (such as the active speech segment): Allocate a higher fine-grained frame rate;

[0076] - Low-entropy value region (such as the silent segment): Adopt medium and coarse-grained coding to reduce the frame rate.

[0077] In the Quantization stage, the features screened by the binary mask are quantized, mapping the continuous features to discrete codebook indices.

[0078] In the decoder, the quantized multi-scale features are sequentially fed into the decoder backbone, and deconvolution is performed through transposed convolution and upsampling operations. Furthermore, a bottom-up conditional hierarchical design is adopted to gradually fuse features of different resolutions from coarse to fine, and finally a high-fidelity speech signal is reconstructed.

[0079] Through the TFC method, the overall sequence length is shortened at a given total bit rate to achieve dynamic frame rate allocation, thereby obtaining better temporal compactness and flexibility.

[0080] Figure 5 The schematic diagram of the simulation effect showing the VFR allocation example of the temporal flexible coding according to the embodiment of the present application at different granularity ratios is shown.

[0081] As Figure 5 shown, it clearly shows that the coarse-grained frames are mainly allocated to the silent segments, and the granularity ratio can be conveniently adjusted to adapt to different frame rates.

[0082] Through the TFC method provided by the embodiment of the present application, the bit rate and the sequence length are reduced, the reconstruction quality is improved under the same bit rate configuration, dynamic frame rate adjustment (18.75Hz - 75Hz) is supported, and different scenario requirements can be adapted. In addition, the generation model (such as VALL-E) processes shorter sequences, reduces the inference delay, and helps to accelerate downstream tasks. Moreover, the low frame rate silent segments greatly reduce the data volume, are orthogonal to the existing codebook compression techniques (such as codebook dropout), and can be used in an additive manner. It promotes the evolution of speech codec from "fixed timing" to "semantic-driven timing", and provides new ideas for multi-modal generation (speech + text alignment).

[0083] It should be noted that it is proposed in this article that a temporal flexible coding (TFC) technique with variable frame rate (VFR) is introduced into the neural speech codec, which can seamlessly adjust the average frame rate and dynamically allocate the frame rate according to the temporal entropy value. The experimental results show that the codec using TFC achieves the optimal reconstruction quality while maintaining a high degree of flexibility, and remains competitive even at a low frame rate.

[0084] 1. Introduction

[0085] By dynamically adjusting the receptive field of the output encoding along the time axis, TFC first introduced VFR encoding into the neural speech codec, utilizing a non-parametric entropy calculation technique to evaluate the entropy value of the speech passage for dynamic frame rate allocation. Regions with high entropy values indicate information-dense content and will be allocated a higher frame rate; while for low-entropy regions such as silence, a lower frame rate will be allocated. In short, TFC enables a single model to support user-specified average frame rates from 18.75 Hz to 75 Hz (based on the DAC framework for 24 kHz audio) and achieve time-variable frame rate allocation within the specified total bitrate range.

[0086] Overall, TFC achieves flexible frame rate control and interpretable granularity allocation in a unified framework. This work makes a new contribution to speech coding by first introducing VFR into the neural speech codec. In addition, the design of TFC can be orthogonally integrated with other techniques aimed at reducing the sequence length and overall bitrate, showing broad and promising application prospects.

[0087] 2. Background

[0088] In this section, in the broader context of the development of feature compression, terms such as CBR, VBR, CFR, and VFR are comprehensively discussed and necessary definitions are provided for further understanding.

[0089] 2.1. Constant Bitrate and Variable Bitrate in Neural Codecs

[0090] In the field of audio and speech coding and decoding, the concepts of Constant Bitrate (CBR) and Variable Bitrate (VBR) are defined from the time domain. CBR means that the codec maintains a constant number of bits per unit time, while VBR allows dynamic adjustment of the bit allocation in the time domain according to the complexity of the signal.

[0091] Formally, let z e = f θ (x) ∈ R D×T represent the continuous latent representation of audio x, obtained through the encoder f θ (·), where T represents the number of downsampled time frames and D represents the dimension of the latent vector. At a specific frame t, the latent vector z e [t] is quantized to obtain z q [t]. In a codec based on Residual Vector Quantization (RVQ), z q [t] is calculated as:

[0092]

[0093] where Q i(·) is the i-th quantizer, and the residual vector r i [t] is recursively defined as:

[0094]

[0095] Subsequently, the decoder g ψ (·) decodes z q to reconstruct the audio

[0096] In some current related technologies, a single RVQ-based codec can operate at multiple target bitrates through a quantizer discard strategy. Specifically, by randomly sampling n q ~{1,…,N q}, and only using the first n q quantizers for different samples during training, the user can flexibly select n q during inference. A smaller n q will result in lower reconstruction quality but also reduce the bitrate usage. However, this technique itself does not change the fact that the RVQ codec allocates the same number of codebooks to all time frames at a fixed frame rate, so they are still CBR.

[0097] The demand for speech generation has driven new concepts in neural codecs, including the design of single-codebook, multi-resolution, and low-frame-rate or low-bitrate codecs. Although these designs have advantages in real-time conversation applications, they still belong to the CBR paradigm.

[0098] Recently, VRVQ introduced a VBR strategy into the RVQ-based neural audio codec. This method assigns a time-varying n q to different downsampled frames based on the importance mapping of neural prediction during inference. The design of VRVQ performs well in reducing the overall bitrate consumption but does not reduce the number of frames.

[0099] 2.2. Variable Frame Rate Compression

[0100] Although VRVQ achieves VBR encoding, it still operates at a fixed frame rate (CFR), i.e., each encoded frame represents a fixed time window of the signal. Low frame rate is particularly important for autoregressive modeling, which has prompted the clarification of a broader concept of VBR and the discussion of variable frame rate (VFR) codecs.

[0101] In the field of audio compression, VBR does not necessarily imply VFR, especially for RVQ-based codecs, where each frame can contain multiple quantizers. Here, VFR is defined as an encoding strategy such that each encoded frame can adaptively cover different time spans, meaning the model can effectively adjust the downsampling rate according to the information density. For example, during a silent segment, a VFR encoding can cover a longer time span, while during a rapid phoneme transition, it covers a shorter span.

[0102] The concept of VFR has been implemented in semantic tokens derived from speech self-supervised learning models, which are mainly used for discriminative tasks such as automatic speech recognition (ASR). In these tasks, preserving fine-grained acoustic details is less important than semantic information, so VFR compression can be achieved by extracting information from pre-trained models or using unit discovery techniques. However, the acoustic tokens generated by neural codecs inherently encode more detailed information, so it becomes more difficult to allocate appropriate proportions for different encoding granularities because they lack the explicit supervision labels that semantic tokens have. This requires the development of effective heuristic strategies to allocate variable frame rates.

[0103] Inspired by variable-rate image compression methods that adjust spatial resolution according to content complexity, this paper introduces a VFR strategy into speech codecs to adaptively manage the temporal resolution.

[0104] 3. Temporal Flexible Encoding

[0105] This section describes the proposed Temporal Flexible Encoding (TFC) strategy. TFC is a pluggable module that can be integrated into various codec backbones such as DACs and does not introduce additional optimization losses.

[0106] 3.1. Information Density Measurement Based on Temporal Entropy

[0107] To measure the information density of speech segments for the granularity allocation of VFR encoding, a non-parametric entropy-based method is selected. Although it is possible to choose to use a neural router to allocate granularity, this leads to instability in gradient estimation, and the implementation shows that the entropy-based method performs better. Therefore, this paper improves the algorithm originally designed for spatial entropy to adapt to the temporal information of speech signals.

[0108] Suppose τ is a time period, and x[t] ∈ [-1, 1] represents the amplitude value of the normalized speech signal x at t ∈ τ. Define N uniformly distributed intervals {u1, u2,..., u N}, covering the interval [-1, 1], where and \(i = 0, 1, \ldots, N - 1\). To estimate the probability density of \(x[t]\) over these intervals, the Gaussian similarity \(p\) of \(x[t]\) with each interval \(u\) is calculated i as t,i :

[0109]

[0110] where \(\sigma\) controls the sharpness of the distribution. This describes the likelihood that \(x[t]\) "spreads" into the interval \(u\) i .

[0111] For a certain segment \(\tau\), the similarities of all samples \(|\tau|\) in \(\tau\) are averaged and normalized to form a probability distribution:

[0112]

[0113] where \(\epsilon\) is used to ensure numerical stability. Then, the temporal entropy of \(\tau\) is calculated through Shannon entropy:

[0114]

[0115] These calculated entropy values reflect the information density of a specific time window \(\tau\) and are used for the dynamic frame rate allocation described in the next section.

[0116] 3.2. Encoder Supporting Variable Frame Rate Allocation

[0117] Let the encoder backbone output of the codec represent the sequence \(z\) f , with a receptive field of \(w\) f = \(W\) sampling points and a stride of \(s\). Here, \(F\) is used to denote its frame rate, corresponding to the high - resolution representation \(z\) f . Then, two additional subsampled representation sequences are generated through a CNN module:

[0118] - Medium - resolution representation \(z\) m , downsampled from \(z\) f , with a frame rate of \(F / 2\) and a stride of \(2s\), and a receptive field of \(w\) m = \(W + s\) sampling points per frame;

[0119] - Low - resolution representation \(z\) c , downsampled from \(z\) m , with a frame rate of \(F / 4\) and a stride of \(4s\), and a receptive field of \(w\) c = \(W + 3s\) sampling points per frame.

[0120] For each resolution, entropy values are calculated based on the method in Section 3.1, where the size of \(\tau\) is set to the corresponding receptive field and the time segment is slid with the corresponding stride. Denote the scalar entropy sequences of high, medium, and low resolutions as \(h\) f , \(h\) m , \(h\)c , and they are respectively consistent with the frame rate of z f , z m , z c .

[0121] Then, these three resolutions are combined into a VFR representation. The user-predefined non-negative scalar granularity ratio is r f , r m , r c and r f + r m + r c = 1. First, a binary sequence granularity mask b c , b m , b f :

[0122]

[0123]

[0124] where 1(·) is the indicator function, Quantile q (h) represents the q-quantile in each speech, (·)↑k represents the repeated operation k times, and h′ m represents the part not covered by the b c mask. The frame rate of the binary mask b is consistent with the resolutions of the corresponding h and z. In this way, a 1-second speech segment will generate Fr c / 4 coarse-grained frames, Fr m / 2 medium-grained frames, and Fr f high-grained frames before quantization, and their time spans follow the ratio r c : r m : r f .

[0125] Subsequently, for each resolution, the continuous representation is quantized to generate using the RVQ process described by Equation (1). To fuse these representations into a unified feature sequence aligned with the finest time resolution the quantization vectors of the medium and coarse resolutions are repeated and combined with their corresponding masks through element-wise multiplication ⊙:

[0126]

[0127] Generally speaking, this hierarchical design allows multiple time resolutions to be represented by one codebook-indexed frame, thereby reducing the coding bit rate in low-entropy regions through a smaller frame rate.

[0128] 3.3. Decoder for Conditional Hierarchical Design

[0129] A conditional hierarchical design is adopted in the decoder to refine the fused latent representation after quantization by gradually integrating the scale information from coarse to fine. Given the binary granularity masks {b f ,b m ,b c}, the decoder processes the latent representation as follows

[0130] - The coarse-resolution representation y c is obtained by average pooling on : - The mid-resolution representation y m is conditioned on y c and combined with the original quantized mid-resolution latent features. Specifically, a transposed CNN block is used to upsample y c by a factor of 2, and the upsampled sequence is added to .

[0131] - The finest-resolution representation y f integrates y m and y c and fuses with the original quantized fine-resolution latent features . Similarly, y m is upsampled by a transposed CNN and then added to .

[0132] Overall, each layer refines the output of its previous layer by aggregating the reconstructed features with the earliest quantized features, avoiding the error propagation that may occur by simply relying on the previous reconstruction result. Finally, y c with the original frame rate F is passed into the decoder backbone of the codec to recover the waveform.

[0133] 4. Experiments

[0134] 4.1. Architecture and Setup

[0135] TFC is implemented on the framework of DAC, which achieves high-fidelity reconstruction performance among various existing codecs. Here, using its official configuration, 24kHz audio is processed to generate RVQ coding at a frame rate of 75Hz. Thus, in the DAC+TFC framework, the finest granularity is F = 75Hz. Here, the hierarchical design (see Section 3.2 for details) uses CNN blocks for 2 times of 2-fold downsampling to obtain the medium and coarse granularity representations, and the upsampling is reversely implemented in the transposed CNN blocks in Section 3.3. The encoded representation is quantized using 8-dimensional RVQ, each quantization codebook contains 1024 (10-bit) entries, and the maximum number of codebooks is N q= 32. Since this setting results in a relatively high bitrate (24 kbps), which is not practical for downstream generation applications, a codebook discarding strategy was adopted during the inference stage, limiting the number of codebooks used to a maximum of 8 (6 kbps).

[0136] The DAC baseline model and our proposed DAC+TFC model were trained on the 960-hour speech data of LibriTTS. To ensure sufficient training at different temporal resolutions, the frame rate ratio r was heuristically assigned in the batch input of DAC+TFC. f = 0.4, r m = 0.3, r c = 0.3. Each model was batched with 1-second segments, with a batch size of 20, and trained for 1 million iterations.

[0137] The reconstruction quality was evaluated using Mel and STFT distances, with the latter being more capable of capturing high-frequency fidelity. Additionally, UTMOS (Universal Target Mean Opinion Score, a MOS prediction system highly correlated with human ratings), STOI (Short-Time Objective Intelligibility), and WER (Word Error Rate) metrics were used to ensure comprehensive evaluation.

[0138] 4.2. Result A: CFR Performance at Different Bitrates

[0139] DAC+TFC supports frame rate selection from 18.75 Hz to 75 Hz by adjusting the granularity ratio during inference. This can be regarded as a bitrate control method parallel to codebook discarding. For the DAC baseline model, different bitrates can be selected by specifying the number of codebooks during inference. Therefore, a fair comparison between DAC and the proposed DAC+TFC can be made at the same bitrate. For example, if the DAC baseline model is used for inference with 4 codebooks, the bitrate is 3 kbps; then, it can be compared with DAC+TFC with a frame rate of 37.5 Hz and 8 codebooks, which also has a bitrate of 3 kbps.

[0140] Table 1: CFR Reconstruction Performance of DAC and the Proposed DAC+TFC at Different Bitrates.

[0141]

[0142] The relevant results are shown in Table 1. Significantly, DAC+TFC outperforms DAC in most metrics, especially in terms of WER. These results indicate that TFC effectively controls the bitrate, maintains the audio quality, and achieves the original intention of reducing the sequence length. For downstream generative models, reducing the number of frames will significantly accelerate the generation process, while increasing the number of codebooks has a negligible impact on the generation speed.

[0143] 4.3. Result B: VFR vs CFR Performance Comparison

[0144] In Section 4.2, TFC demonstrated basic temporal flexibility but still operated in CFR mode because the frame rate was directly allocated through a single granularity. In this section, the performance of TFC under true VFR conditions is evaluated, i.e., when at least two granularities are combined during inference.

[0145] First, the VFR and CFR cases are compared at the same average frame rate. A setting with an average bitrate of 3 kbps and a frame rate of 37.5 Hz is selected because this allows TFC to mix 18.75 Hz and 75 Hz components in a 2:1 ratio, resulting in VFR results (e.g., setting r f , r m , r c to 0.2, 0.7, 0.1 respectively).

[0146] Figure 6 Fig. shows a schematic diagram of the simulation effect of an example of the performance of DAC+TFC in VFR mode. As Figure 6 shown in part (a) of, compared with pure 37.5 Hz (i.e., 0% of 75 Hz granularity) and DAC at 3 kbps, TFC brings significant improvements by mixing other frame rates, as reflected in the Mel distance, UTMOS, and STOI metrics. Additionally, the lowest WER occurs when the 75 Hz component is 6%. Overall, these results indicate that TFC has the flexibility to customize the frame rate and achieves superior performance through VFR encoding.

[0147] To further explore the potential of TFC, its low-frame-rate VFR results are compared with the 75 Hz CFR setting. Since the reconstruction quality of the same architecture inevitably decreases at lower bitrates, the UTMOS and WER metrics can be focused on to evaluate the perceptual quality. The results (both using N q = 8) include the 75 Hz DAC baseline model, as well as different configurations of DAC+TFC, with the frame rate ratios r c : r m : r f gradually changing from 0:0:100% (pure 75 Hz frame rate) to 0:50%:50% (average 56.25 Hz frame rate). AsFigure 6 As shown in part (b) of , except for one outlier in WER, the DAC+TFC configuration is generally always better than the 75Hz DAC baseline. Notably, the UTMOS scores for some lower frame rate settings are even higher than those of the 75Hz DAC+TFC configuration, indicating that the VFR design has significant potential in more compactly compressing speech.

[0148] 5. Conclusion

[0149] This paper proposes Temporally Flexible Coding (TFC), which first introduces Variable Frame Rate (VFR) into neural speech codecs and dynamically adjusts the temporal resolution according to the information density. While reducing the sequence length, TFC can effectively balance the bit rate and audio quality.

[0150] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0151] In some embodiments, the embodiments of this application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used to execute any one of the methods for training a large language model described above in this application.

[0152] In some embodiments, the embodiments of this application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is made to execute any one of the methods for training a large language model described above.

[0153] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a method for training a large language model.

[0154] Figure 7 FIG. 4 is a schematic hardware structure diagram of an electronic device for executing a method for training a large language model according to another embodiment of the present application. As Figure 7 shown, the device includes:

[0155] One or more processors 710 and a memory 720. Figure 7 Here, one processor 710 is taken as an example.

[0156] The device for executing the method for training a large language model may further include: an input device 730 and an output device 740.

[0157] The processor 710, the memory 720, the input device 730, and the output device 740 may be connected through a bus or other means. Figure 7 Here, connection through a bus is taken as an example.

[0158] The memory 720, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for training a large language model in the embodiments of the present application. The processor 710 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 720, that is, implements the method for training a large language model in the above method embodiments.

[0159] The memory 720 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 720 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 720 may optionally include a memory remotely disposed relative to the processor 710, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0160] The input device 730 can receive input digital or character information and generate signals related to the user settings and function controls of the electronic device. The output device 740 can include display devices such as a display screen.

[0161] The one or more modules are stored in the memory 720 and, when executed by the one or more processors 710, perform the method for training a large language model in any of the above method embodiments.

[0162] The above product can execute the method provided in the embodiments of the present application and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0163] The electronic device in the embodiments of the present application exists in various forms, including but not limited to:

[0164] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0165] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.

[0166] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.

[0167] (4) Other on-board electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment.

[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A voice reconstruction method, comprising: Obtaining original voice data to be reconstructed, and extracting granular voice frame features of the original voice data at each frame rate; Calculating the voice information density of each of the granular voice frame features within each time window; Screening granular frame feature segments that meet a preset granularity ratio from each of the granular voice frame features; the granularity ratio predefines the time proportion of the segment duration of each granular frame feature segment in a unit period; Performing voice reconstruction based on each of the granular frame feature segments to obtain corresponding reconstructed voice data.

2. The method according to claim 1, wherein, The voice information density is determined according to the Shannon entropy and / or energy amplitude of the granular voice frame features within the corresponding time window.

3. The method according to claim 1, wherein, The screening of granular frame feature segments that meet a preset granularity ratio from each of the granular voice frame features according to the voice information density includes: Obtaining a first voice information density sequence and a second voice information density sequence corresponding to a first granular voice frame feature and a second granular voice frame feature respectively; the second frame rate corresponding to the second granular voice frame feature is higher than the first frame rate corresponding to the first granular voice frame feature; Performing a first sorting on each voice information density in the first voice information density sequence, and screening first granular frame feature segments according to the result of the first sorting and the granularity ratio; For a subsequence corresponding to the remaining time period in the second voice information density sequence, performing a second sorting on each voice information density in the subsequence, and screening second granular frame feature segments according to the result of the second sorting and the granularity ratio; the remaining time period is a time period obtained by subtracting the segment duration of the first granular frame feature segment from the total voice period.

4. The method according to any one of claims 1-3, wherein The time window corresponding to the low frame rate is larger than the time window corresponding to the high frame rate.

5. The method according to claim 3, wherein The screening of first granular frame feature segments according to the result of the first sorting and the granularity ratio includes: Generating a binary mask corresponding to the first frame rate according to the result of the first sorting and the granularity ratio; Performing mask processing on the first granular voice frame feature according to the binary mask to obtain a corresponding first granular frame feature segment.

6. The method according to claim 1, wherein The extracting of granular voice frame features of the original voice data at each frame rate includes: Performing feature extraction on the original voice data based on multiple cascaded downsampling layers in a convolutional neural network to determine granular voice frame features of the original voice data corresponding to multiple downsampled frame rates.

7. The method according to claim 5, wherein, The performing of voice reconstruction based on each of the granular frame feature segments to obtain corresponding reconstructed voice data includes: Performing residual vector quantization processing on each of the granular frame feature segments to determine corresponding quantization vectors; Performing frame rate alignment processing on each of the quantization vectors according to a base frame rate, and splicing the quantization vectors after frame rate alignment processing to obtain a corresponding fused latent representation; the base frame rate is the frame rate of the original voice data; Performing voice reconstruction based on the fused latent representation to obtain corresponding reconstructed voice data.

8. The method according to claim 7, wherein The performing of voice reconstruction based on the fused latent representation to obtain corresponding reconstructed voice data includes: Perform average pooling on the fusion latent representation to obtain a corresponding pooled feature vector; Perform transposed convolution and upsampling on the pooled feature vector to obtain a corresponding upsampled feature sequence; Based on the fusion of the upsampled feature sequence and the quantization vector matching the upsampled frame rate, obtain the corresponding reconstructed speech data.

9. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-8.

10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method according to any one of claims 1-8.