Voice compression method and system based on multi-scale back projection feature fusion
By using multi-scale projection feature fusion layers and residual convolutional networks, the method enhances voice synthesis quality by capturing detailed speech features and reducing redundancy, addressing the limitations of existing low-rate voice coding methods.
Patent Information
- Application Number
- CN202510787110.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing speech encoding technology cannot fully capture the multi-scale features of speech signals, resulting in the loss of speech details and the lack of detailed compensation mechanism, resulting in limited synthetic speech quality, especially at low bit rates.
Multi-scale back projection feature fusion layer and residual convolution feedforward network are used to capture speech detail features through cross-learning of convolution kernels of different scales, compress and reconstruct using the back projection mechanism, and improve speech quality through feature fusion and residual connection.
It improves the quality of speech synthesis, enhances detailed information and reduces information redundancy, improves the model's local modeling ability and timing expression ability, and improves the speech synthesis effect at low bit rate.
Smart Images

Figure CN120319218A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of speech signal processing, and particularly to a speech compression method and system based on multi-scale back-projection feature fusion. Background Art
[0002] Low-rate speech coding technology has wide application requirements in many key fields such as satellite communication, short-wave communication, underwater acoustic communication, and secure communication.
[0003] Current methods mainly synthesize the original speech through a speech coding model, which includes an encoding end and a decoding end; the encoding end performs downsampling operations on the original speech through a convolutional neural network to complete the feature extraction of the speech signal, and after quantifying and dequantifying the features extracted by the encoding end, the dequantified speech is obtained. The decoding end uses a convolutional neural network to perform upsampling operations on the dequantified speech to restore the input speech and obtain the speech synthesis result.
[0004] The speech signal is a complex signal that contains rich information such as pitch, rhythm, and timbre, and the speech signal itself has multi-scale features that cover speech information in different time and frequency ranges. When current methods perform speech synthesis, they only capture the features of the speech signal through a convolutional neural network, often unable to fully capture complex multi-scale features, resulting in the loss of speech detail information, and current methods lack a corresponding detail compensation mechanism for the upsampled and downsampled speech, resulting in limited quality of the final synthesized speech. Summary of the Invention
[0005] The embodiments of the present application provide a speech compression method and system based on multi-scale back-projection feature fusion, which can improve the quality of speech synthesis. The technical solutions are as follows: On the one hand, a speech compression method based on multi-scale back-projection feature fusion is provided, including: Calling multiple multi-scale back-projection feature fusion layers in the encoder to perform multi-layer feature encoding on the speech signal to be synthesized to obtain speech features; Calling multiple multi-scale back-projection feature fusion layers in the decoder to perform multi-layer feature decoding on the speech features to obtain synthesized speech; Among them, the process of encoding or decoding the input features by the multi-scale back-projection feature fusion layer includes: using convolutional kernels of different scales to perform cross-learning on the input features to obtain multi-scale features, respectively performing back-projection on the extracted multi-scale features to obtain back-projection features, fusing the back-projection features and the multi-scale features to obtain a fusion result, and fusing the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
[0006] Optionally, the multi-scale back-projection feature fusion layer performs multi-path convolution processing on the input features using convolutional kernels of different scales, splices the features after convolutional processing of all paths to obtain spliced features of multiple paths, and performs feature extraction on each spliced feature to obtain multi-scale features.
[0007] Optionally, the multi-scale back-projection feature fusion layer first downsamples each multi-scale feature and then upsamples it to obtain an upsampled feature, calculates the residual between the multi-scale feature and its upsampled feature, performs weighted fusion on the multi-scale feature and the residual to obtain a fusion result, splices all the fusion results to obtain a spliced feature, and performs residual connection on the spliced feature and the feature input to the multi-scale back-projection feature fusion layer to obtain the output feature of the multi-scale back-projection feature fusion layer.
[0008] Optionally, the weights used for weighted fusion of the multi-scale feature and the residual are obtained by calculating the attention vectors of the multi-scale feature and the residual and normalizing the two attention vectors.
[0009] Optionally, after multiple multi-scale back-projection feature fusion layers in the encoder perform multi-layer feature encoding on the speech signal to be synthesized, the input is fed into a residual convolutional feed-forward network. After the output features of the residual convolutional feed-forward network are processed by convolution, speech features are output. After multiple multi-scale back-projection feature fusion layers in the decoder perform multi-layer feature decoding on the speech features, the input is fed into a residual convolutional feed-forward network. After the output features of the residual convolutional feed-forward network are processed by convolution, the synthesized speech is output. Among them, the process of the residual convolutional feed-forward network processing the features input to this network includes: Performing convolution processing and dropout processing on the features input to this network to obtain extended features, performing depthwise separable convolution processing on the extended features to obtain channel-enhanced features, adding the channel-enhanced features to the extended features to obtain residual-enhanced features, performing convolution processing on the residual-enhanced features and performing residual connection with the features input to the residual convolutional feed-forward network to obtain speech features.
[0010] Optionally, the convolutional layer in the encoder is called to extract the basic temporal features of the speech signal to be synthesized. After that, multiple multi-scale back-projection feature fusion layers are used to perform multi-layer feature extraction on the basic temporal features to obtain speech features.
[0011] On the other hand, a speech compression system based on multi-scale back-projection feature fusion is provided, including: An encoding module for calling multiple multi-scale back-projection feature fusion layers in the encoder to perform multi-layer feature encoding on the speech signal to be synthesized to obtain speech features. A decoding module, configured to call multiple multi-scale back-projection feature fusion layers in a decoder to perform multi-layer feature decoding on speech features to obtain synthesized speech; Wherein, the process of the multi-scale back-projection feature fusion layer encoding or decoding input features includes: the multi-scale back-projection feature fusion layer uses convolutional kernels of different scales to perform cross-learning on the input features to obtain multi-scale features, performs back-projection on the extracted multi-scale features respectively to obtain back-projection features, fuses the back-projection features and the multi-scale features to obtain a fusion result, and fuses the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
[0012] On the other hand, a computer device is also provided, and the device includes: A processor, adapted to execute a computer program; A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, it implements the speech compression method based on multi-scale back-projection feature fusion provided in the above aspect.
[0013] On the other hand, a computer-readable storage medium is also provided, and the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the speech compression method based on multi-scale back-projection feature fusion provided in the above aspect.
[0014] On the other hand, a computer program product is also provided, and the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the speech compression method based on multi-scale back-projection feature fusion provided in the above aspect.
[0015] In the speech compression method and system based on multi-scale back-projection feature fusion provided in the embodiments of the present application, in the process of encoding a speech signal through an encoder to obtain speech features and decoding the speech features through a decoder to obtain synthesized speech, multi-scale back-projection feature fusion layers are both adopted. The multi-scale back-projection feature fusion layer uses convolutional kernels of different scales to perform cross-learning on the input features, captures speech detail features of different scales through cross-learning to obtain multi-scale features, performs back-projection on the extracted multi-scale features respectively to obtain back-projection features, uses the back-projection mechanism to compress, reconstruct and perform residual feedback on the captured multi-scale features to obtain finer features, fuses the back-projection features and the multi-scale features to obtain a fusion result, and fuses the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer, and solves the problem of inconsistent information levels in multi-scale feature fusion through feature fusion; thereby improving the quality of speech synthesis.
[0016] In addition, the method proposed in the embodiments of the present application also adds a residual convolutional feed-forward network to the encoder and decoder, introduces depthwise separable convolution and residual connection in the residual convolutional feed-forward network, improves the local modeling ability of the model and the temporal expression ability of speech, enhances detailed information and reduces information redundancy, and further improves the speech synthesis quality by combining the residual convolutional feed-forward network with the multi-scale back-projection feature fusion layer. Brief Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0018] Figure 1 It is a schematic diagram of the overall process of the speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application; Figure 2 It is a flowchart of speech feature quantization in the speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application; Figure 3 It is a block diagram of the speech synthesis model in the speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application; Figure 4 It is a block diagram of the multi-scale back-projection feature fusion layer in the speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application; Figure 5 It is a block diagram of the back-projection mechanism provided by the embodiments of the present application; Figure 6 It is a block diagram of the feature fusion block provided by the embodiments of the present application; Figure 7 It is a block diagram of the residual convolutional feed-forward network in the speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application. Detailed Description of the Embodiments
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0020] Low-rate speech coding technology has extensive application requirements in many key fields such as satellite communication, short-wave communication, underwater acoustic communication, and secure communication. Especially in bandwidth-constrained or extremely harsh environments, the demand for speech synthesis technology in these fields is increasing. For example, satellite communication often needs to cross the atmosphere and space, and the signal transmission bandwidth is extremely limited. In this case, using low-rate speech coding can significantly save transmission bandwidth and improve communication efficiency; in encrypted communication, low-rate speech synthesis technology can not only compress data but also enhance the security of speech through encryption and coding to prevent data from being stolen or tampered with; in extremely harsh mountain communication environments, communication devices not only have to withstand the tests of extreme low temperatures and strong winds but also operate stably with limited energy supply, and the coding rate of speech synthesis is often extremely low. As the coding rate decreases, the quality of speech synthesis will be affected. Therefore, how to ensure the quality of speech synthesis at low bit rates has become a key issue in current speech coding and synthesis technology research.
[0021] The current method mainly synthesizes the original speech through a speech coding model. The process of using the speech coding model to synthesize speech by the current method includes: (1) Encoding end: Use a convolutional neural network to perform downsampling on the original speech so as to retain the main feature information while compressing and complete the extraction of acoustic information.
[0022] (2) Quantization end: Design a codebook to quantize the features extracted by the encoding end, update the codebook by calculating the Euclidean distance between the input vector and the initialized codebook; calculate the codeword with the closest Euclidean distance to the input vector and replace it with the codebook index for transmission.
[0023] (3) Dequantization end: Dequantize the quantized features generated by the quantization end according to the quantization index and restore the acoustic features according to the codebook.
[0024] (4) Decoding end: Use a convolutional neural network to perform upsampling on the dequantized speech to restore the input speech.
[0025] (5) Discriminator: The discriminator is responsible for evaluating the generated speech and promoting the generator to generate more realistic speech through adversarial training, thereby improving the quality of the reconstructed speech.
[0026] A voice signal is a complex signal that contains rich information such as pitch, rhythm, timbre, etc. Moreover, the voice signal itself has multi-scale characteristics, covering voice information within different time and frequency ranges. However, current methods often cannot fully capture these complex multi-scale characteristics, resulting in the loss of voice detail information, or problems such as inconsistent hierarchical information and information redundancy during multi-scale information fusion, limiting the quality of voice synthesis. Additionally, the encoder and decoder in current methods will lose high-frequency detail information during the upsampling and downsampling processes, leading to poor quality of the synthesized voice. During downsampling, especially for low-bitrate voice coding, compressing too many details will cause distortion of the voice signal, making it difficult for the decoder to recover the originally clear and natural voice; upsampling can restore the features to the original resolution, but due to information loss, the voice signal recovered by the decoder often becomes blurred and cannot perfectly restore the details. Thus, current methods lack a corresponding detail compensation mechanism for the voice after upsampling and downsampling, resulting in limited quality of the finally generated synthesized voice.
[0027] In order to improve the quality of voice synthesis, the embodiments of this application propose a voice compression method based on multi-scale back-projection feature fusion, introducing a multi-scale back-projection feature fusion layer and a residual convolutional feed-forward network in the encoder and decoder to improve the quality of the synthesized voice. Among them, the multi-scale back-projection feature fusion layer uses convolutional kernels of different scales to cross-learn the features input to this layer, captures more complete voice feature details, obtains multi-scale features, performs back-projection on the extracted multi-scale features respectively to obtain back-projection features, and uses the back-projection mechanism to compress, reconstruct, and residually feedback the captured multi-scale features to obtain finer features; the feature fusion block adaptively adjusts and fuses the weights from different voice branches to solve the problem of inconsistent information levels in multi-scale feature fusion; secondly, depthwise separable convolution and residual connection are introduced in the residual convolutional feed-forward network to enhance the local modeling ability of the model and the temporal expression ability of the voice, enhance detail information and reduce information redundancy, and further improve the quality of voice synthesis by combining the residual convolutional feed-forward network with the multi-scale back-projection feature fusion layer.
[0028] The following details the voice compression method based on multi-scale back-projection feature fusion provided by the embodiments of this application.
[0029] The voice compression method based on multi-scale back-projection feature fusion provided by the embodiments of this application, as Figures 1-7 shown, includes: Call multiple multi-scale back-projection feature fusion layers in the encoder to perform multi-layer feature encoding on the voice signal to be synthesized, and obtain voice features; Call multiple multi-scale back-projection feature fusion layers in the decoder to perform multi-layer feature decoding on the voice features, and obtain the synthesized voice; Among them, the process of encoding or decoding the input features by the multi-scale back-projection feature fusion layer includes: cross-learning the input features using convolutional kernels of different scales to obtain multi-scale features, respectively performing back-projection on the extracted multi-scale features to obtain back-projection features, fusing the back-projection features and the multi-scale features to obtain a fusion result, and fusing the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
[0030] In a possible implementation, the multi-scale back-projection feature fusion layer performs multi-path convolution processing on the input features using convolutional kernels of different scales, concatenates the features after convolution processing of all paths to obtain the concatenated features of multiple paths, and extracts features from each concatenated feature to obtain multi-scale features. By using convolutional kernels of different scales for cross-learning, more complete details of speech features can be captured.
[0031] In a possible implementation, the multi-scale back-projection feature fusion layer first downsamples each multi-scale feature and then upsamples it to obtain upsampled features, calculates the residuals between the multi-scale features and their upsampled features, weights and fuses the multi-scale features and the residuals to obtain a fusion result, concatenates all the fusion results to obtain concatenated features, and performs residual connection between the concatenated features and the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
[0032] Among them, the weights used for weighted fusion of the multi-scale features and the residuals are obtained by calculating the attention vectors of the multi-scale features and the residuals and normalizing the two attention vectors. By adaptively adjusting and fusing the weights from different speech branches, the problem of inconsistent information levels in multi-scale feature fusion can be solved.
[0033] In a possible implementation, after multiple multi-scale back-projection feature fusion layers in the encoder perform multi-layer feature encoding on the speech signal to be synthesized, the input is fed into the residual convolutional feed-forward network. After the output features of the residual convolutional feed-forward network are processed by convolution, speech features are output; After multiple multi-scale back-projection feature fusion layers in the decoder perform multi-layer feature decoding on the speech features, the input is fed into the residual convolutional feed-forward network. After the output features of the residual convolutional feed-forward network are processed by convolution, the synthesized speech is output; Among them, the process of the residual convolutional feed-forward network processing the features input to this network includes: Perform convolution processing and dropout processing on the features input to the network to obtain extended features, perform depthwise separable convolution processing on the extended features to obtain channel-enhanced features; add the channel-enhanced features to the extended features to obtain residual-enhanced features; perform convolution processing on the residual-enhanced features and perform residual connection with the features input to the residual convolutional feedforward network to obtain the output features of the residual convolutional feedforward network.
[0034] By introducing depthwise separable convolution and residual connection, the local modeling ability and the temporal expression ability of speech are improved to enhance detailed information and reduce information redundancy.
[0035] In a possible implementation, call the convolutional layer in the encoder to extract the basic temporal features of the speech signal to be synthesized; After that, perform multi-layer feature extraction on the basic temporal features through multiple multi-scale back-projection feature fusion layers to obtain speech features.
[0036] The speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application, at the encoding end, the encoder uses multi-scale back-projection feature fusion layers to extract features from the speech signal, then performs back-projection mapping on the extracted multi-scale features, corrects them by feeding back residual information, further refines the extracted multi-scale features, and uses feature fusion blocks for fusion to improve the model's understanding of features of different scales. After the encoder performs downsampling, the features are enhanced through a convolutional feedforward network to learn the feature context information and obtain latent vectors. At the quantization end, multi-stage vector quantization is used to quantize the latent vectors, and the encoded indexes are packed into a binary byte stream for transmission. At the decoding end, the decoder obtains the quantization vectors of the codebook according to the indexes, and then synthesizes speech through upsampling, multi-scale back-projection feature fusion layers and residual convolutional feedforward networks. In addition, a discriminator is introduced during the training process to distinguish the authenticity of the synthesized speech, and through adversarial training, the generator is prompted to generate more realistic speech, thereby improving the quality of the reconstructed speech.
[0037] As Figure 1 shown, the specific implementation process of the speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of the present application includes: (1) Obtain the speech signal to be synthesized. Among them, the speech signal to be synthesized is a speech signal sampled at 8KHz, and this speech signal can be represented by where is the number of speech channels, T is the total number of sample points of the speech, where d is the duration of the speech signal, is the sampling rate of the speech signal.
[0038] (2) At the encoding end, multiple multi-scale back-projection feature fusion layers in the encoder are called to perform multi-layer feature encoding on the speech signal to be compressed, and speech features are obtained.
[0039] As Figure 3 shown, the encoder includes two convolutional layers, four multi-scale back-projection feature fusion layers, and a residual convolutional feed-forward network. The core is the introduction of the multi-scale back-projection feature fusion layer and the residual convolutional feed-forward network. Among them, both convolutional layers are one-dimensional convolutions.
[0040] The process of the encoder performing feature encoding on the speech signal to be synthesized to obtain speech features includes: First, the speech signal to be synthesized input to the encoder is convolved through one of the convolutional layers to obtain the basic temporal features of the speech signal to be synthesized; this convolutional layer is a one-dimensional convolution with 32 channels and a kernel size of 7.
[0041] Then, the basic temporal features are subjected to multi-layer feature encoding through four multi-scale back-projection feature fusion layers. For the first three multi-scale back-projection feature fusion layers, the output features of each multi-scale back-projection feature fusion layer are used as the input features of the next multi-scale back-projection feature fusion layer. The output features of the last multi-scale back-projection feature fusion layer are used as the input features of the residual convolutional feed-forward network.
[0042] Each multi-scale back-projection feature fusion layer includes three parts: a multi-scale convolution block, a back-projection mechanism, and a feature fusion block.
[0043] As Figure 4 shown, the multi-scale convolution block processes the features input to the multi-scale back-projection feature fusion layer through two different receptive field paths. One path convolves the feature x through a convolutional layer with a kernel size of 3 to obtain the first convolution-processed feature; the other path convolves the feature x through a convolutional layer with a kernel size of 5 to obtain the second convolution-processed feature. The convolution-processed features extracted from the two paths are concatenated in the channel dimension to obtain two concatenated features. Subsequently, one-dimensional convolution feature extraction is performed on the two concatenated features through convolutional layers (3 and 5) with matching sizes, and then one-dimensional convolution processing is performed on the features extracted from the two concatenated features to reduce the features back to the original dimension, generating multi-scale feature x1 and multi-scale feature x2.
[0044] The multi-scale features x1 and x2 are processed through the back-projection mechanism to correct the detailed features in the multi-scale features. As Figure 5As shown below, taking the multi-scale feature x1 as an example, the back-projection mechanism is described. A convolutional layer with normalization and causal convolution and a stride of 2 is used to downsample the multi-scale feature x1 to compress the length of the time dimension and obtain the downsampling result of the multi-scale feature x1. Subsequently, the corresponding transposed convolutional layer is used to upsample the downsampling result of the multi-scale feature x1 and attempt to restore it to a temporal length close to the original fused feature to obtain the upsampled feature. Then, the residual between the multi-scale feature x1 and the upsampled feature is calculated, and the residual is input into the residual block to further model the error signal and perform non-linear correction to obtain the residual R1.
[0045] Subsequently, the multi-scale feature x1 and the residual R1 corrected by the back-projection residual are fed into the feature fusion block to dynamically allocate different weights and adjust the importance of the multi-scale features. Specifically, as Figure 6 shown, first, the two features of the multi-scale feature x1 and the residual R1 are added as the main input feature U, and then the two features are stacked along the last dimension into a new four-dimensional tensor [B, C, T, 2], where B is the batch size, C is the number of channels, T is the time dimension, and 2 represents two branches. Global average pooling and global max pooling are respectively performed on the main input feature U in the channel dimension to extract the global semantic feature and the local prominent feature and expand the last dimension to make the shape [B, C, 1] to obtain the average pooling feature and the max pooling feature. Subsequently, the average pooling feature and the max pooling feature are concatenated, reduced to the original dimension through one-dimensional convolution, and the feature is feature-mapped with PReLU activation to form the hidden vector z. Then, for each of the two branches, the attention vector is calculated through the corresponding convolutional layer, and then the attention vectors of each branch are concatenated along the branch dimension to form a tensor of [B, 2, C]. Subsequently, softmax normalization is performed on the attention vectors of the two branches to make the sum of the weights between the branches equal to 1 to enhance the discriminability and obtain the multi-scale feature weight and the residual weight. Finally, the multi-scale feature x1 and the residual R1 are weighted and summed according to the above weights to generate the fusion result T1. For the multi-scale feature x2 and the residual R2, the same steps are taken to obtain the fusion result T2. Then, the fusion result T1 and the fusion result T2 are concatenated in the channel dimension and the number of channels is reduced back to the original dimension through one-dimensional convolution to obtain the feature T. The identity mapping I of the input multi-scale back-projection feature fusion layer feature is obtained through skip connection. Then, T and the identity mapping I are connected by residual connection to obtain the output feature of the multi-scale back-projection feature fusion layer.
[0046] Behind each multi-scale back-projection feature fusion layer in the encoder, there is immediately followed by a one-dimensional convolution with downsampling, with a stride of S and a convolutional kernel twice the stride S. During the downsampling process, the number of channels is doubled. Where S = (1, 4, 8, 10).
[0047] After the output features of the last multi-scale back-projection feature fusion layer in the encoder are downsampled by a one-dimensional convolution, they are input into the residual convolutional feed-forward network, which further performs non-linear modeling and local receptive field enhancement on the encoded speech features. Its structure mainly consists of two 1×1 convolutional layers, a depthwise separable convolution, an activation function, and a dropout mechanism, and a residual connection is introduced to enhance the training stability. Let the feature input into the residual convolutional feed-forward network be B, the dimension of feature B be C, and the time step be T, then the input tensor size of the residual convolutional feed-forward network is [B, C, T].
[0048] As Figure 7 shown, the processing process of the residual convolutional feed-forward network for feature B includes: First, a 1×1 convolutional layer is used to perform convolutional processing on feature B, expanding the number of channels of feature B from C to 2C to achieve the upsampling operation of the channel dimension. After convolution, dropout is added to prevent overfitting, and the expanded feature B1 is obtained. Subsequently, a depthwise separable convolution processing is performed on the expanded feature B1. The convolutional kernel size of this depthwise separable convolution processing is 3, the padding is 1, and the number of groups is equal to the number of channels (i.e., each channel is independently convolved) to enhance the local perception ability of each channel, and the channel-enhanced feature B* is obtained. Then, a residual enhancement strategy is adopted, and the expanded feature B1 and the channel-enhanced feature B* are added to obtain the residual-enhanced feature B2. After that, the GELU activation function is used to enhance the non-linear expression ability of the residual-enhanced feature B2, and then the number of channels is restored to the original dimension C through a 1×1 convolutional layer to complete the channel compression mapping. Finally, dropout is used again and a residual connection is made with the input feature B to obtain the output feature Y of the residual convolutional feed-forward network. The design parameters of the residual convolutional feed-forward network are as follows: the input channel number C = 512, the expansion factor = 2, the corresponding intermediate channel is 2C = 1024, and the dropout ratio is 0.2. Then, through a one-dimensional convolutional layer with a kernel size of 7 and 512 channels, the speech features output by the encoder are obtained.
[0049] Quantize the speech features output by the encoder at the quantization end.
[0050] The quantization end includes a multi-stage vector quantization system. There are M codebooks in this multi-stage vector quantization system, M = 256. The vector dimension C of each codebook is 512, and the number N of potential frame vectors is 25. Each codebook can encode bits. To ensure the quantization performance and convergence speed, the K-Means clustering algorithm is used to partition the latent vector space encoded by the encoder to obtain the initialized codebook set , and the set of speech features input to the quantization end is , corresponding to the low-dimensional feature representations obtained by processing each frame through the encoder. In the embodiments of the present application, a multi-stage quantizer is used to quantize the speech features output by the encoder. The multi-stage quantizer consists of several stages, and each stage quantizes the quantization residuals of the previous stage step by step to approximate the original vector. Each stage of the quantizer uses the Euclidean distance metric as the similarity index to quickly and accurately locate the most matching codebook vector, thereby ensuring the minimization of the quantization error. The first-stage quantizer finds the quantization vector in the codebook set C that is closest to the speech feature z , and the second-stage quantizer quantizes the residuals of the first-stage quantization vector to obtain the second-stage quantization vector , ; The third-stage quantizer continues to quantize the remaining residuals, that is, the third-stage quantizer quantizes the residuals of the second-stage quantization vector and the first-stage quantization vector to obtain the second-stage quantization vector , . The remaining quantizers are iterated in turn until all quantization is completed. The quantization vectors are stored as indices; the obtained quantization vector indices are compressed and transmitted to the decoding end.
[0051] Preferably, the quantized speech features are packed into a binary byte stream and transmitted to the decoding end.
[0052] At the decoding end, the dequantized features are obtained in the multi-stage quantizer according to the indices, and are added in turn to obtain the final latent frame vector . The dequantized speech features are input into the decoder to synthesize speech. The decoder is symmetric to the encoder structure. First, a one-dimensional convolution with a kernel size of 7 and 512 channels is performed. Followed by four multi-scale back-projection feature fusion layers. Different from the encoding end, each multi-scale back-projection feature fusion layer is followed by a one-dimensional transposed convolution with a stride of E and a kernel twice the stride of E. During the upsampling process, the number of channels is halved, where E = (10, 8, 4, 1). Then, local feature enhancement is performed through a residual convolutional feed-forward network, and then a one-dimensional convolution with a kernel of 7 and 32 channels is performed. Finally, the reconstructed speech, that is, the synthesized speech, is obtained.
[0053] To further improve the speech synthesis quality of the generator, in the training stage, the embodiments of this application use a multi-scale STFT discriminator (MS-STFT) and a multi-period discriminator (MPD) to discriminate the quality of the synthesized speech, which is only used in the training stage. The generator refers to the overall network model composed of an encoder, a quantizer, and a decoder. Specifically, the main role of the discriminator is to judge the quality of the generator's generation effect, and the generator is trained to better generate speech signals. The input of the discriminator is the input speech of the encoder and the synthesized speech output by the decoder. The input speech is the original speech. By predicting the authenticity of the synthesized speech, it learns how to distinguish the original speech from the synthesized speech and feeds back the prediction result of the synthesized speech to the generator, competing and collaborating with each other in an adversarial training manner. The MS-STFT discriminator consists of multiple sub-discriminators with the same structure, which act on the complex STFT representations at different scales. Specifically, it uses 5 sub-discriminators, corresponding to 5 short-time Fourier transform (STFT) window lengths in sequence: 2048, 1024, 512, 256, and 128, and the sliding step sizes are 512, 256, 128, 64, and 32 respectively. Each sub-discriminator acts on different time scales and frequency bandwidths to fully extract the local details and global structure information of the speech signal. After the input speech signal undergoes the STFT transform, the real part and the imaginary part of its complex spectrum are respectively extracted and concatenated in the channel dimension as a two-dimensional input tensor with the shape of [B, 2, F, T], where F is the frequency dimension and T is the number of time frames. This complex spectrum is then input into a two-dimensional convolutional network with a consistent structure for discrimination. The structure of this two-dimensional convolutional network includes: a two-dimensional convolutional layer with a kernel size of 3×8 and 32 channels; several two-dimensional convolutional layers with different time dimension dilation rates D and a frequency dimension stride of 2; and a final two-dimensional convolutional layer (kernel size 3×3, stride 1×1) for judging authenticity. The discriminator at each scale will output a set of true / false classifications and intermediate feature maps, and the results of all sub-discriminators are jointly used for training.
[0054] MPD is a hybrid architecture composed of multiple sub-discriminators. Each sub-discriminator focuses on modeling speech patterns under a specific period p, thereby enhancing the perception ability of different periodic features. Specifically, the sub-discriminators capture different implicit structures by looking at different parts of the input audio. The period p is set to [2, 3, 5, 7, 11] to avoid overlap and structural redundancy. First, the one-dimensional original speech of length T is reshaped into a two-dimensional matrix with height T / p and width p. Then, two-dimensional convolution is applied to the reshaped two-dimensional matrix, and the LeakyReLU activation function is connected after each layer of convolution. All convolutional layers operate in the time dimension, and the kernel size on the width axis is limited to 1 to independently process periodic samples. Each sub-discriminator outputs the feature maps of the real and generated samples after each layer of convolution, and the outputs of all sub-discriminators will be unified for training.
[0055] In summary, MS-STFT and MPD form a complementary discrimination mechanism in the frequency domain and time domain, providing a more stable and multi-angle training signal for the decoder, which helps to improve the naturalness and clarity of the synthesized speech.
[0056] The speech compression method based on multi-scale back-projection feature fusion provided by the embodiments of this application introduces multi-scale back-projection feature fusion technology and residual convolutional feed-forward network technology into the encoder and decoder. It uses multi-scale feature extraction and back-projection residual correction to learn richer and more complete features, and uses the residual convolutional feed-forward network to locally enhance the information after upsampling and downsampling to improve the quality of the synthesized speech. To achieve the above purpose, a multi-scale back-projection feature fusion layer and a residual convolutional feed-forward network are proposed. The multi-scale back-projection feature fusion layer mainly includes three parts: multi-scale feature extraction, back-projection correction mechanism, and selective feature fusion. The multi-scale convolution captures speech detail features at different scales through cross-learning; the back-projection mechanism compresses, reconstructs, and residually feedbacks the captured detail features to obtain finer features; the feature fusion block adaptively adjusts the weights from different branches and fuses them. The residual convolutional feed-forward network improves the local modeling ability of the model and the temporal expression ability of speech by introducing depthwise separable convolution and residual connection in the feed-forward structure, enhancing detail information and reducing information redundancy. This solution improves the quality of speech synthesis by introducing these two technologies of multi-scale back-projection feature fusion and residual convolutional feed-forward network.
[0057] The embodiments of this application also provide a speech compression system based on multi-scale back-projection feature fusion, including: An encoding module, configured to call multiple multi-scale back-projection feature fusion layers in the encoder to perform multi-layer feature encoding on the speech signal to be synthesized, and obtain speech features; A decoding module, configured to call multiple multi-scale back-projection feature fusion layers in a decoder to perform multi-layer feature decoding on speech features to obtain synthesized speech; Among them, the process of the multi-scale back-projection feature fusion layer encoding or decoding the input features includes: the multi-scale back-projection feature fusion layer uses convolutional kernels of different scales to perform cross-learning on the input features to obtain multi-scale features, performs back-projection on the extracted multi-scale features respectively to obtain back-projection features, fuses the back-projection features and the multi-scale features to obtain a fusion result, and fuses the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
[0058] The embodiment of the present application further provides a computer device, which includes: A processor, adapted to execute a computer program; A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, it implements the speech compression method based on multi-scale back-projection feature fusion provided by the embodiment of the present application.
[0059] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the speech compression method based on multi-scale back-projection feature fusion provided by the embodiment of the present application.
[0060] The embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the speech compression method based on multi-scale back-projection feature fusion provided by the embodiment of the present application.
[0061] The method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0062] Those of ordinary skill in the art can realize that, in combination with the units and algorithm steps of the examples described in this embodiment, they can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0063] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, they do not limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A speech compression method based on multi-scale back-projection feature fusion, characterized in that Including: Calling multiple multi-scale back-projection feature fusion layers in the encoder to perform multi-layer feature encoding on the speech signal to be synthesized, obtaining speech features; Calling multiple multi-scale back-projection feature fusion layers in the decoder to perform multi-layer feature decoding on the speech features, obtaining the synthesized speech; Among them, the process of the multi-scale back-projection feature fusion layer encoding or decoding the input features includes: using convolution kernels of different scales to perform cross learning on the input features to obtain multi-scale features, performing back-projection on the extracted multi-scale features respectively to obtain back-projection features, fusing the back-projection features and the multi-scale features to obtain a fusion result, and fusing the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
2. The speech compression method based on multi-scale back-projection feature fusion according to claim 1, characterized in that The multi-scale back-projection feature fusion layer performs multi-path convolution processing on the input features using convolution kernels of different scales, splices the features after convolution processing of all paths to obtain the spliced features of multiple paths; and extracts features from each spliced feature to obtain multi-scale features.
3. The speech compression method based on multi-scale back-projection feature fusion according to claim 1, wherein The multi-scale back-projection feature fusion layer first downsamples and then upsamples each multi-scale feature to obtain the upsampled features, calculates the residuals between the multi-scale features and their upsampled features, and performs weighted fusion on the multi-scale features and the residuals to obtain a fusion result; splices all the fusion results to obtain the spliced features, and performs residual connection on the spliced features and the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
4. The speech compression method based on multi-scale back-projection feature fusion according to claim 3, wherein The weights used for weighted fusion of the multi-scale features and the residuals are obtained by calculating the attention vectors of the multi-scale features and the residuals and normalizing the two attention vectors.
5. The speech compression method based on multi-scale back-projection feature fusion according to claim 1, wherein After multiple multi-scale back-projection feature fusion layers in the encoder perform multi-layer feature encoding on the speech signal to be synthesized, the input is fed into the residual convolutional feed-forward network. After the output features of the residual convolutional feed-forward network are processed by convolution, the speech features are output; After multiple multi-scale back-projection feature fusion layers in the decoder perform multi-layer feature decoding on the speech features, the input is fed into the residual convolutional feed-forward network. After the output features of the residual convolutional feed-forward network are processed by convolution, the synthesized speech is output; Among them, the process of the residual convolutional feed-forward network processing the features input to the network includes: Performing convolution processing and dropout processing on the features input to the network to obtain the extended features, performing depthwise separable convolution processing on the extended features to obtain the channel-enhanced features; adding the channel-enhanced features to the extended features to obtain the residual-enhanced features; performing convolution processing on the residual-enhanced features and performing residual connection with the features input to the residual convolutional feed-forward network to obtain the output features of the residual convolutional feed-forward network.
6. The speech compression method based on multi-scale back-projection feature fusion according to claim 1, wherein Calling the convolutional layer in the encoder to extract the basic temporal features of the speech signal to be synthesized; After that, multiple multi-scale back-projection feature fusion layers are used to perform multi-layer feature extraction on the basic temporal features to obtain speech features.
7. A speech compression system based on multi-scale back-projection feature fusion, characterized in that, Including: An encoding module for calling multiple multi-scale back-projection feature fusion layers in the encoder to perform multi-layer feature encoding on the speech signal to be synthesized, obtaining speech features; A decoding module, configured to call multiple multi-scale back-projection feature fusion layers in a decoder to perform multi-layer feature decoding on speech features to obtain synthesized speech; Among them, the process of the multi-scale back-projection feature fusion layer encoding or decoding input features includes: the multi-scale back-projection feature fusion layer uses convolutional kernels of different scales to perform cross-learning on the input features to obtain multi-scale features, performs back-projection on the extracted multi-scale features respectively to obtain back-projection features, fuses the back-projection features and the multi-scale features to obtain a fusion result, and fuses the fusion result with the features input to the multi-scale back-projection feature fusion layer to obtain the output features of the multi-scale back-projection feature fusion layer.
8. An electronic device, characterized in that, The device includes: A processor, adapted to execute a computer program; A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, it implements the speech compression method based on multi-scale back-projection feature fusion according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the speech compression method based on multi-scale back-projection feature fusion according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the speech compression method based on multi-scale back-projection feature fusion according to any one of claims 1-6.
Citation Information
Patent Citations
Image compression framework combining super-resolution and residual coding technology
CN107181949A
Image super-resolution reconstruction method based on back projection attention network
CN112215755A
Voice compression method and system based on deep learning and vector prediction
CN117423348A
End-to-end speech coding method and system based on selective back projection feature fusion
CN118136024A
Voice compression method and system based on multi-scale residual attention
CN118335092A
Cited By
Cerebral stroke early diagnosis model construction method and device, electronic equipment and storage medium
CN120745732A
Voice coding and decoding method based on principal component analysis and multi-scale depth attention
CN121617406A
Voice compression method and system based on feature weighted residual vector quantization
CN121789696A