Expandable bone conduction voice signal transmission method based on diffusion probability model
Through the bone conduction speech codec network based on the diffusion probability model, the high-frequency loss and low bit rate encoding and decoding problems of bone conduction speech signals are solved, and high-quality voice signal transmission and adaptive packet transmission are realized.
Patent Information
- Application Number
- CN202510190617.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Due to the low-pass filtering characteristics of the human skeleton, high-frequency components are lost, resulting in a reduced speech clarity, which makes it difficult for traditional speech codecs to effectively code at lower bit rates.
A bone conduction speech codec network based on diffusion probability model is adopted. By combining the generative model and the diffusion probability model, the characteristics of the speech signal are extracted and high-frequency reconstruction is carried out to achieve coding and enhancement of the speech signal.
The calculation complexity of the speech encoding and decoding process is reduced, high-quality speech encoding and decoding is realized at a lower bit rate, and through the scalable packet transmission mechanism, it adapts to network load changes and ensures high-quality transmission of voice signals.
Smart Images

Figure CN120164485A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bone-conducted voice signal transmission, and particularly to a scalable bone-conducted voice signal transmission method based on a diffusion probability model. Background Art
[0002] Air-conducted voice records voice by converting sound waves propagated in the air into electrical signals, and bone-conducted voice signals record voice by converting vibrations around the speaker's skull into electrical signals, which means that in a noisy environment, such as public transportation or a noisy workplace, using bone-conducted devices can transmit voice information more clearly and accurately. However, due to the limitations of its propagation mechanism, bone-conducted voice has the following disadvantages: Since the human bone is equivalent to a low-pass filter, the high-frequency components in the bone-conducted voice signal are lost, and its energy distribution is usually limited below 2 kHz. This will greatly reduce the clarity of the voice and make the bone-conducted voice sound dull. During voice communication, the sending end needs to encode and compress the bone-conducted voice signal into a binary bitstream and send it to the receiving end. In order to make full use of the advantages of bone-conducted voice and overcome its limitations, some research efforts have tried to develop high-frequency compensation technologies for bone-conducted voice signals and use existing voice codecs to encode the enhanced bone-conducted voice. This enables bone-conducted voice to be more widely applied.
[0003] Traditional voice codecs are used to encode the enhanced bone-conducted voice obtained by the above method for transmission over a communication channel, and then decoded and played by the receiving end. However, this process includes two steps: enhancement of bone-conducted voice and encoding of the enhanced voice. The two independent steps result in a relatively high computational complexity. Traditional voice codecs are also unable to perform voice encoding and decoding at a low bit rate. Summary of the Invention
[0004] According to the problems existing in the prior art, the present invention discloses a scalable bone-conducted voice signal transmission method based on a diffusion probability model, which specifically includes the following steps:
[0005] Collect the bone-conducted voice signal to be sent and the air-conducted voice signal corresponding to the bone-conducted voice, and establish a training set and a test set based on the bone-conducted voice signal and the air-conducted voice signal;
[0006] Construct a bone-conducted voice codec network based on a diffusion probability model. The bone-conducted voice codec network includes a voice encoder, a residual vector quantizer, and a voice decoder. The voice encoder includes a generative model and a diffusion probability model, wherein the generative model includes a feature extraction module, a high-frequency reconstruction module, and an overall optimization module;
[0007] Train the bone conduction speech codec network on the divided training set and update the network parameters. Use the backpropagation algorithm to transmit the gradient values of the network, and repeatedly train and update the network parameters to control the generated speech signal to approximate the air conduction speech signal, thereby obtaining the optimal model parameters and completing the training process of the bone conduction speech codec;
[0008] Enter the deployment stage, control the speech encoder and speech decoder separately, and input the bone conduction speech signal in the test set into the bone conduction speech codec network loaded with the optimal model parameters to obtain the binary bitstream to be sent;
[0009] Transmit the binary bitstream from the sending end to the receiving end, and use the speech decoder to decode and synthesize the air conduction speech;
[0010] Play the continuous air conduction speech through the speaker.
[0011] Furthermore, the generative model generates a roughly enhanced speech signal from the bone conduction speech signal with a sampling rate of 16 kHz, and inputs this speech into the diffusion probability model for enhancement processing to generate an air conduction speech signal with a sampling rate of 48 kHz. The audio signal with a sampling rate of 48 kHz is divided into five subbands by a subband filter, namely 0 - 4 kHz, 4 - 8 kHz, 8 - 12 kHz, 12 - 16 kHz, and 16 - 20 kHz. Each subband undergoes an improved discrete cosine transform as the input speech feature of the residual vector quantizer.
[0012] Furthermore, the residual vector quantizer includes five residual vector quantizer branches with exactly the same structure but different model parameters. The speech features generated by the speech encoder are respectively input into the five residual vector quantizer branches for quantization processing to obtain five quantization results.
[0013] Furthermore, the feature extraction module includes two types of convolutions: 3×3 convolution + ReLU and 1×1 convolution + ReLU, where convolution + ReLU means that the convolution is followed by the ReLU activation function. Specifically, 3×3 convolution + ReLU is used when the layer number is odd, and 1×1 convolution + ReLU is used when the layer number is even. Residual learning is used to integrate hierarchical information and retain shallow information, and this process is expressed as
[0014] O kf = R((Conv1(Conv2(O full )))+Conv3(Conv4(O key ))) (1)
[0015] where, R represents the ReLU function, Conv i represents the convolution operation, Ofull and O key represent the global and local key features respectively, and O kf represents the output information of the feature extraction module.
[0016] Furthermore, the high-frequency reconstruction module is based on a generator model with a U-Net structure, adopts symmetric encoder blocks and decoder blocks, and has a skip connection structure, where the layout of the encoder is symmetric with that of the decoder, and the skip connections are added between each encoder block and its symmetric decoder block. The outermost skip connection only connects the necessary voice channels of the output. The encoder block and the decoder block each contain four blocks and are respectively sandwiched between two conventional convolutional layers. The encoder block includes four downsampling layers, where the downsampling multiples are 2, 2, 8, and 8 respectively, while the decoder block performs upsampling in the reverse order; each time of downsampling, the number of channels doubles; each time of upsampling, the number of channels halves. Each decoder block consists of an upsampling layer and three residual units, and each unit contains one-dimensional convolutions with dilation rates of 1, 3, and 9 respectively. The encoder block corresponds to the decoder block and is composed of the same residual units, and realizes downsampling through one-dimensional convolution.
[0017] Furthermore, the overall optimization module in the generation model contains 1 layer of ReLU, 4 layers of Conv+ReLU, and 1 layer of Conv; each Conv+ReLU layer consists of a convolutional layer and a ReLU activation function, where the filter size of the convolutional layer is 3×3, and the number of input and output channels is 64. The filter size of the final convolutional layer is 3×3, and the number of input and output channels is 64 and 1 respectively. This process is expressed as
[0018] O rb = Cv1(R(Cv2(R(Cv3(R(Cv4(R(Cv5(O hf ))))))))) (2)
[0019] where, R represents the ReLU function, and Cv i represents the convolution operation, and O hf represents the output of the high-frequency reconstruction module, and O rb represents the output of the feature extraction module.
[0020] Furthermore, the diffusion probability model is a generative model of a generative trend, and generates complex data from a simple noise distribution (such as a normal distribution) through iterative Monte Carlo Markov chain sampling. The work of the diffusion probability model includes two processes: the forward process and the reverse process. Among them, the forward process is a step-by-step generation mechanism based on a probability model, and sequentially generates subsequent states starting from the initial state according to the time step or sequence dependence relationship, which can be denoted as q(y 1:N |y), indicating from y1 to y under the given condition y NThe joint distribution, i.e., The conditional probability q(y n |y n-1 ) represents the likelihood of the current state y n-1 given the previous state y n . The fixed noise variance is λ 1:N = [λ1, λ2,..., λ N which is set from 0.001 to 0.02 by linear interpolation, and this process is composed of the multiplication of multiple Gaussian distributions. The forward process can be expressed as
[0021]
[0022] where y represents the input signal, I represents the identity matrix which plays a scaling role in this formula to ensure that the covariance matrix of the Gaussian distribution is a diagonal matrix and each element is equal to λ n . This setting allows the model to control the noise intensity through different scaling factors, thus providing greater flexibility in adjusting the variability of the generated data.
[0023] The inverse process is represented by the parameter ω and is defined as the reverse process of diffusion. The model samples y from the Gaussian distribution n , which represents a random noise. The inverse process is defined as the product of the transition probabilities. Through the transition probabilities of the inverse process, the model gradually removes the noise to obtain samples close to the real data. The transition probability p ω (y n-1 |y n ) represents the conditional probability of restoring from the current state y n to the previous state y n-1 . This probability distribution is modeled as a multivariate Gaussian distribution.
[0024]
[0025] where μ ω (y n , n) and are the mean and variance of y n-1 estimated by the model respectively. The latent variable y n-1 is sampled from the distribution p(y n-1 |y n ).
[0026] The reverse process can be expressed as
[0027]
[0028] where ∈ ω is the noise estimated by the model. where φ n is defined as 1 - λ n , is defined as
[0029] Due to the above - mentioned technical solution, a scalable bone - conduction voice signal transmission method disclosed by the present invention has the following advantages: By adopting a deep - learning model, a neural voice codec architecture suitable for bone - conduction voice is proposed, which can achieve voice enhancement while performing voice coding, reducing the usage complexity.
[0030] The idea of scalability is introduced. During the transmission process, data packets can be dynamically adjusted according to the change of network load. The transmitted data packets are divided into five levels: code1, code2, code3, code4, and code5, where code1 is the base layer, and code2, code3, code4, and code5 are the enhancement layers. When the network load is low, data packets can be transmitted completely. When the network load is high, on the premise of ensuring that the base layer is not lost, data packets are discarded sequentially starting from code5 to meet the network transmission requirements. Even if some data packets are lost, high - quality voice signals can still be decoded, thus realizing the scalable bone - conduction voice signal transmission method.
[0031] The residual vector quantizer adopts five residual vector quantizers with exactly the same structure but different model parameters. The voice features generated by the encoder are simultaneously sent into five residual vector quantizers with exactly the same structure but different model parameters for quantization processing to obtain five quantization results.
[0032] The voice decoder part adopts a generative adversarial network vocoder for high - fidelity voice synthesis as the decoder architecture of the scalable neural bone - conduction voice codec. The generative adversarial network for high - fidelity voice synthesis is one of the most advanced neural vocoders currently, which can achieve highly realistic voice synthesis.
[0033] The feature extraction module is adopted in the architecture. This module first extracts the low - frequency features of the voice signal and gradually aggregates these features through residual learning, which can enhance the network's attention to the low - frequency voice signal and improve the memory ability of the shallow layer to the deep layer. In addition, residual connections are also an effective way to promote the network's feature learning and fusion ability because they can improve performance without introducing too many additional parameters.
[0034] The high - frequency reconstruction module is adopted in the architecture. This module first uses a convolutional network to convert the global features and the extracted key features into non - linear. Then, residual learning is used to integrate the global and local features. Next, the fused features are input into the encoder - decoder network to effectively reconstruct the high - frequency components of the voice signal.
[0035] The overall optimization module is adopted in the architecture to further refine the high-frequency part of the speech signal reconstructed by the high-frequency reconstruction module, thereby improving the speech quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of the scalable bone conduction speech signal transmission method based on the diffusion probability model according to the embodiment of the present invention;
[0038] Figure 2 It is a schematic diagram of the architecture of the bone conduction speech codec network according to the embodiment of the present invention;
[0039] Figure 3 It is a schematic diagram of the architecture of the bone conduction speech encoder according to the embodiment of the present invention;
[0040] Figure 4 It is a schematic diagram of the architecture of the generator model according to the embodiment of the present invention;
[0041] Figure 5 It is a schematic diagram of the architecture of the feature extraction module according to the embodiment of the present invention;
[0042] Figure 6 It is a schematic diagram of the architecture of the high-frequency reconstruction module according to the embodiment of the present invention
[0043] Figure 7 It is a schematic diagram of the architecture of the overall optimization module according to the embodiment of the present invention;
[0044] Figure 8 It is a schematic diagram of the architecture of the diffusion probability model according to the embodiment of the present invention;
[0045] Figure 9 It is the STOI objective evaluation index of the decoded air conduction speech signal under the low bit rate condition according to the embodiment of the present invention;
[0046] Figure 10 It is the STOI objective evaluation index of the decoded air conduction speech signal under the medium bit rate condition according to the embodiment of the present invention;
[0047] Figure 11 It is the STOI objective evaluation index of the decoded air conduction speech signal under the high bit rate condition according to the embodiment of the present invention;
[0048] Figure 12 This is the subjective evaluation index (MOS) of the decoded air-conducted speech signal in the embodiments of the present invention; Detailed implementation manners
[0049] To make the technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention:
[0050] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0052] As Figure 1 shown, a scalable bone-conducted speech signal transmission method based on a diffusion probability model,
[0053] S1: Collect the bone-conducted speech signal to be transmitted and the corresponding air-conducted speech signal of the bone-conducted speech, and establish a training set and a test set based on the bone-conducted speech signal and the air-conducted speech signal.
[0054] Specifically, in implementation, the training set and the test set are obtained from the ESMB dataset. This dataset includes 128 hours of speech spoken by 131 male and 156 female speakers. Each speaker read approximately 20 minutes of Chinese speech, and the speech sampling frequency was 16 kHz stereo, with the left channel being the air-conducted speech signal and the right channel being the bone-conducted speech signal.
[0055] S2: Construct a bone conduction voice codec network based on a diffusion probability model, including a voice encoder, a residual vector quantizer, and a voice decoder. The voice encoder includes a generative model and a diffusion probability model. The generative model includes a feature extraction module, a high-frequency reconstruction module, and an overall optimization module. Figure 2 Shows the overall structure of a scalable bone conduction voice codec network based on a diffusion probability model.
[0056] The voice encoder network is as Figure 3 shown. The voice encoder includes a generative model and a diffusion probability model. The generative model is used to generate a roughly enhanced voice signal from a bone conduction voice signal with a sampling frequency of 16 kHz. Subsequently, the voice is input into the diffusion model for further enhancement to generate an air conduction voice signal with a sampling frequency of 48 kHz. Then, the audio with a sampling frequency of 48 kHz is divided into five sub-bands by a sub-band filter, namely 0 - 4 kHz, 4 - 8 kHz, 8 - 12 kHz, 12 - 16 kHz, and 16 - 20 kHz. Finally, each sub-band undergoes a modified discrete cosine transform as the voice feature for subsequent residual vector quantization.
[0057] The quantizer part includes five residual vector quantizer branches with exactly the same structure but different model parameters. The voice features generated by the encoder are simultaneously sent into five residual vector quantizers with exactly the same structure but different model parameters for quantization processing to obtain five quantization results.
[0058] The voice decoder architecture uses a generative adversarial network vocoder for high-fidelity voice synthesis as the decoder architecture of the scalable neural bone conduction voice codec. The generative adversarial network vocoder for high-fidelity voice synthesis is one of the state-of-the-art neural vocoders and can achieve highly realistic voice synthesis.
[0059] As Figure 4 shown, the generative model includes a feature extraction module, a high-frequency reconstruction module, and an overall optimization module. The feature extraction module is as Figure 5 shown. The feature extraction module involves two types of convolutions: 3×3 convolution + ReLU and 1×1 convolution + ReLU, where convolution + ReLU means that the convolution is followed by a ReLU activation function. Specifically, 3×3 convolution + ReLU is used when the layer number is odd, while 1×1 convolution + ReLU is used when the layer number is even. This heterogeneous method reduces the overall computational cost and memory consumption of the network. At the same time, to retain shallow information, residual learning is used to integrate hierarchical information, and this process can be expressed as
[0060] O kf =R( ( Conv1(Conv2(O full )))+Conv3(Conv4(Okey ))) (11)
[0061] Among them, R represents the ReLU function, Conv i represents the convolution operation, O full and O key represent the global and local key features respectively, and O kf represents the output of the feature extraction module.
[0062] As Figure 6 shown, the high-frequency reconstruction module is based on a generator model with a U-Net structure, adopts a symmetric encoder-decoder network, and has skip connections. The layout of the encoder is symmetric with that of the decoder, and skip connections are added between each encoder block and its symmetric decoder block. The outermost skip connection only connects the necessary speech channels of the output. Both the encoder and the decoder contain four blocks and are respectively sandwiched between two conventional convolutional layers. The encoder has four downsampling layers with downsampling multiples of 2, 2, 8, and 8 respectively, while the decoder performs upsampling in the reverse order. Each time downsampling is performed, the number of channels doubles; each time upsampling is performed, the number of channels is halved. Each decoder block consists of an upsampling layer and three residual units, and each unit contains one-dimensional convolutions with dilation rates of 1, 3, and 9 respectively. The encoder blocks correspond to the decoder blocks and are composed of the same residual units, and downsampling is achieved through one-dimensional convolutions.
[0063] As Figure 7 shown, the overall optimization module contains 1 layer of ReLU, 4 layers of Conv+ReLU, and 1 layer of Conv. Each Conv+ReLU layer consists of a convolutional layer and a ReLU activation function, where the filter size of the convolutional layer is 3×3, and the number of input and output channels is 64. The filter size of the final convolutional layer is 3×3, and the number of input and output channels is 64 and 1 respectively. This process can be expressed as
[0064] O rb = Cv1(R(Cv2(R(Cv3(R(Cv4(R(Cv5(O hf ))))))))) (12)
[0065] Among them, R represents the ReLU function, Cv i represents the convolution operation, O hf represents the output of the high-frequency reconstruction module, and O rb represents the output of the feature extraction module.
[0066] As Figure 8As shown, the diffusion probability model is a generative model that generates trends. Through iterative Monte Carlo Markov chain sampling, it generates complex data from a simple noise distribution (such as a normal distribution). The operation of the diffusion probability model involves two processes: the forward process and the reverse process. Among them, the forward process is a step-by-step generation mechanism based on a probability model. According to the time steps or sequence dependencies, it sequentially generates subsequent states starting from the initial state, which can be denoted as q(y 1:N |y), representing the joint distribution from y1 to y N under the given condition y, that is the conditional probability q(y n |y n-1 ) represents the likelihood of the current state y n-1 after knowing the previous state y n . Among them, the fixed noise variance is λ 1:N = [λ1, λ2,..., λ N is set by linear interpolation from 0.001 to 0.02, and this process is composed of the multiplication of multiple Gaussian distributions. The forward process can be expressed as
[0067]
[0068] where y represents the input signal, and I represents the identity matrix, which plays a scaling role in this formula to ensure that the covariance matrix of the Gaussian distribution is a diagonal matrix, and each element is equal to λ n . This setting allows the model to control the noise intensity through different scaling factors, thus providing greater flexibility in adjusting the variability of the generated data.
[0069] The inverse process is represented by the parameter ω and is defined as the reverse process of diffusion. The model samples y from the Gaussian distribution n , which represents a random noise. The inverse process is defined as the product of transition probabilities. Through the transition probabilities of the inverse process, the model gradually removes the noise to obtain samples close to the real data. The transition probability p ω (y n-1 |y n ) represents the conditional probability of restoring from the current state y n to the previous state y n-1 . This probability distribution is modeled as a multivariate Gaussian distribution.
[0070]
[0071] Among them, μ ω (y n , n) and σ n 2 are the estimates of y by the model respectivelyn-1 The mean and variance. The latent variable y n-1 is sampled from the distribution p(y n-1 |y n ).
[0072] The reverse process can be expressed as
[0073]
[0074] where, ∈ ω is the noise estimated by the model, where φ n is defined as 1 - λ n , is defined as
[0075] The model proposed in the present invention operates directly in the time domain, eliminating the need to extract any speech features. The proposed diffusion model takes the downsampled signal y and the noise level s as inputs, and outputs the estimated diffusion noise 10 residual layers and 64 channels are used, and a U-Net architecture is introduced after the dilated convolution in both directions of each residual layer. The dilated convolution in both directions uses a convolution kernel size of 3 and a dilation period of [1, 2,..., 512]×3. The U-Net architecture has achieved good results in tasks such as speech super-resolution and speech enhancement. However, to the best of our knowledge, U-Net has not been integrated into the field of diffusion probability models. Therefore, in order to further improve the ability of the diffusion probability model in speech super-resolution, a new diffusion model architecture combining the two is proposed.
[0076] S3: Construct the loss function of the network, train the scalable bone conduction speech codec network based on the diffusion probability model constructed in S2 on the divided training set, and update the parameters of the proposed network through the loss function. The parameter update specifically refers to using the backpropagation algorithm to transmit gradients and iterating repeatedly to reduce the error, so that the predicted depth approximates the air conduction speech signal, and finally obtain the optimal model parameters.
[0077] In this embodiment, the training mode of the generative adversarial network is adopted, and two different discriminators are used: a multi-resolution convolutional discriminator that takes a single waveform as input; a discriminator based on STFT that takes the complex-valued STFT of the input waveform as input.
[0078] The multi-resolution convolutional discriminator uses three discriminators with the same structure applied to audio inputs at different resolutions: the original, 2x downsampled, and 4x downsampled. Each discriminator consists of an initial ordinary convolution followed by four grouped convolutions, each group of size 4, with a downsampling factor of 4 and a channel multiplier of 4, up to 1024 output channels.
[0079] The STFT-based discriminator involves a 2D convolutional layer followed by a series of 2D convolutions. The first 2D convolution has a kernel size of 7×7 and 32 channels. This is followed by a series of residual blocks. Each block starts with a 3×3 convolution, then is followed by a 3×4 or 4×4 convolution with a stride of (1,2) or (2,2), where (s t ,s f ) represents the downsampling factors along the time and frequency axes. The downsampling alternates between strides of (1,2) and (2,2), with a total of 6 residual blocks. The number of channels gradually increases with the network depth. At the output of the final residual block, the activation has a shape of T / (H·2 3 )×F / 2 6 , where T is the number of samples in the time domain and F = W / 2 is the number of frequency bins. The last layer uses a fully connected layer (implemented as a 1×F / 2 6 convolution) to aggregate the frequency bins to obtain a one-dimensional signal in the time domain.
[0080] Let G(x) denote the decoded speech signal. The adversarial loss function of the discriminator is
[0081]
[0082] The loss function of the generator is
[0083]
[0084] The weight values λ adv , λ feat , λ rec are set to 1, 100, and 1 respectively.
[0085]
[0086] where, let k ∈ {0,..., k} index the individual discriminators at different resolutions, where k = 0 corresponds to the STFT-based discriminator, and k ∈ {1,..., k} represents the multi-resolution convolutional discriminators (in this embodiment, k = 3), and T k represents the output of the k-th discriminator along the time dimension.
[0087]
[0088] where, L is the number of internal layers, is the output of discriminator k at the l-th layer at time t, and T k,l represents the time dimension of this layer.
[0089]
[0090] Among them, S t s S(x) represents the mel spectrogram with a window length of s, a hop length of s / 4, and 64 mel filter coefficients at the t-th frame.
[0091] S4: In the deployment stage, load the best parameter model obtained in step S3, and use the scalable bone conduction speech codec network based on the diffusion probability model constructed in S2.
[0092] In the specific implementation process, the speech encoder and the speech decoder need to be used separately. Input the bone conduction speech signal in the ESMB test set into the bone conduction speech encoder loaded with the best parameters to obtain the binary bitstream to be sent;
[0093] S5: Transmit the binary bitstream from the sending end to the receiving end through the data communication module;
[0094] S6: Use the speech decoder to decode and synthesize the air conduction speech from the binary bitstream;
[0095] S7: Play the restored continuous speech signal through the speaker via the speech playback module.
[0096] For the convenience of those skilled in the art to implement specifically, the hardware platform used in the present invention is NVIDIA GeForce RTX 4090 GPU, and the software platform is the PyTorch deep learning framework. During the training process, the learning rate is 0.0002, and the adaptive learning rate optimization algorithm with weight decay is used to optimize the network parameters, where the decay coefficient of the first-order momentum estimate is 0.9, the decay coefficient of the second-order momentum estimate is 0.9, and the numerical stability term is 10 -6 .
[0097] To verify the effectiveness of the present invention, a large number of bone conduction speeches are tested under three bitrate conditions (low bitrate 3 kbps, medium bitrate 6 kbps, high bitrate 12 kbps) using objective evaluation metrics. In addition, a common bone conduction speech coding method is selected, that is, first perform speech enhancement on the bone conduction speech, and then perform encoding and decoding operations. A popular traditional speech codec and neural speech codec on the current market are selected for comparison. Among them, Scalable Neural BCM Codec is the method of the present invention, Figures 9 to 11 in which are the objective test results,Figure 12 These are the results of subjective experiments. Both sets of experimental results show that the proposed method achieves good performance in bone conduction speech coding and decoding at different rates. The method of the present invention can achieve speech enhancement while performing speech coding, reducing the complexity of use. It avoids the problem that traditional speech codecs cannot perform speech coding and decoding at low bit rates. In addition, the idea of scalability is introduced, and data packets can be dynamically adjusted according to changes in network load during transmission to meet the transmission requirements of the network. Even if some data packets are lost, high-quality speech signals can still be decoded, thus realizing a scalable bone conduction speech signal transmission method.
[0098] A scalable bone conduction speech signal transmission method based on a diffusion probability model disclosed in the present invention can dynamically adjust data packets according to changes in network load during transmission. The transmitted data packets are divided into five levels: code1, code2, code3, code4, and code5, where code1 is the base layer and code2, code3, code4, and code5 are the enhancement layers. When the network load is low, this method can transmit data packets completely. On the contrary, when the network load is high, on the premise of ensuring that the base layer is not lost, data packets can be discarded sequentially starting from code5 to meet the transmission requirements of the network. Even if some data packets are lost, high-quality speech signals can still be decoded, thus realizing a scalable bone conduction speech signal transmission method.
[0099] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A scalable bone conduction speech signal transmission method based on a diffusion probability model, characterized in that include: Collecting the bone conduction voice signal to be sent and the air conduction voice signal corresponding to the bone conduction voice, and establishing a training set and a test set based on the bone conduction voice signal and the air conduction voice signal; Constructing a bone conduction speech codec network based on a diffusion probability model, wherein the bone conduction speech codec network includes a speech encoder, a residual vector quantizer and a speech decoder, wherein the speech encoder includes a generation model and a diffusion probability model, wherein the generation model includes a feature extraction module, a high frequency reconstruction module and an overall optimization module; Training the bone conduction speech codec network on the divided training set and updating the network parameters, using the back propagation algorithm to transfer the gradient value of the network, repeatedly training and updating the network parameters, controlling the generated speech signal to approach the air conduction speech signal, thereby obtaining the optimal model parameters, and completing the training process of the bone conduction speech codec; Entering the deployment phase, the speech encoder and speech decoder are controlled separately, and the bone conduction speech signal in the test set is input into the bone conduction speech codec network loaded with the optimal model parameters to obtain a binary code stream to be sent; Transmitting the binary code stream to the receiving end through the transmitting end, and using a speech decoder to decode the binary code stream to synthesize air-conducted speech; Play continuous air-conducted speech through a loudspeaker.
2. The method for transmitting an expandable bone conduction voice signal based on a diffusion probability model according to claim 1, characterized in that: The generative model generates a roughly enhanced speech signal from a bone conduction speech signal with a sampling rate of 16kHz, and inputs the speech into a diffusion probability model for enhancement processing to generate an air conduction speech signal with a sampling rate of 48kHz, wherein the audio signal with a sampling rate of 48kHz is divided into five sub-bands by a sub-band filter, namely 0-4kHz, 4-8kHz, 8-12kHz, 12-16kHz and 16-20kHz, wherein each sub-band is subjected to an improved discrete cosine transform as the input speech feature of the residual vector quantizer.
3. The scalable bone conduction voice signal transmission method based on the diffusion probability model according to claim 2, characterized in that: The residual vector quantizer includes five residual vector quantizer branches with completely the same structure and different model parameters. The speech features generated by the speech encoder are respectively input into the five residual vector quantizer branches for quantization processing to obtain five quantization results.
4. The scalable bone conduction voice signal transmission method based on the diffusion probability model according to claim 1, characterized in that: The feature extraction module includes two types of convolutions: 3×3 convolution + ReLU and 1×1 convolution + ReLU, where convolution + ReLU means convolution followed by ReLU activation function. Specifically, 3×3 convolution + ReLU is used when the number of layers is odd, and 1×1 convolution + ReLU is used when the number of layers is even. Residual learning is used to integrate hierarchical information and retain shallow information. The process is expressed as THE kf =R((Conv1(Conv2(O full )))+Conv3(Conv4(The key ))) (1) Among them, R represents the ReLU function, Conv i represents the convolution operation, O full and O key Represent the global and local key features respectively, O kf Represents the output information of the feature extraction module.
5. The scalable bone conduction voice signal transmission method based on a diffusion probability model according to claim 1, characterized in that: The high-frequency reconstruction module is based on a generator model of a U-Net structure, adopts symmetrical encoder blocks and decoder blocks, and has a jump connection structure, wherein the layout of the encoder is symmetrical to that of the decoder, a jump connection is added between each encoder block and its symmetrical decoder block, the outermost jump connection only connects the necessary speech channels of the output, the encoder block and the decoder block each contain four blocks, and are respectively sandwiched between two conventional convolutional layers, the encoder block includes four downsampling layers, wherein the downsampling multiples are 2, 2, 8 and 8 respectively, and the decoder block performs upsampling in the opposite order; The number of channels doubles with each downsampling and halves with each upsampling. Each decoder block consists of an upsampling layer and three residual units, each of which contains one-dimensional convolutions with dilation rates of 1, 3, and 9, respectively. The encoder block corresponds to the decoder block, consists of the same residual units, and downsampling is achieved through one-dimensional convolution.
6. The scalable bone conduction voice signal transmission method based on a diffusion probability model according to claim 1, characterized in that: The overall optimization module in the generative model contains 1 layer of ReLU, 4 layers of Conv+ReLU and 1 layer of Conv; each Conv+ReLU layer consists of a convolution layer and a ReLU activation function, where the filter size of the convolution layer is 3×3, the number of input and output channels are both 64, and the final convolution layer filter size is 3×3, the number of input and output channels are 64 and 1 respectively. The process is expressed as O rb =Cv1(R(Cv2(R(Cv3(R(Cv4(R(Cv5(O hf ))))))))) (2) Among them, R represents the ReLU function, Cv i represents the convolution operation, O hf represents the output of the high-frequency reconstruction module, O rb Represents the output of the feature extraction module.
Citation Information
Patent Citations
Bone conduction voice signal transmission method based on deep learning architecture
CN118692473A
Systems and methods for audio signal generation
WO2022236803A1
Speech signal processing method, neural network training method and device
WO2024050802A1