Extremely low rate voice communication method and related device

By extracting semantic and acoustic features from speech information, and using variational posterior distribution entropy value and dynamic modulation for quantization indexing, the high intelligibility and sound quality fidelity problems of voice communication at extremely low rates are solved, and an efficient and accurate voice communication method is realized.

CN120564733APending Publication Date: 2025-08-29BEIJING UNIV OF POSTS & TELECOMM +1

Patent Information

Application Number
CN202510555243.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The prior art is difficult to take into account the high intelligibility and sound quality fidelity of speech at extremely low rates. The sounds reconstructed by traditional methods or deep learning solutions at hundreds of bit rates or even lower lack natural tone and rhythm.

Method used

Semantic and acoustic features are extracted from the speech information, and the entropy value calculation and dynamic modulation of the variational posterior distribution are calculated and dynamically modulated, and the details of the speech information are gradually restored, and the reconstruction is carried out in combination with conditional distribution sampling.

Benefits of technology

High natural and high fidelity voice communication is achieved at very low rates, and the bit requirements are significantly reduced through the method of feature separation and quantization indexing, improving channel robustness and decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564733A_ABST
    Figure CN120564733A_ABST
Patent Text Reader

Abstract

The invention provides an extremely low rate voice communication method and related equipment. The method comprises the following steps: extracting a first feature and a second feature from acquired voice information; the first feature reflects semantic information of the voice information; the second feature reflects the acoustic feature of the voice information; calculating an entropy value of variational posteriori distribution of the first feature, and adjusting the first feature according to the entropy value to obtain a third feature; performing semantic quantization on the third feature to obtain a first index; performing acoustic quantization on the second feature to obtain a second index; and gradually restoring details of the voice information based on the first index and the second index so as to complete transmission of the voice information. According to the embodiment of the invention, through semantic and acoustic feature fusion and entropy modulation and quantization, the compression efficiency is improved, the bit demand is reduced, and high-quality voice transmission and restoration are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to an extremely low-rate voice communication method and related equipment. Background Art

[0002] In traditional speech codec technology, common methods are primarily based on direct compression of time-domain waveforms or frequency-domain features. With the continuous advancement of communication technology, speech coding schemes for bandwidths of several kilobits per second (kbps) or higher have become relatively mature, capable of achieving relatively stable compressed transmission while ensuring good intelligibility and sound quality. However, when systems face more stringent bandwidth constraints (e.g., transmission rates of only a few hundred bits per second or even lower), existing traditional and deep learning-based speech codec solutions still struggle to achieve both high speech intelligibility and sound fidelity.

[0003] However, the concept of semantic communication in existing technologies aims to extract high-level semantic information from speech and restore or reconstruct the lost acoustic details through more powerful generative models at the decoding end, thereby maintaining the core content of the language at extremely low bit rates. However, if only semantic information is relied upon, the reconstructed sound often lacks natural timbre and rhythm, and sounds rather stiff. To this end, some studies have attempted to simultaneously obtain semantic and acoustic features at the encoding end, and to compensate for the loss of sound quality through a streamlined acoustic bitstream. However, in practical applications, without a suitable entropy modulation strategy and a high-performance back-end generative model, it is still difficult to achieve the goal of high naturalness and high fidelity at bit rates of hundreds of bits or even lower. Summary of the Invention

[0004] In view of this, the purpose of this application is to propose an extremely low rate voice communication method and related equipment.

[0005] Based on the above objectives, the present application provides an extremely low rate voice communication method, comprising:

[0006] Extracting a first feature and a second feature from the acquired speech information; wherein the first feature reflects semantic information of the speech information; and the second feature reflects acoustic features of the speech information;

[0007] Calculating an entropy value of a variational posterior distribution of the first feature, and adjusting the first feature according to the entropy value to obtain a third feature;

[0008] Performing semantic quantization on the third feature to obtain a first index; performing acoustic quantization on the second feature to obtain a second index;

[0009] Based on the first index and the second index, the details of the voice information are gradually restored to complete the transmission of the voice information.

[0010] In a possible implementation, extracting the first feature and the second feature from the acquired voice information includes:

[0011] Performing a temporal context analysis on the voice information to obtain the first feature;

[0012] The voice information is convolved, and the convolution weight is dynamically adjusted in combination with modulated convolution to obtain the second feature; the modulated convolution can enhance the extraction of acoustic details.

[0013] In one possible implementation, calculating an entropy value of a variational posterior distribution of the first feature and adjusting the first feature according to the entropy value to obtain a third feature includes:

[0014] Calculating the mean and variance of the first feature; constructing a variational posterior distribution based on the mean and variance;

[0015] Calculating the entropy of the variational posterior distribution;

[0016] Calculate the modulation coefficient according to the entropy value;

[0017] The first feature is adjusted using the modulation coefficient to obtain the third feature.

[0018] In a possible implementation, performing semantic quantization on the third feature to obtain the first index includes:

[0019] Calculate the distance between the third feature and each codeword in the acquired codebook, and use the index of the codeword corresponding to the shortest distance as the first index.

[0020] In a possible implementation, acoustically quantizing the second feature to obtain a second index includes:

[0021] The second feature is quantized to obtain the second index.

[0022] In a possible implementation, gradually restoring details of the voice information based on the first index and the second index to complete the transmission of the voice information includes:

[0023] Concatenate the first index and the second index to obtain a third index;

[0024] Restoring the first index in the third index to obtain a fourth feature; restoring the second index in the third index to obtain a fifth feature;

[0025] The fourth feature and the fifth feature are combined to obtain a sixth feature;

[0026] Based on the sixth feature, reverse reasoning is performed to gradually reduce noise and sample speech features from the conditional distribution to complete the transmission of the speech information.

[0027] Based on the same inventive concept, an embodiment of the present application further provides an extremely low rate voice communication device, comprising:

[0028] an extraction module configured to extract a first feature and a second feature from the acquired voice information; the first feature reflects the semantic information of the voice information; the second feature reflects the acoustic feature of the voice information;

[0029] an entropy calculation module, configured to calculate an entropy value of a variational posterior distribution of the first feature, and adjust the first feature according to the entropy value to obtain a third feature;

[0030] An index acquisition module is configured to perform semantic quantization on the third feature to obtain a first index; and perform acoustic quantization on the second feature to obtain a second index;

[0031] The restoration module is configured to gradually restore the details of the voice information based on the first index and the second index to complete the transmission of the voice information.

[0032] Based on the same inventive concept, an embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, an extremely low-rate voice communication method as described in any one of the above items is implemented.

[0033] Based on the same inventive concept, an embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute any of the above-mentioned extremely low-rate voice communication methods.

[0034] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, which includes computer program instructions, and the computer instructions are used to enable the computer program product to execute any of the above-mentioned extremely low-rate voice communication methods.

[0035] As can be seen from the above, the ultra-low rate voice communication method and related equipment provided by the present application extract a first feature and a second feature from the acquired voice information; the first feature reflects the semantic information of the voice information; the second feature reflects the acoustic feature of the voice information; the entropy value of the variational posterior distribution of the first feature is calculated, and the first feature is adjusted according to the entropy value to obtain a third feature; the third feature is semantically quantized to obtain a first index; the second feature is acoustically quantized to obtain a second index; based on the first index and the second index, the details of the voice information are gradually restored to complete the transmission of the voice information. The embodiment of the present application first extracts semantic features and acoustic features from the voice. The semantic features capture semantic information through temporal context analysis, and the acoustic features are combined with dynamic modulated convolution to enhance detail extraction. The separated features are more robust and adaptable. For the semantic features, the entropy value of the variational posterior distribution is calculated, and the semantic representation is dynamically modulated to reduce information redundancy and enhance quantization performance. The quantization step maps high-dimensional semantic and acoustic features to discrete indices, significantly reducing bit requirements while retaining content consistency and timbre characteristics. The concatenated index structure simplifies transmission, avoids desynchronization issues, and improves channel robustness. During the decoding phase, semantic and acoustic features are gradually restored through indexing, and details are gradually reconstructed through conditional distribution sampling, effectively improving speech naturalness and quality. This method achieves an optimal balance between compression efficiency, information preservation, and sound quality restoration, providing an efficient and accurate solution for extremely low-rate voice communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0037] Figure 1 This is a flowchart of an extremely low-rate voice communication method according to an embodiment of the present application;

[0038] Figure 2 This is a schematic diagram of the overall process of an embodiment of the present application;

[0039] Figure 3 This is a schematic diagram of the encoding process of an embodiment of the present application;

[0040] Figure 4 This is a schematic diagram of the decoding process of an embodiment of the present application;

[0041] Figure 5 This is a schematic diagram of the structure of an extremely low-rate voice communication device according to an embodiment of the present application;

[0042] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0044] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0045] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0046] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.

[0047] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0048] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0049] As described in the background technology section, in traditional speech coding and decoding (speech codec) technology, common methods are mainly based on direct compression of time domain waveforms or frequency domain features. With the continuous development of communication technology, speech coding schemes for bandwidth environments of several kilobits per second (kbps) or higher have become relatively mature, capable of achieving relatively stable compressed transmission while ensuring good intelligibility and sound quality. However, when the system faces more stringent bandwidth constraints (for example, transmission rates of only a few hundred bits per second or even lower), existing traditional and deep learning-based speech coding and decoding schemes still find it difficult to simultaneously take into account high speech intelligibility and sound quality fidelity.

[0050] However, the concept of semantic communication in existing technologies aims to extract high-level semantic information from speech and restore or reconstruct the lost acoustic details through more powerful generative models at the decoding end, thereby maintaining the core content of the language at extremely low bit rates. However, if only semantic information is relied upon, the reconstructed sound often lacks natural timbre and rhythm, and sounds rather stiff. To this end, some studies have attempted to simultaneously obtain semantic and acoustic features at the encoding end, and to compensate for the loss of sound quality through a streamlined acoustic bitstream. However, in practical applications, without a suitable entropy modulation strategy and a high-performance back-end generative model, it is still difficult to achieve the goal of high naturalness and high fidelity at bit rates of hundreds of bits or even lower.

[0051] Taking the above into consideration, an embodiment of the present application proposes an extremely low-rate voice communication method, which extracts a first feature and a second feature from the acquired voice information; the first feature reflects the semantic information of the voice information; the second feature reflects the acoustic feature of the voice information; the entropy value of the variational posterior distribution of the first feature is calculated, and the first feature is adjusted according to the entropy value to obtain a third feature; the third feature is semantically quantized to obtain a first index; the second feature is acoustically quantized to obtain a second index; based on the first index and the second index, the details of the voice information are gradually restored to complete the transmission of the voice information. The embodiment of the present application first extracts semantic features and acoustic features from the voice. The semantic features capture semantic information through temporal context analysis, and the acoustic features are combined with dynamic modulated convolution to enhance detail extraction. The separated features are more robust and adaptable. For the semantic features, the entropy value of the variational posterior distribution is calculated to dynamically modulate the semantic representation, reduce information redundancy and enhance quantization performance. The quantization step maps high-dimensional semantic and acoustic features to discrete indices, significantly reducing bit requirements while retaining content consistency and timbre characteristics. The concatenated index structure simplifies transmission, avoids desynchronization issues, and improves channel robustness. During the decoding phase, semantic and acoustic features are gradually restored through indexing, and details are gradually reconstructed through conditional distribution sampling, effectively improving speech naturalness and quality. This method achieves an optimal balance between compression efficiency, information preservation, and sound quality restoration, providing an efficient and accurate solution for extremely low-rate voice communications.

[0052] The technical solutions of the embodiments of the present application are described in detail below through specific examples.

[0053] refer to Figure 1 The extremely low rate voice communication method of the embodiment of the present application includes the following steps:

[0054] Step S101: extracting a first feature and a second feature from the acquired voice information; the first feature reflects the semantic information of the voice information; the second feature reflects the acoustic feature of the voice information;

[0055] Step S102, calculating the entropy of the variational posterior distribution of the first feature, and adjusting the first feature according to the entropy to obtain a third feature;

[0056] Step S103: performing semantic quantization on the third feature to obtain a first index; performing acoustic quantization on the second feature to obtain a second index;

[0057] Step S104: based on the first index and the second index, gradually restore the details of the voice information to complete the transmission of the voice information.

[0058] Regarding the above steps, in an embodiment of the present application, steps S101 to S103 may be performed in an encoder, and step S104 may be performed in a decoder.

[0059] refer to Figure 2 , which is a schematic diagram of the overall process of an embodiment of the present application.

[0060] like Figure 2 As shown, after obtaining the original speech (speech information), silence detection is first performed, and then the original speech is feature extracted to obtain speech features, and then the speech features are input into the encoder for encoding processing. In the embodiment of the present application, the encoder is divided into a semantic encoder and an acoustic encoder. In the semantic encoder, the corresponding speech information is subjected to variational entropy modulation processing, and then the modulated features are subjected to semantic quantization processing. The quantized results are semantic codes and semantic indexes. For the acoustic encoder, in the acoustic encoder, the features are acoustically quantized to directly obtain the corresponding acoustic indexes, and then the semantic indexes and acoustic indexes are spliced. After splicing, they are transmitted to the corresponding decoder through the channel. In the decoder, the semantic index and acoustic index can be directly obtained based on the spliced ​​index, and the semantic index is restored by looking up the table to obtain the corresponding semantic features. The acoustic index is subjected to μ-law inverse quantization (u-law inverse quantization) processing to restore the corresponding acoustic features. After that, the semantic features and acoustic features are subjected to self-attention feature fusion processing, and then input into the conditional diffusion decoder to achieve the recovery of the speech information.

[0061] The following is a detailed description of each step in the technical solution of this application.

[0062] refer to Figure 3 , which is a schematic diagram of the encoding end process of an embodiment of the present application.

[0063] It should be noted that the overall very low rate voice communication method of this application can be implemented in a trained communication model, which may include an encoder and a decoder. The following detailed description of the embodiments of this application is based on the training of this communication model. That is, all the following features are assumed to be used for training and will not be repeated here.

[0064] For step S101, voice information needs to be acquired first.

[0065] In this embodiment, a large amount of speech data sets can be obtained and expanded and processed through various data enhancement methods to construct a high-quality training data set to provide support for subsequent model training.

[0066] Specifically, the embodiment of the present application uses an open source speech dataset, and then adopts a common data enhancement method in the field of deep learning. Different noises, different signal-to-noise ratios, and different reverberation impulse responses are added to the speech to generate noisy speech, and finally a speech-noisy speech training dataset is obtained.

[0067] The data preprocessing process is as follows:

[0068] The 16kHz mono audio is divided into frames with a 20ms frame length and a 20ms frame shift. The 40-dimensional logarithmic Mel spectrum is calculated for each frame. The Mel frequency is closer to the human hearing mechanism than the normal frequency mechanism. It grows rapidly in the low frequency range, but grows slowly in the high frequency range. Its corresponding relationship with the normal frequency is as follows:

[0069]

[0070] Where f is the normal frequency and m is the corresponding Mel frequency. The energy spectrum of each frame is passed through a Mel-scale triangular filter bank to obtain the Mel filter bank features. Triangular bandpass filters are used to smooth the spectrum, eliminate harmonics, and highlight speech formants. This also enables feature extraction from the original spectrum to reduce computational complexity. The resulting time-frequency features are transformed using one-dimensional convolution, mapped to a 128-channel (T×128) temporal representation, and passed to the semantic and acoustic branches in parallel.

[0071] Furthermore, feature extraction is performed on the speech information to obtain a first feature and a second feature, wherein the first feature reflects the semantic information of the speech information; and the second feature reflects the acoustic feature of the speech information.

[0072] In some embodiments, the extracting of the first feature and the second feature from the acquired voice information includes: performing temporal context analysis on the voice information to obtain the first feature; performing convolution processing on the voice information and dynamically adjusting the convolution weight in combination with modulation convolution to obtain the second feature; the modulation convolution can enhance the extraction of acoustic details.

[0073] In this embodiment, a parallel coding module based on semantic and acoustic features is constructed. Deep semantic features are extracted through a semantic encoder, while detailed features of speech are obtained through an acoustic encoder. For the extracted semantic features, a variational entropy modulation (VEM) module is introduced to estimate the distribution of semantic vectors and adaptively adjust the bit allocation during the encoding process to minimize redundant information. Through semantic quantization (Semantic Vector Quantization, Semantic VQ), the continuous semantic representation is discretized into a semantic code (c_s) and the corresponding index (index) is obtained. For the extracted acoustic features, acoustic quantization (Acoustic VQ) is used to quantize them into an acoustic code (c_a). Finally, the index (index) and the acoustic code (c_a) are spliced ​​and transmitted together.

[0074] Specifically, for the semantic branch, the preprocessed features are fed into the Transformer semantic encoder, and a self-attention mechanism is applied to the (T×128) input to model the global context and output a high-level semantic vector sequence.

[0075] The self-attention mechanism dynamically captures global context by capturing the relationship between elements in the input sequence. First, the input sequence X (of dimension Dx×N, where Dx is the feature dimension and N is the sequence length) undergoes a linear transformation to generate the query (Q), key (K), and value (V) matrices:

[0076] Q=XW Q

[0077] K=XW K

[0078] V=XW V

[0079] Among them, W Q Represents the query weight matrix, W K represents the key weight matrix, W V Represents the value weight matrix.

[0080] Next, the attention score is obtained by calculating the dot product of the query and the key and scaling it:

[0081]

[0082] where d k is the dimension of the key.

[0083] The attention scores are converted into attention weights through a soft maximization function:

[0084]

[0085] Finally, the output is the weighted sum of the attention weight and the value:

[0086] Output=Attention Weights·V

[0087] The self-attention mechanism effectively captures the global context by modeling the relationship between each element in the sequence and other elements, thereby addressing long-range dependencies and improving feature extraction accuracy. The final output here is the corresponding first feature.

[0088] The temporal dimension is downsampled during the encoding process to reduce the number of frames for subsequent quantization.

[0089] At the same time, for the acoustic branch, the acoustic encoder uses a lightweight convolutional network (CNN) combined with modulated convolution to supplement the encoding of acoustic details such as timbre, rhythm, and fundamental frequency.

[0090] Further, with respect to step S102, in some embodiments, the calculation of the entropy value of the variational posterior distribution of the first feature and the adjustment of the first feature according to the entropy value to obtain the third feature include: calculating the mean and variance of the first feature; constructing a variational posterior distribution based on the mean and variance; calculating the entropy value of the variational posterior distribution; calculating the modulation coefficient according to the entropy value; and adjusting the first feature using the modulation coefficient to obtain the third feature.

[0091] In this embodiment, variational entropy modulation (VEM) is then performed on the third feature to optimize its statistical distribution, making it more compact and reducing unnecessary information redundancy. The first step in this process is to extract the feature's probability distribution information through a neural network. The mean and variance of the feature are calculated to form a variational posterior distribution. Next, the entropy of this distribution is calculated, which measures the amount of information contained in the feature. A large entropy value indicates that the feature is relatively dispersed and may contain redundant information, which is not conducive to subsequent quantization. A small entropy value indicates that the feature is overly concentrated, which may lose key information. To optimize the feature distribution, the system calculates a modulation coefficient, which is dynamically adjusted based on the feature's entropy value to ensure that the feature is neither too dispersed nor too concentrated. Subsequently, the feature is variationally sampled according to the optimized mean and variance to obtain an adjusted feature vector. This optimized feature is more suitable for subsequent vector quantization, making the feature more compact during encoding, improving quantization efficiency, and reducing subsequent quantization error.

[0092] The following is the specific calculation formula:

[0093] Variational posterior distribution:

[0094] q(z)=N(μ,σ 2 )

[0095] Where q(z) represents the variational posterior distribution;

[0096] represents the mean learned by the neural network, represents the center of the feature, Represents the corresponding mean calculation function, z e Indicates the first feature.

[0097] σ s =g φ (z e ) represents the variance learned by the neural network, represents the uncertainty of the feature, g φ Represents the corresponding variance calculation function.

[0098] Information entropy:

[0099]

[0100] Among them, H(q(z e )) represents information entropy.

[0101] Reconstruction features:

[0102] λ(H)=α+β*tanch(γ(H-H0))

[0103]

[0104] Among them, λ(H) represents the modulation coefficient, H represents the information entropy, H0 represents the target entropy value (hyperparameter), z′ e Represents the third feature, represents the standard normal distribution sampling noise, α, β, and γ represent different parameters that control the modulation intensity.

[0105] Further, with respect to step S103, the semantically quantizing the third feature to obtain the first index includes: calculating the distance from the third feature to each codeword in the acquired codebook, and taking the index of the codeword corresponding to the shortest distance as the first index.

[0106] After completing the variational entropy modulation, the optimized feature vector will enter the vector quantization (VQ) process to further reduce the data dimension and convert it into a discrete representation for easy storage or transmission.

[0107] In this embodiment, during the VQ process, the system first presets a vector codebook containing multiple codewords representing different categories in the feature space. For each optimized feature vector, the system calculates its distance to all codewords in the codebook, finds the closest codeword, and uses the index of that codeword as the final discretized representation. Ultimately, the combination of VEM and VQ enables semantic features to maintain core information at low bit rates while significantly reducing storage and transmission overhead through efficient quantization, thereby achieving ultra-low bitrate speech coding or semantic information compression.

[0108] Specifically, calculate the distance from the third feature after variational entropy modulation to the codebook, find the nearest codeword index k, and replace the input feature with the nearest codeword. The distance calculation formula is:

[0109]

[0110] d(z′ e ,c k ) is the third feature z′ e To the kth codeword c k The Euclidean distance, K is the codebook size

[0111] Semantic quantization obtains the semantic code (c_s) and records the index k for transmission.

[0112] Regarding the second feature, in some embodiments, performing acoustic quantization on the second feature to obtain the second index includes: performing quantization processing on the second feature to obtain the second index.

[0113] In this embodiment, acoustic features are quantized into acoustic codes (c_a) (i.e., the second index) using u-law quantization through acoustic quantization (Acoustic VQ). The quantization formula is:

[0114]

[0115] x is the input acoustic feature vector, κ is the non-uniform parameter, typically set to 255 (8-bit quantization), and c_a is the quantized integer index. Acoustic quantization effectively reduces bit requirements while ensuring naturalness and timbre. The concatenation of the first index generated by semantic quantization and the second index generated by acoustic quantization is then transmitted to the decoder.

[0116] For further reference, Figure 4 , which is a schematic diagram of the decoding end process of an embodiment of the present application.

[0117] For step S104, in some embodiments, based on the first index and the second index, the details of the voice information are gradually restored to complete the transmission of the voice information, including: splicing the first index and the second index to obtain a third index; restoring the first index in the third index to obtain a fourth feature; restoring the second index in the third index to obtain a fifth feature; splicing the fourth feature and the fifth feature to obtain a sixth feature; based on the sixth feature, performing reverse reasoning, gradually reducing noise, and sampling voice features from the conditional distribution to complete the transmission of the voice information.

[0118] In this embodiment, the conditional diffusion model is used to gradually restore the audio details, the fourth feature is restored by table lookup according to the first index, and the second index is restored to the fifth feature through μ-law inverse transform. The fourth feature and the fifth feature are fused using the attention mechanism. Then, the conditional diffusion model (Diffusion Decoder) is used to gradually restore the details of the audio. First, the fourth feature and the fifth feature are spliced ​​to form the input of the conditional diffusion network. In the diffusion decoding process, starting from random noise, after multiple steps of iterative denoising and feature repair, the acoustic texture and time-frequency details of the audio are restored layer by layer, and finally a high-fidelity time-frequency representation is output. Afterwards, a waveform decoder based on a variational autoencoder (VAE) is used to convert the above time-frequency representation into a time-domain audio signal to obtain a reconstructed speech that can be played directly.

[0119] Specifically, during the diffusion decoding process, the conditional diffusion network begins with random noise and, through multiple denoising iterations and feature refinement, gradually restores the audio's spectral details and acoustic texture. Each iteration, guided by the conditional input, allows the decoding process to gradually approach the time-frequency representation of the real audio. This gradual approximation method effectively avoids the distortion common in traditional generative models, ensuring the finesse and realism of the reconstructed audio.

[0120] During the diffusion decoding process, a denoising diffusion model is used to model the conditional probability distribution p(x|z,s), where x represents the output speech features, and z and s represent the fourth and fifth features recovered at the decoder, respectively. The denoising network, based on a U-Net architecture, combines a multi-layer perceptron (MLP) and a residual network. These features are used for conditional inference, resulting in more efficient reconstruction of speech features.

[0121] In the reverse inference phase of the diffusion process, the model samples speech features from the conditional distribution by gradually reducing the noise, ultimately generating a high-fidelity speech signal. To this end, the diffusion loss function is defined as a denoising criterion, which is expressed as:

[0122] L(z,s,x)=E ∈,t [||ε-ε θ (x t |t,z,s)|| 2 ]

[0123] Among them, ε∈R d is a noise vector sampled from the standard normal distribution N(0,I), x t is the speech feature vector after noise perturbation, defined as is the parameter of the noise schedule. The noise estimation network ε θ By optimizing the parameter θ, we receive the conditional information t, z, and s as input to predict the noise distribution. This loss function can be regarded as a scoring matching method that specifically handles the scoring function loss of the conditional distribution p(x|z,s), which is expressed as

[0124] After diffusion decoding is complete, the resulting spectral representation typically takes the form of a Mel-spectrogram. Subsequently, this embodiment uses a waveform decoder based on a variational autoencoder (VAE) to perform waveform synthesis on the Mel-spectrogram representation output by the diffusion network, converting the spectral information into a time-domain audio signal. This waveform decoder fully considers the smoothness and consistency of spectral features during the decoding process, further improving the fidelity and naturalness of the reconstructed audio.

[0125] To comprehensively evaluate the decoding effect, this implementation calculates multiple metrics between the decoded audio and the original audio, including: Perceptual Evaluation of Speech Quality (PESQ): used to objectively evaluate the quality and distortion of speech. Short-Time Objective Intelligibility (STOI): used to evaluate speech intelligibility. Mean Opinion Score (MOS): based on the average score of subjective listening experience, measuring the naturalness and listening experience of the reconstructed audio.

[0126] A comprehensive analysis of the above indicators shows that the decoding scheme based on the conditional diffusion model and VAE waveform decoder can effectively restore the acoustic details and semantic information of the audio, ensuring that the reconstructed audio reaches a high level in clarity, naturalness and fidelity.

[0127] It can be seen from the above embodiments that the extremely low-rate voice communication method described in the embodiments of the present application extracts a first feature and a second feature from the acquired voice information; the first feature reflects the semantic information of the voice information; the second feature reflects the acoustic feature of the voice information; the entropy value of the variational posterior distribution of the first feature is calculated, and the first feature is adjusted according to the entropy value to obtain a third feature; the third feature is semantically quantized to obtain a first index; the second feature is acoustically quantized to obtain a second index; based on the first index and the second index, the details of the voice information are gradually restored to complete the transmission of the voice information. The present application proposes an efficient extremely low-rate voice communication method, which achieves the goal of compressing the amount of information while maintaining the naturalness of the voice and the accuracy of restoration by extracting, quantizing, transmitting and reconstructing the semantic features and acoustic features of the voice information respectively. This method integrates the modulation of the variational posterior distribution, adaptive quantization and a carefully designed decoding method. The technical effects of each part are analyzed one by one below.

[0128] First, the present application extracts the first feature and the second feature, namely the semantic feature and the acoustic feature, from the speech information. The first feature (semantic feature) can effectively capture the semantic content contained in the speech information, that is, the abstract language expression part, while the second feature (acoustic feature) is responsible for describing the acoustic details of the speech, such as pitch, timbre, duration, etc. Through this feature separation framework, the method realizes semantic-level information compression at the analysis level, decouples the semantic representation of natural language from the acoustic information, thereby reducing the coupling between redundant features and improving the efficiency and robustness of subsequent processing. When extracting the first feature, the semantic relationship is captured through temporal context analysis, which greatly improves the quality of the semantic representation used for compression and quantization. The convolution processing and dynamic modulation of the second feature not only improves the accuracy of the feature representation, but also strengthens the coverage of acoustic details. Its dynamic modulation mechanism effectively copes with the complex and diverse acoustic changes in speech, and improves the adaptability of the extracted features to downstream tasks.

[0129] In the technical solution of the second step, by calculating the entropy of the variational posterior distribution of the first feature and adjusting it, the traditional semantic model optimization is introduced into the framework of information theory to obtain the third feature. The technical effect of this step is mainly reflected in two aspects: on the one hand, the calculation of the mean and variance and the variational posterior distribution constructed based on this can fully capture the essential statistical characteristics of the first feature, providing a better statistical basis for subsequent quantization; on the other hand, through the calculation and modulation of the entropy of the posterior distribution, the semantic features are dynamically adjusted, further reducing information redundancy. The modulation process balances the compactness of feature representation and information retention rate by introducing entropy optimization. This modulation mechanism based on variational distribution enhances the discriminative ability of the third feature and provides a higher quality basic feature for the subsequent quantization step.

[0130] The semantic quantization step performed on the third feature calculates its distance to the codewords in the codebook and selects the codeword index corresponding to the closest distance as the first index. This process uses a pre-trained semantic codebook to achieve effective mapping, compressing continuous high-dimensional semantic feature vectors into fixed discrete index representations. This quantization operation greatly reduces the number of bits required for semantic feature transmission while maintaining the accuracy of the semantic content. At the same time, due to the use of distance-optimal codeword matching, each quantization step can find the codeword expression that is closest to the semantic feature, minimizing quantization error and improving the reconstruction quality of the semantic coding.

[0131] Acoustic feature quantization achieves similar results to semantic feature quantization by quantizing the second feature (acoustic feature) into discrete indices (i.e., second indices). Acoustic quantization uses non-uniform distributions, such as algorithms like \muμ-law, to effectively reduce bit requirements while preserving the naturalness of the sound, thereby significantly reducing the transmission overhead of the acoustic features. Furthermore, because the acoustic features extracted by convolution contain localized time-frequency information, the quantization results can preserve the timbre and pitch characteristics of the speech, further ensuring the quality of speech restoration during decoding.

[0132] During the information assembly and transmission phase, the data is further unified by concatenating the first index (semantic index) with the second index (acoustic index). This concatenated third index not only simplifies the data transmission structure but also avoids the desynchronization issues that can arise from separate transmissions. Transmitting both indexes as a whole also reduces channel loss during transmission, improving the integrity and robustness of information delivery.

[0133] Finally, during the decoding phase, each feature is gradually restored from the concatenated index (the third index), followed by detailed reconstruction and speech generation, helping to achieve high-quality transmission and reconstruction of speech information. During decoding, the first index (semantic index) is restored from the third index to extract the fourth feature (semantically restored feature), and the second index (acoustic index) is restored to extract the fifth feature (acoustically restored feature). The semantic and acoustic features are concatenated to form the sixth feature, and noise is gradually reduced during the reverse inference process, allowing conditional distribution sampling. This design replaces direct prediction with a gradually guided reverse inference process, fully utilizing the randomness of the sampling noise to ensure natural speech generation and a high signal-to-noise ratio. By combining semantic and acoustic features, the decoder can more accurately restore speech details, achieving the greatest approximation to the original speech.

[0134] Overall, the technical effect of this application is significant. Efficient compression is achieved through entropy modulation and quantization operations, minimizing the bit requirements for data transmission; and combined with feature extraction methods such as temporal context analysis and dynamic modulation convolution, the core information of semantic and acoustic features is fully captured, significantly enhancing transmission efficiency and decoding quality. Combined with the step-by-step reasoning decoding strategy, the generated restored speech performs well in naturalness, timbre and content consistency. This method has a high reference value and broad application prospects for practical applications in the field of voice communication.

[0135] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.

[0136] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0137] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an extremely low-rate voice communication device.

[0138] refer to Figure 5 , the extremely low rate voice communication device comprises:

[0139] The extraction module 51 is configured to extract a first feature and a second feature from the acquired voice information; the first feature reflects the semantic information of the voice information; the second feature reflects the acoustic feature of the voice information;

[0140] an entropy calculation module 52 configured to calculate an entropy value of a variational posterior distribution of the first feature, and adjust the first feature according to the entropy value to obtain a third feature;

[0141] The index acquisition module 53 is configured to perform semantic quantization on the third feature to obtain a first index; and perform acoustic quantization on the second feature to obtain a second index;

[0142] The restoration module 54 is configured to gradually restore the details of the voice information based on the first index and the second index to complete the transmission of the voice information.

[0143] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0144] The apparatus of the above embodiment is used to implement the corresponding extremely low rate voice communication method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0145] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the extremely low-rate voice communication method described in any of the above embodiments is implemented.

[0146] Figure 6 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0147] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0148] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0149] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0150] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0151] The bus 1050 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).

[0152] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0153] The electronic device of the above embodiment is used to implement the corresponding extremely low rate voice communication method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0154] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the extremely low-rate voice communication method described in any of the above embodiments.

[0155] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0156] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the extremely low rate voice communication method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0157] Based on the same inventive concept, corresponding to the extremely low-rate voice communication method described in any of the above embodiments, the present disclosure further provides a computer program product comprising computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the extremely low-rate voice communication method. For the execution entities corresponding to the steps in each embodiment of the extremely low-rate voice communication method, the processors that execute the corresponding steps can belong to the corresponding execution entities.

[0158] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the extremely low rate voice communication method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0159] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0160] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.

[0161] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.

[0162] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.

Claims

1. A very low rate voice communication method, characterized in that: include: Extracting a first feature and a second feature from the acquired voice information; The first feature reflects the semantic information of the speech information; the second feature reflects the acoustic feature of the speech information; Calculating an entropy value of a variational posterior distribution of the first feature, and adjusting the first feature according to the entropy value to obtain a third feature; Performing semantic quantization on the third feature to obtain a first index; performing acoustic quantization on the second feature to obtain a second index; Based on the first index and the second index, the details of the voice information are gradually restored to complete the transmission of the voice information.

2. The method according to claim 1, characterized in that The extracting the first feature and the second feature from the acquired voice information includes: Performing a temporal context analysis on the voice information to obtain the first feature; The voice information is convolved, and the convolution weight is dynamically adjusted in combination with modulated convolution to obtain the second feature; the modulated convolution can enhance the extraction of acoustic details.

3. The method according to claim 1, characterized in that The calculating an entropy value of a variational posterior distribution of the first feature and adjusting the first feature according to the entropy value to obtain a third feature includes: Calculating the mean and variance of the first feature; constructing a variational posterior distribution based on the mean and variance; Calculating the entropy of the variational posterior distribution; Calculate the modulation coefficient according to the entropy value; The first feature is adjusted using the modulation coefficient to obtain the third feature.

4. The method according to claim 1, wherein The performing semantic quantization on the third feature to obtain a first index includes: Calculate the distance between the third feature and each codeword in the acquired codebook, and use the index of the codeword corresponding to the shortest distance as the first index.

5. The method according to claim 1, characterized in that The acoustically quantizing the second feature to obtain a second index includes: The second feature is quantized to obtain the second index.

6. The method according to claim 1, characterized in that The step of gradually restoring details of the voice information based on the first index and the second index to complete the transmission of the voice information includes: Concatenate the first index and the second index to obtain a third index; Restoring the first index in the third index to obtain a fourth feature; restoring the second index in the third index to obtain a fifth feature; The fourth feature and the fifth feature are combined to obtain a sixth feature; Based on the sixth feature, reverse reasoning is performed to gradually reduce noise and sample speech features from the conditional distribution to complete the transmission of the speech information.

7. An extremely low rate voice communication device, characterized in that: include: an extraction module, configured to extract a first feature and a second feature from the acquired voice information; The first feature reflects the semantic information of the speech information; the second feature reflects the acoustic feature of the speech information; an entropy calculation module, configured to calculate an entropy value of a variational posterior distribution of the first feature, and adjust the first feature according to the entropy value to obtain a third feature; An index acquisition module is configured to perform semantic quantization on the third feature to obtain a first index; and perform acoustic quantization on the second feature to obtain a second index; The restoration module is configured to gradually restore the details of the voice information based on the first index and the second index to complete the transmission of the voice information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.

10. A computer program product comprising computer program instructions, which, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice compression method and system based on deep learning and vector prediction

    CN117423348A

Cited By

  • Voice communication method, device and equipment based on three-group decomposition

    CN121811890A