Low bit rate cross high loss medium speech communication method based on semantic feature separation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-06-17
- Publication Date
- 2026-08-04
AI Technical Summary
[0006]鉴于上述问题,本发明提供了一种基于语义特征分离的低码率跨高损媒介语音通信方法,解决了现有技术中高损媒介语音通信的语音质量和性能不够好的技术问题
(1)本发明通过“32维Mel频谱++清浊音”与“32维→18维倒谱+2维参数”的双层表示,将语音的语义可懂度、节律结构与音色细节在特征层面合理分离,为不等错误保护与低码率压缩提供基础。
Smart Images

Figure CN122511269A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication and signal processing technology under extreme communication environments, specifically to a low-bit-rate voice communication method across high-loss media based on semantic feature separation. Background Technology
[0002] In confined propagation environments characterized by complex obstructions, strong attenuation, and multi-interface reflections, electromagnetic waves and sound waves are easily affected by strong absorption, strong scattering, and time-varying channel conditions during propagation, making it difficult for conventional mid-to-high frequency wireless systems to achieve stable coverage. To improve penetration and communication capabilities in such environments, low-frequency magneto-electric mechanical antennas are commonly used in engineering practice to achieve very low frequency (VLF) and extremely low frequency (ELF) communication. However, their usable bandwidth is extremely narrow, and the information bit rate they can carry is typically only about [missing information]. This makes it difficult to meet the demand for direct transmission of high-fidelity voice.
[0003] Traditional hierarchical designs for speech coding and error correction channel coding typically use speech encoders such as CELP and MelP as the front end and error correction codes such as convolutional codes and LDPC as the back end, often based on simplified path loss and Rayleigh / Rayse fading models for optimization. However, these methods struggle to accurately characterize layered propagation and effects such as guided modes and side waves in multi-layered blocking, non-uniform media, and complex reflection / diffraction environments. Consequently, they are prone to mismatches between code rate allocation and protection strategies and real channel characteristics; under extremely low signal-to-noise ratio and strict code rate constraints, speech intelligibility will significantly decrease.
[0004] Most existing technical solutions are designed for single radio environments or statistical channel models, lacking specific designs for low-frequency voice communication across high-loss media; while existing channel modeling research has proposed parameterized descriptions based on electromagnetic simulation and experimental calibration, it has not yet been combined with low-bit-rate speech source-channel coding. Deep coupling, and not in It balances speech intelligibility, perceptual quality, bit rate stability, and engineering feasibility at bit rates of several orders of magnitude.
[0005] Under the dual constraints of extremely narrow bandwidth and extremely low signal-to-noise ratio, how to reasonably separate semantic intelligibility-related information from timbre and detail-related information while simplifying channel state representation, and combine hash compression and low-frequency modulation structure to complete voice communication across high-loss media, is a key problem that urgently needs to be solved. Summary of the Invention
[0006] In view of the above problems, the present invention provides a low bit rate voice communication method across high loss media based on semantic feature separation, which solves the technical problem that the voice quality and performance of high loss media voice communication in the prior art are not good enough.
[0007] This invention provides a low-bit-rate voice communication method across high-loss media based on semantic feature separation, characterized by the following steps: Step S1: Preprocess and transform the speech to be transmitted to obtain the Mel spectrum image and fundamental frequency voiced / unvoiced features; Step S2: Based on the real-time channel state, perform joint source channel coding on the Mel spectrum image and the fundamental frequency voiced / unvoiced features to obtain Mel continuous hidden features and fundamental frequency voiced / unvoiced continuous hidden features. Step S3: Align the Mel continuous latent features and the fundamental voiced / unvoiced continuous latent features in time and combine them into a hash input vector; input the hash input vector into the trainable hash bottleneck, and obtain a unified bitstream by K-dimensional hash space projection and binarization operation; Step S4: Perform keying modulation based on the unified bit stream and the channel state to obtain a modulated signal; transmit the modulated signal and pilot signal through the transmitting antenna; Step S5: Update the channel state according to the pilot signal received by the receiving antenna and feed it back to the transmitting end; Step S6: Demodulate, hash decode, and perform joint source channel decoding on the modulated signal received by the receiving antenna in sequence, corresponding to those at the transmitting end, to obtain the Mel spectrum image and fundamental frequency voiced / unvoiced features; Based on the Mel cepstral coefficients and fundamental voiced / unvoiced features of the Mel spectrogram image, parameterized speech features are obtained by concatenating them; the parameterized speech features are then input into... A frame neural vocoder is used to obtain the reconstructed speech waveform.
[0008] Preferably, in step S1, the preprocessing specifically includes performing pre-emphasis, frame windowing, and endpoint detection operations on the original speech signal in sequence; The feature transformation specifically includes performing frequency domain transformation on each frame of speech signal using short-time Fourier transform, calculating the power spectrum and mapping it to the Mel scale to obtain a 32-dimensional logarithmic Mel spectrum sequence, which serves as the Mel spectrum image. For each frame of speech, the fundamental frequency F0 and the voiced / unvoiced probability VUV are extracted. The voiced / unvoiced probability VUV takes the value of 0 or 1, where 1 represents unvoiced and 0 represents voiced.
[0009] Preferably, in step S2, joint source-channel coding of the Mel spectrum image is performed based on the real-time channel state, specifically including: Two-dimensional convolution is performed on the Mel spectrum image. The features obtained from the two-dimensional convolution are input into a state space network. The output features are then subjected to gated linear attention processing to obtain the Mel continuous latent features. During this process, the channel state is used... Dynamically adjust feature redundancy allocation, channel weights, and encoding strategies; In step S2, joint source-channel coding is performed on the fundamental frequency voiced / unvoiced tone features based on the real-time channel state, specifically including: The fundamental frequency and voiced / unvoiced features are input into a one-dimensional sequence modeling network to generate continuous latent features of the fundamental frequency and voiced / unvoiced sounds; during this process, coding redundancy and quantization precision are adjusted by channel state. The channel state This includes: normalized path loss factor, normalized delay spread scale, normalized noise power, and scene outline number; The channel state The prior values are obtained as follows: Transceiver antennas are deployed in a typical high-loss medium scenario to collect measured data on path loss, delay spread, and noise power. Parameters are fitted based on a hierarchical medium propagation model, and the channel state is obtained after normalization and quantization. The prior value; The expression for the joint source-channel coding is:
[0010]
[0011] in, This is a continuous latent feature of Mel. This is a continuous hidden feature of the fundamental frequency voiceless / unvoiced tone. and These represent the Mel-image joint source-channel coding network and the fundamental frequency voiced / unvoiced one-dimensional joint source-channel coding network, respectively. Indicates the fundamental frequency. Indicates the voicing or unvoicing of consonants.
[0012] Preferably, in step S3, the training process of the trainable hash bottleneck adopts a pass-through estimator; in addition, during the training process, the average bitrate of the unified bitstream is kept within a preset range through time window statistics and bitrate constraints. In step S3, after K-dimensional hash space projection and binarization, the expression for obtaining the unified bitstream is:
[0013]
[0014]
[0015] in, For the hash input vector, and For trainable parameters, For the K-dimensional continuous hash projection result, Represents a symbolic function. This represents a binary bit vector.
[0016] Preferably, step S4 specifically includes: according to The corresponding modulation parameters are used to map the unified bit stream output from the hash bottleneck using amplitude shifting, frequency shifting, or multilevel shifting to generate a modulated signal. Generate pilot signals and insert them into the modulated signal according to the frame structure; The modulation signal and pilot signal are transmitted via a low-frequency magneto-electric mechanical antenna.
[0017] Preferably, step S5 specifically includes: The receiving antenna extracts the pilot signal, compares it with the locally stored transmit pilot reference signal, and obtains the amplitude attenuation of the pilot signal; the channel state is updated based on the amplitude attenuation of the pilot signal. The expression is:
[0018] in, This is the updated channel state vector. The channel state vector before the update. To estimate the channel state based on the current pilot signals, To smoothly update the coefficients, This indicates normalization and quantization processing; The updated It is sent to the transmitter via the feedback link.
[0019] Preferably, step S6 includes: The receiver recovers the bitstream from the high-loss medium channel through demodulation, and then uses a hash decoding network and continuous features. Decoder, in the updated Under the given conditions, the Mel continuous latent features and the fundamental voiced / unvoiced continuous latent features are reconstructed; the reconstructed Mel continuous latent features are then restored to a 32-dimensional Mel spectrum image by a decoding network, and the fundamental voiced / unvoiced continuous latent features are restored to fundamental voiced / unvoiced features. The 32-dimensional Mel spectrum is then mapped to 18-dimensional Mel cepstral coefficients, which are concatenated with the fundamental voiced / unvoiced features to obtain 20-dimensional parametric speech features. The 20-dimensional parametric speech features are then input into a neural vocoder based on an LPCNet-like framework to obtain the reconstructed speech waveform.
[0020] Preferably, end-to-end training is performed on the low-bitrate voice communication process across high-loss media based on semantic feature separation, and the expression for the comprehensive loss function of the end-to-end training is as follows:
[0021] in, Represents the comprehensive loss function. , , , , as well as These represent the Mel spectrum reconstruction loss, multi-resolution spectrum loss, fundamental frequency bias loss, voiced / unvoiced consistency loss, 20-dimensional parametric feature reconstruction loss, and bit rate bias penalty term, respectively. , , , , , These are the weights of each loss term; During the end-to-end training process, different additive noise, frequency-selective fading, multipath convolution, bandwidth limitation, and path loss are applied to the transmission process of continuous latent features.
[0022] Compared with the prior art, the present invention has at least the following beneficial effects: (1) This invention utilizes "32-dimensional Mel spectrum + The dual-layer representation of "voiced / unvoiced sounds" and "32-dimensional → 18-dimensional cepstral + 2-dimensional parameters" reasonably separates the semantic intelligibility, rhythmic structure and timbre details of speech at the feature level, providing a foundation for unequal error protection and low bit rate compression.
[0023] (2) This invention implements the joint source-channel coding function in the Mel image-based joint source-channel coding unit, the "F0 / VUV" one-dimensional joint source-channel coding unit, and the decoding unit structure, through... Injection enables adaptation to channel conditions across high-loss media, and has the characteristic of gradually smoothing degradation when channel conditions deteriorate, thus avoiding a significant cliff effect in voice quality.
[0024] (3) This invention introduces a trainable hash bottleneck to achieve direct mapping from features to bitstream, and works as a bit-level compression and bitrate control module in an end-to-end training framework, achieving approximately The average bitrate is kept stable within a certain range, and bitrate fluctuations are limited within a short time window.
[0025] (4) This invention adopts the imitation The framework's lightweight neural vocoder, conditioned on 20-dimensional parametric speech features, models the vocal tract response and excitation residuals separately, achieving performance comparable to traditional... Approximate perceived quality and naturalness, while also being similar to the front end. Structural co-formation exhibits progressive degradation characteristics under extremely low bit rate and extremely low signal-to-noise ratio conditions.
[0026] (5) This invention adopts a simplified engineering-feasible scheme in terms of channel characterization and physical layer implementation. Composed of a small number of measurable parameters, the physical layer modulation adopts a low-complexity keying method, which is compatible with the hardware structure of low-frequency magneto-electric mechanical antennas and has good engineering feasibility. Attached Figure Description
[0027] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.
[0028] Figure 1 This is a schematic diagram of the overall architecture of the low bit rate voice communication method across high loss media based on semantic feature separation provided by the present invention.
[0029] Figure 2 The speech front-end feature extraction and 32-dimensional Mel spectrum provided by this invention Diagram showing the relationship between voiced and unvoiced sounds.
[0030] Figure 3 Mel latent features and voiced / unvoiced sounds provided by this invention Schematic diagram of dual-path representation and Mel-image-based joint source-channel coding unit and "F0 / VUV" one-dimensional joint source-channel coding unit structure.
[0031] Figure 4 This is a schematic diagram of the trainable hash bottleneck compression and bitstream generation process provided by the present invention.
[0032] Figure 5 This is a schematic diagram illustrating physical layer modulation and transmission on a low-frequency magnetoelectric mechanical antenna provided by the present invention.
[0033] Figure 6 The receiver hash decoding, Mel-image-based joint source-channel decoding unit, "F0 / VUV" one-dimensional joint source-channel decoding unit, and simulation provided by this invention are all part of this invention. A schematic diagram of the speech reconstruction process using a frame neural vocoder.
[0034] Figure 7 The flowchart of the low bit rate voice communication method across high loss media based on semantic feature separation provided by the present invention is shown. Detailed Implementation
[0035] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0036] In order to illustrate the effectiveness of the method proposed in this invention, a specific embodiment of the present invention will be used to describe the above technical solution of the present invention in detail below, such as... Figure 7 As shown, this invention discloses a low-bit-rate voice communication method across high-loss media based on semantic feature separation. The specific implementation steps are as follows: Step S1: Preprocess and transform the speech to be transmitted to obtain the Mel spectrum image and fundamental frequency voiced / unvoiced features; In this step, such as Figure 1 and Figure 2 As shown, the present invention first performs preprocessing and feature transformation on the original speech signal to be transmitted at the transmitting end. Specifically, the preprocessing includes sequentially performing pre-emphasis, frame windowing, and endpoint detection operations on the original speech signal.
[0037] In some embodiments, pre-emphasis can be achieved using a high-pass filter to compensate for high-frequency components of the speech signal; the framing operation can be performed by dividing the original signal into sliding windows by setting the window length and frame shift length; the window function for windowing can be a Hamming window or a Hanning window; the endpoint detection can be based on information such as signal energy to remove silence segments and mark the effective start and end positions of speech.
[0038] To further illustrate the above preprocessing and feature transformation process, in some embodiments, the pre-emphasis processing can be expressed as:
[0039] in, The original speech signal with index n. For the original speech signal with index n-1, This is the pre-emphasized speech signal. This is the pre-emphasis coefficient.
[0040] The feature transformation operation includes: performing a frequency domain transformation on each frame of speech signal based on the short-time Fourier transform (STFT), calculating the power spectrum and mapping it to the Mel scale to obtain a 32-dimensional logarithmic Mel spectrum sequence that evolves over time. The 32-dimensional logarithmic Mel spectrum is nonlinearly divided on the frequency axis according to the characteristics of human ear perception, with higher resolution in the low-frequency region and lower resolution in the high-frequency region.
[0041] Specifically, performing a short-time Fourier transform on the windowed, framed speech yields:
[0042] in, This represents the frequency domain signal with index k. For index The pre-emphasized speech signal, For window functions, For frame shift, For the length of the window, It is the imaginary unit.
[0043] Furthermore, the first Frame number The dimensional logarithmic Mel spectrum can be expressed as:
[0044] in, For the first Mel filters, Indicates amplitude. A constant to prevent logarithmic operations from resulting in zero values.
[0045] The 32-dimensional Mel spectrum that evolves over time is regarded as a two-dimensional speech image, with one dimension being the time axis and the other dimension being the frequency axis, thus forming the Mel spectrum image.
[0046] Meanwhile, the fundamental frequency F0 and voiced / unvoiced probability VUV are extracted from each frame of speech on a unified time axis. The voiced / unvoiced probability VUV takes a value of 0 or 1, where 1 represents unvoiced sound and 0 represents voiced sound.
[0047] Step S2: Based on the real-time channel state, perform joint source channel coding on the Mel spectrum image and the fundamental frequency voiced / unvoiced features to obtain Mel continuous hidden features and fundamental frequency voiced / unvoiced continuous hidden features. like Figure 3 As shown, this invention feeds the Mel spectrum image and F0 / VUV features into two independent joint source channel coding networks for processing.
[0048] For Mel spectrogram images, this invention employs a Mel image-based JSCC coding unit for processing. The JSCC coding unit performs two-dimensional convolution on the Mel spectrogram image to extract local textures and energy patterns on the time-frequency plane. The features obtained from the two-dimensional convolution are input into a state space network, which can model long-range dependencies in the time dimension and capture the rhythmic structure and dynamic evolution of speech. Finally, gated linear attention processing is performed to obtain Mel continuous latent features. Gated linear attention processing can aggregate key information over a large time span while maintaining the linear growth characteristic of computational complexity, thus characterizing both the rhythmic structure and spectral envelope structure of speech while keeping the parameter scale lightweight.
[0049] The Mel-image JSCC encoding unit of the present invention also uses real-time channel conditions during the encoding process. According to the above at each layer Adjust feature redundancy allocation, channel weights, and encoding strategies. Specifically, this can be achieved through feature scaling, channel attention mechanisms, or adaptive instance normalization. Injected into each coding layer, enabling the network to dynamically adjust feature redundancy allocation, channel weights, and coding strategies based on real-time channel conditions.
[0050] Channel state of the present invention It is a vector, and the channel state vector includes the normalized path loss factor, the normalized delay spread scale, the normalized noise power, and the scene contour number, which are used to characterize the signal attenuation level, the degree of temporal dispersion, the background noise intensity, and the scene type across high-loss media, respectively.
[0051] Accordingly, the channel state vector can be represented as:
[0052] in, Represents the channel state vector. This represents the normalized path loss factor. Indicates the normalized delay spread scale. Represents normalized noise power. Indicates the scene outline number.
[0053] In some embodiments, channel state The prior values can be obtained through actual measurements, specifically including: deploying transceiver antennas in typical high-loss medium scenarios, collecting measured data on path loss, delay spread, and noise power, fitting parameters based on a hierarchical medium propagation model, and obtaining the channel state after normalization and quantization processing. The prior value.
[0054] By employing the Mel image-based JSCC coding unit, joint source-channel coding and inequality error protection for Mel continuous latent features can be completed within the image-based JSCC coding module, providing Mel continuous latent features with a small code rate redundancy. Thus, under the condition of limited total bit rate, the intelligibility and reliable transmission of speech rhythm information are given priority.
[0055] For the fundamental frequency voiced / unvoiced tone features, this invention employs independent one-dimensional JSCC coding units for processing. Specifically, the fundamental frequency and voiced / unvoiced tone features are input into a one-dimensional sequence modeling network, and temporal modeling is performed while maintaining frame-level temporal resolution. The one-dimensional sequence modeling network has multiple neural network layers, and each layer is also based on the channel state vector. By adjusting coding redundancy and quantization precision, the continuous latent features of fundamental frequency voiced and unvoiced tones are finally generated, achieving robust joint source-channel coding of fundamental frequency trajectories and voiced / unvoiced boundaries.
[0056] The above dual-path joint source-channel coding process can be further expressed as:
[0057]
[0058] in, This is a continuous latent feature of Mel. This is a continuous hidden feature of the fundamental frequency voiceless / unvoiced tone. and These represent the Mel-image joint source-channel coding network and the fundamental frequency voiced / unvoiced one-dimensional joint source-channel coding network, respectively. Indicates the fundamental frequency. Indicates the voicing or unvoicing of consonants.
[0059] Through the aforementioned dual-path independent coding structure, semantic feature separation is achieved, with Mel spectrum focusing on semantics and timbre structure, and F0 / VUV focusing on prosody and voiced / unvoiced boundaries, providing a foundation for subsequent trainable hash compression and low bit rate transmission.
[0060] Step S3: Align the Mel continuous latent features and the fundamental voiced / unvoiced continuous latent features in time and combine them into a hash input vector; input the hash input vector into the trainable hash bottleneck, and obtain a unified bitstream by K-dimensional hash space projection and binarization operation; like Figure 1 , Figure 4 As shown, a trainable hash bottleneck is set at the output of the Mel image-based JSCC encoding module and the independent one-dimensional JSCC encoding module to achieve bit-level compression and bit rate control.
[0061] In this step, the Mel continuous latent features and the fundamental voiced / unvoiced continuous latent features are first aligned along the time axis.
[0062] After alignment, the hash input vector corresponding to each time step can be represented as:
[0063] in, This is the hash input vector.
[0064] The trainable hash bottleneck sequentially performs continuous latent feature projection and binarization on the hash input vector. Specifically, the Mel continuous latent feature corresponding to each time step is concatenated with the fundamental voiced / unvoiced tone continuous latent feature to form a hash input vector, which is then mapped to a K-dimensional hash logarithm space to obtain a K-dimensional continuous value vector. Then, a binarization operation is performed on the K-dimensional continuous value vector to obtain a binary bit vector with a value of -1 or +1.
[0065] The K-dimensional hash space projection can be expressed as:
[0066] in, and For trainable parameters, This is the result of a K-dimensional continuous hash projection.
[0067] Binarization can be represented as:
[0068] in, Represents a symbolic function. Represents a binary bit vector. .
[0069] Through the aforementioned trainable hashing bottleneck, the continuous latent features at each time step are compressed into a binary bit vector of length K. The binary bit vectors from all time steps are concatenated in chronological order to form the unified bit stream. Therefore, the unified bit stream can be represented as: ,in, This represents the number of time steps corresponding to the current speech segment.
[0070] The above describes the forward propagation inference process of the trainable hash bottleneck. Furthermore, the trainable hash bottleneck of this invention can employ a straight-through estimator (STE) or other differentiable approximation methods (such as Gumbel-Softmax, soft sign function, etc.) during the backpropagation training process, allowing the gradient to pass through the discretization operation, thereby enabling the binarization process to participate in end-to-end backpropagation.
[0071] The trainable hash bottleneck of this invention maintains the average bitrate of the unified bitstream within a preset range during training through a time window statistics and constraint mechanism. Specifically, the duration of each frame of speech is statistically set to... Allocate per frame If there are 100 bits, then the average code rate can be expressed as: When calculating the average bitrate within a time window, it can be represented as: ,in, To calculate the length of the statistical time window, For the first The number of bits corresponding to a frame. By adjusting the value of K, this invention controls the average bitrate within the range of approximately 0.8 to 1.2 kbps, limiting bitrate fluctuations over short time scales.
[0072] Through the above steps, the JSCC encoding module and the trainable hash bottleneck of this invention work together. The joint source-channel coding network is responsible for semantic feature separation, inequality error protection, and channel adaptive coding within the continuous feature domain, while the trainable hash bottleneck is responsible for mapping continuous latent features into a discrete bit stream with a stable bit rate. This division of labor enables the system to achieve a combination of progressive degradation, low bit rate, and stable bit stream performance across high-loss media channels.
[0073] Step S4: Perform keying modulation based on the unified bit stream and the channel state to obtain a modulated signal; transmit the modulated signal and pilot signal through the transmitting antenna; like Figure 5 As shown, in this step, the present invention is based on the unified bitstream and real-time channel state vector. Physical layer keying modulation is performed to generate a modulated signal suitable for transmission across high-loss media.
[0074] Specifically, based on the unified bitstream output by the hash bottleneck and The corresponding modulation parameters are used to generate a modulation waveform on a low-frequency magneto-electric mechanical antenna. A simple keying method is employed to map the bitstream to the amplitude, frequency, or multi-level states of a low-frequency carrier, generating a modulation signal to adapt to bandwidth-constrained and hardware implementation limitations in high-loss media environments. The transmitter inserts pilot fields according to the frame structure, and the receiver performs channel estimation and synchronization based on the pilots, updating accordingly. This provides feedback information for adjusting the coding strategy of subsequent speech segments. The joint source-channel coding and rate control function focuses on continuous features. The encoder is a bottleneck for trainable hashing, and the physical layer modulation is responsible for robustly radiating a given bitstream.
[0075] The above keying modulation process can be uniformly represented as:
[0076] in, Indicates keying modulation signal, This represents the mapping function for amplitude keying, frequency keying, or multilevel keying. This represents the set of modulation parameters determined based on the channel state.
[0077] In some embodiments, the modulation parameters include carrier frequency, symbol rate, modulation order, pulse shaping filter parameters, etc., and the modulation parameters are based on... Dynamic adjustment. Specifically, when When the path loss is high, the symbol rate can be reduced to increase the energy per bit; when When the indication delay spread is large, the protection interval can be increased to avoid inter-symbol interference; when When the noise power is high, binary FSK is used instead of multilevel modulation.
[0078] Pilot signals are inserted into the modulated signal for channel estimation, synchronization, and channel state updates at the receiver. The pilot signals employ a predetermined pseudo-random sequence or sine wave sequence, periodically inserted between data symbols at fixed time intervals; typical pilot overhead is 5% to 15%. The position, power, and sequence pattern of the pilot symbols are pre-agreed upon at both the transmitting and receiving ends. The receiver estimates parameters such as the channel's amplitude response, phase response, delay spread, and noise power spectral density based on correlation calculations between the received pilot signals and the local reference signal.
[0079] The transmit frame structure after pilot insertion can be represented as:
[0080] in, Indicates the transmit frame structure, For the first One pilot symbol, For the first A data signal symbol.
[0081] The modulated signal and pilot signal are transmitted via a transmitting antenna. The transmitting antenna is a low-frequency magneto-electric mechanical antenna, whose structure includes a magnetic core, coil windings, a resonant capacitor, and a power amplifier. The modulated signal, after being amplified, drives the antenna to radiate, generating a time-varying magnetic or electric field that propagates as an electromagnetic wave through a high-loss medium. Due to its extremely low operating frequency, it can penetrate high-loss media, enabling long-distance or deep-penetration communication.
[0082] Step S5: Update the channel state according to the pilot signal received by the receiving antenna and feed it back to the transmitting end; like Figure 1 and Figure 5 As shown, the receiving antenna receives a composite modulated signal containing pilot and data signals. The receiving antenna is also a low-frequency magneto-electric mechanical antenna, which captures electromagnetic field energy in space through magnetic induction or electric field coupling. After processing by a preamplifier, filter, and analog-to-digital converter, a digital baseband signal is obtained.
[0083] For pilot signals in digital baseband signals, the receiver of this invention estimates and updates the channel state.
[0084] In some embodiments, the receiver can estimate the amplitude response based on pilot correlation calculations, and the estimated amplitude response value can be expressed as:
[0085] in, This represents the magnitude response estimate. For locally stored transmit pilot reference signals, To receive pilot signals, This indicates the calculation of the inner product. Represents the norm.
[0086] In some embodiments, the received pilot signal can be compared with a locally stored transmitted pilot reference signal to obtain the amplitude attenuation of the pilot signal; the path loss can be estimated using the amplitude attenuation of the pilot signal, the delay spread can be estimated using the delay distribution of the multipath components, and the noise power spectral density can be estimated using the power of the received signal outside the pilot symbol. The estimated values are then normalized and quantized to update the channel state vector. The various components.
[0087] The path loss estimate can be expressed as:
[0088] in, This represents the estimated path loss. Furthermore, the channel state update process can be represented as:
[0089] in, This is the updated channel state vector. To estimate the channel state based on the current pilot signals, To smoothly update the coefficients, This indicates normalization and quantization processing.
[0090] The updated channel state vector is sent to the transmitter via a feedback link. In some embodiments, in a bidirectional communication system, the feedback information is transmitted to the transmitter via a reverse channel. The transmitter receives the feedback... Then, when encoding the next speech segment, the updated version will be used. By injecting Mel-based image-based joint source-channel coding network and F0 / VUV one-dimensional joint source-channel coding network, as well as trainable bottlenecks, adjusting coding strategies and redundancy allocation, and updating physical layer modulation parameters, adaptive tracking of time-varying channels can be achieved.
[0091] Through the above channel state estimation, update and feedback mechanism, the present invention can perceive the dynamic changes of the channel across high loss medium in real time, and adjust the coding and modulation strategies accordingly. When the channel conditions deteriorate, priority is given to ensuring the transmission of semantic intelligibility-related features, and when the channel conditions improve, the fidelity of timbre and detail is improved.
[0092] Step S6: Demodulate, hash decode, and perform joint source channel decoding on the modulated signal received by the receiving antenna in sequence, corresponding to those at the transmitting end, to obtain the Mel spectrum image and fundamental frequency voiced / unvoiced features; Based on the Mel cepstral coefficients and fundamental voiced / unvoiced features of the Mel spectrogram image, parameterized speech features are obtained by concatenating them; the parameterized speech features are then input into... A frame neural vocoder is used to obtain the reconstructed speech waveform.
[0093] In this step, such as Figure 6 As shown, after the receiving antenna receives the modulated signal, it sequentially performs demodulation, hash decoding, and joint source channel decoding processing corresponding to the transmitting end to recover the Mel spectrum image and the voiced / unvoiced characteristics of the fundamental frequency.
[0094] This invention first recovers the bitstream from the high-loss medium channel through demodulation at the receiving end, and then recovers it through a hash decoding network and continuous features. Decoder, in the updated Under certain conditions, Mel continuous latent features and fundamental voiced / unvoiced continuous latent features are reconstructed. The reconstructed Mel continuous latent features are then decoded by a decoding network to obtain a 32-dimensional Mel spectrum image, and the fundamental voiced / unvoiced continuous latent features are restored to fundamental voiced / unvoiced features.
[0095] The 32-dimensional Mel spectrum is then mapped to 18-dimensional Mel-frequency cepstral coefficients (MFCCs), which are concatenated with the fundamental frequency voiced / unvoiced features to obtain 20-dimensional parametric speech features. The 20-dimensional parametric speech features are then input into an LPCNet-like neural vocoder to reconstruct the time-domain speech waveform.
[0096] The The framework neural vocoder constructs a linear prediction filter based on the 20-dimensional parameterized speech features to obtain a time-varying vocal tract response. It then uses a lightweight neural network to perform autoregressive or near-autoregressive modeling of the excitation signal, generating sub-band excitations and performing inter-band synthesis in a multi-band structure. This reconstructs the fundamental frequency, formants, and spectral details with a relatively small model size, and integrates with the front-end... The structure together achieves progressive degradation and cliff-effect-free voice transmission characteristics. The LPCNet-like neural vocoder is obtained through pre-training on a large-scale speech dataset. Based on this 20-dimensional feature, the neural vocoder constructs a linear prediction filter to generate a time-varying vocal tract response. It then uses a lightweight neural network to perform autoregressive or near-autoregressive modeling on the excitation residuals, generates sub-band excitations in a multi-band structure, completes inter-band synthesis, and outputs a time-domain speech waveform.
[0097] Due to continuous features The stage assigns a high level of protection to semantic and prosodic features. Under conditions of extremely low signal-to-noise ratio and severe path loss, the present invention maintains intelligibility and continuity in terms of speech rhythm, fundamental frequency profile and large-scale timbre. High-frequency details gradually and smoothly decay as the channel deteriorates, and the entire link exhibits progressive degradation without cliff effect.
[0098] This invention also provides an end-to-end training step for low-bit-rate voice communication across high-loss media based on semantic feature separation. Specifically, all trainable modules, including the voice preprocessing and feature extraction at the transmitting end, the Mel image-based joint source-channel coding network, the F0 / VUV one-dimensional joint source-channel coding network, and the trainable hash bottleneck, as well as the hash decoding network, the joint source-channel decoding network, the parameter feature reconstruction, and the neural vocoder at the receiving end, are cascaded and jointly optimized under a unified loss function constraint.
[0099] In the joint optimization step, the present invention establishes a comprehensive loss function, which specifically includes: Mel spectrum reconstruction loss, multi-resolution spectrum loss, fundamental frequency deviation loss, voiced / unvoiced tone consistency loss, 20-dimensional parametric feature reconstruction loss, and bit rate deviation penalty term.
[0100] In some embodiments, the comprehensive loss function for end-to-end training can be expressed as:
[0101] in, Represents the comprehensive loss function. , , , , as well as These represent the Mel spectrum reconstruction loss, multi-resolution spectrum loss, fundamental frequency bias loss, voiced / unvoiced consistency loss, 20-dimensional parametric feature reconstruction loss, and bit rate bias penalty term, respectively. , , , , , These are the weights of each loss term.
[0102] The Mel spectrum reconstruction loss can be expressed as:
[0103] in, Represents the Mel spectrum. This represents the reconstructed Mel spectrum. This represents the L1 norm.
[0104] The fundamental frequency offset loss can be expressed as:
[0105] in, Indicates the total number of frames. Indicates the fundamental frequency. Indicates the predicted fundamental frequency; The loss of consistency between voiced and unvoiced sounds can be expressed as:
[0106] in, This indicates a marker for predicted voiced or unvoiced sounds.
[0107] The 20-dimensional parametric feature reconstruction loss can be expressed as:
[0108] in, The original 20-dimensional parametric speech features, The reconstructed 20-dimensional parametric speech features.
[0109] The bitrate deviation penalty term can be expressed as:
[0110] in, The average bitrate obtained within the time window. The target bitrate.
[0111] In some embodiments, the end-to-end training dataset may contain speech samples from multiple speakers and in multiple languages, covering different genders, ages, and pronunciation styles. Simultaneously, diverse channel conditions across high-loss media are constructed by randomly sampling the channel state vector. In each training batch, different channel conditions are applied to different speech samples. For example, a parameterized cross-high-loss medium channel model is used to apply additive noise, frequency-selective fading, multipath convolution, bandwidth limitation, and path loss to the transmission process of continuous latent features. This enables the network to learn robust encoding and decoding strategies under a wide range of channel conditions. The training process uses stochastic gradient descent, with a preset number of training rounds, until the comprehensive loss on the validation set converges, ultimately completing end-to-end training.
[0112] While the specific embodiments of the present invention depict actions or steps in a particular order, this should be understood as requiring such actions or steps to be performed in the shown specific order or sequential order, or requiring all illustrated actions or steps to be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations. The above descriptions are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention.
[0113] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A low-bit-rate voice communication method across high-loss media based on semantic feature separation, characterized in that, Includes the following steps: Step S1: Preprocess and transform the speech to be transmitted to obtain the Mel spectrum image and fundamental frequency voiced / unvoiced features; Step S2: Based on the real-time channel state, perform joint source channel coding on the Mel spectrum image and the fundamental frequency voiced / unvoiced features to obtain Mel continuous hidden features and fundamental frequency voiced / unvoiced continuous hidden features. Step S3: Align the Mel continuous latent features and the fundamental voiced / unvoiced continuous latent features in time and combine them into a hash input vector; input the hash input vector into the trainable hash bottleneck, and obtain a unified bitstream by K-dimensional hash space projection and binarization operation; Step S4: Perform keying modulation based on the unified bit stream and the channel state to obtain the modulated signal; The modulation signal and pilot signal are transmitted through the transmitting antenna; Step S5: Update the channel state according to the pilot signal received by the receiving antenna and feed it back to the transmitting end; Step S6: Demodulate, hash decode, and perform joint source channel decoding on the modulated signal received by the receiving antenna in sequence, corresponding to those at the transmitting end, to obtain the Mel spectrum image and fundamental frequency voiced / unvoiced features; Based on the Mel cepstral coefficients and fundamental voiced / unvoiced features of the Mel spectrogram image, parameterized speech features are obtained by concatenating them; the parameterized speech features are then input into... A frame neural vocoder is used to obtain the reconstructed speech waveform.
2. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 1, characterized in that, In step S1, the preprocessing specifically includes performing pre-emphasis, frame windowing, and endpoint detection operations on the original speech signal in sequence; The feature transformation specifically includes performing frequency domain transformation on each frame of speech signal using short-time Fourier transform, calculating the power spectrum and mapping it to the Mel scale to obtain a 32-dimensional logarithmic Mel spectrum sequence, which serves as the Mel spectrum image. For each frame of speech, the fundamental frequency F0 and the voiced / unvoiced probability VUV are extracted. The voiced / unvoiced probability VUV takes the value of 0 or 1, where 1 represents unvoiced and 0 represents voiced.
3. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 2, characterized in that, In step S2, joint source-channel coding is performed on the Mel spectrum image based on the real-time channel state, specifically including: Two-dimensional convolution is performed on the Mel spectrum image. The features obtained from the two-dimensional convolution are input into a state space network. The output features are then subjected to gated linear attention processing to obtain the Mel continuous latent features. During this process, the channel state is used... Dynamically adjust feature redundancy allocation, channel weights, and encoding strategies; In step S2, joint source-channel coding is performed on the fundamental frequency voiced / unvoiced tone features based on the real-time channel state, specifically including: The fundamental frequency and voiced / unvoiced features are input into a one-dimensional sequence modeling network to generate continuous latent features of the fundamental frequency and voiced / unvoiced sounds; during this process, coding redundancy and quantization precision are adjusted by channel state. The channel state This includes: normalized path loss factor, normalized delay spread scale, normalized noise power, and scene outline number; The channel state The prior values are obtained as follows: Transceiver antennas are deployed in a typical high-loss medium scenario to collect measured data on path loss, delay spread, and noise power. Parameters are fitted based on a hierarchical medium propagation model, and the channel state is obtained after normalization and quantization. The prior value; The expression for the joint source-channel coding is: in, This is a continuous latent feature of Mel. This is a continuous hidden feature of the fundamental frequency voiceless / unvoiced tone. and These represent the Mel-image joint source-channel coding network and the fundamental frequency voiced / unvoiced one-dimensional joint source-channel coding network, respectively. Indicates the fundamental frequency. Indicates the voicing or unvoicing of consonants.
4. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 3, characterized in that, In step S3, the training process of the trainable hash bottleneck adopts a pass-through estimator; in addition, during the training process, the average bitrate of the unified bitstream is kept within a preset range through time window statistics and bitrate constraints. In step S3, after K-dimensional hash space projection and binarization, the expression for obtaining the unified bitstream is: in, For the hash input vector, and For trainable parameters, For the K-dimensional continuous hash projection result, Represents a symbolic function. This represents a binary bit vector.
5. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 4, characterized in that, Step S4 specifically includes: according to The corresponding modulation parameters are used to map the unified bit stream output from the hash bottleneck using amplitude shifting, frequency shifting, or multilevel shifting to generate a modulated signal. Generate pilot signals and insert them into the modulated signal according to the frame structure; The modulation signal and pilot signal are transmitted via a low-frequency magneto-electric mechanical antenna.
6. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 5, characterized in that, Step S5 specifically includes: The receiving antenna extracts the pilot signal, compares it with the locally stored transmit pilot reference signal, and obtains the amplitude attenuation of the pilot signal; the channel state is updated based on the amplitude attenuation of the pilot signal. The expression is: in, This is the updated channel state vector. The channel state vector before the update. To estimate the channel state based on the current pilot signals, To smoothly update the coefficients, This indicates normalization and quantization processing; The updated It is sent to the transmitter via the feedback link.
7. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 6, characterized in that, Step S6 includes: The receiver recovers the bitstream from the high-loss medium channel through demodulation, and then uses a hash decoding network and continuous features. Decoder, in the updated Under the given conditions, the Mel continuous latent features and the fundamental voiced / unvoiced continuous latent features are reconstructed; the reconstructed Mel continuous latent features are then restored to a 32-dimensional Mel spectrum image by a decoding network, and the fundamental voiced / unvoiced continuous latent features are restored to fundamental voiced / unvoiced features. The 32-dimensional Mel spectrum is then mapped to 18-dimensional Mel cepstral coefficients, which are concatenated with the fundamental voiced / unvoiced features to obtain 20-dimensional parametric speech features. The 20-dimensional parametric speech features are then input into a neural vocoder based on an LPCNet-like framework to obtain the reconstructed speech waveform.
8. The low-bit-rate voice communication method across high-loss media based on semantic feature separation according to claim 7, characterized in that, End-to-end training is performed on a low-bit-rate voice communication process across high-loss media based on semantic feature separation. The expression for the comprehensive loss function of the end-to-end training is as follows: in, Represents the comprehensive loss function. , , , , as well as These represent the Mel spectrum reconstruction loss, multi-resolution spectrum loss, fundamental frequency bias loss, voiced / unvoiced consistency loss, 20-dimensional parametric feature reconstruction loss, and bit rate bias penalty term, respectively. , , , , , These are the weights of each loss term; During the end-to-end training process, different additive noise, frequency-selective fading, multipath convolution, bandwidth limitation, and path loss are applied to the transmission process of continuous latent features.