Helium speech processing method across deep acoustic feature fusion and dynamic phoneme error correction

By employing a method of cross-depth acoustic feature fusion and dynamic phoneme error correction, the nonlinear distortion and noise interference problems of helium speech in the deep-sea environment were solved, achieving high-fidelity speech reconstruction and low-latency communication, which is suitable for deep-sea saturation diving operations.

CN122224202BActive Publication Date: 2026-08-25SHANGHAI MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610685555.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-25
Estimated Expiration
2046-05-19

AI Technical Summary

Technical Problem

Helium speech distortion affects divers' physiological condition assessment and communication. Traditional methods struggle to cope with nonlinear distortion and complex noise, resulting in insufficient speech clarity in reconstruction. Conventional deep learning models have poor generalization ability and cannot meet the communication requirements of high-risk deep-sea operations.

Method used

A method combining cross-depth acoustic feature fusion and dynamic phoneme error correction is adopted, including helium speech formant nonlinear compensation, frequency domain feature extraction, helium environment characterization matrix construction, helium speech neural inference and synthesis, multi-scale wavelet transform, HiFi-GAN vocoder and other technologies, combined with edge deployment strategy to achieve high-fidelity speech reconstruction.

Benefits of technology

Achieving high-fidelity reconstruction of helium speech in deep-sea environments reduces word error rate to 8.2%, improves speech naturalness and high-frequency harmonic integrity, meets real-time communication requirements, reduces model memory usage by 48%, and is suitable for deep-sea terminal deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224202B_ABST
    Figure CN122224202B_ABST
Patent Text Reader

Abstract

The application provides a helium speech processing method combining deep acoustic features and dynamic phoneme correction, and relates to the technical fields of artificial intelligence and intelligent speech signal processing. The method comprises the following steps: denoising, formant compensation and frequency domain transformation of helium speech to generate two-dimensional time-frequency features; combining depth and helium-oxygen concentration to construct an environment representation matrix, compensating distortion through He-NIS operators, He-DDCA attention, hybrid expert architecture and VB-Adapter, outputting text by combining CTC-AED loss and physical rule phoneme correction; generating mel spectrum based on a high-risk phoneme dictionary and wavelet transform reconstruction of the fundamental harmonic; outputting high-fidelity speech through an optimized HiFi-GAN vocoder, and realizing edge low-latency deployment through mixed precision quantization and streaming scheduling. The application has high recognition accuracy, good speech fidelity and low latency, and is suitable for real-time communication needs of deep-sea saturation diving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and intelligent speech signal processing technology, specifically to a helium speech processing method that integrates cross-depth acoustic feature fusion and dynamic phoneme error correction. Background Technology

[0002] With the rapid development of my country's marine economy, more and more resources will be extracted from the ocean. Saturation diving, a crucial operation for exploring marine resources, involves divers inhaling a mixture of helium and oxygen. This allows them to adapt to the high-pressure environment of the deep sea without experiencing nitrogen sedation. However, this helium-oxygen mixture causes significant changes in the diver's voice, resulting in "helium speech." The distortion caused by helium speech not only affects the ability of ground support personnel to assess the diver's physiological condition but also hinders communication between divers, leading to reduced safety and efficiency in diving operations. Helium speech communication is indispensable for saturation diving operations and is a key technology for deep-sea saturation diving, serving as the only communication method to ensure the smooth execution of such operations. Therefore, solving the problem of helium speech distortion is an urgent issue.

[0003] Various attempts have been made to solve this problem, primarily through traditional methods such as time-domain processing, linear prediction, and homomorphic signal processing. However, these methods have the following significant drawbacks in practical applications: First, traditional helium speech recognition methods are difficult to cope with nonlinear distortion and complex noise. Traditional methods are mostly based on linear acoustic models, which cannot accurately fit the nonlinear shift of vocal cord resonance peaks caused by helium, and are difficult to effectively compensate for high-frequency energy attenuation under stable background noise in the deep sea.

[0004] Secondly, the reconstructed speech lacks clarity and naturalness. Traditional methods rely solely on traditional signal processing to extract and recover acoustic features, resulting in poor sound quality, a mechanical feel, and poor semantic coherence in long contexts. This fails to meet the stringent requirements of high-risk deep-sea operations for communication accuracy and high fidelity.

[0005] Third, existing conventional deep learning speech models have poor generalization ability. When the general end-to-end ASR and TTS concatenated model is directly applied to helium speech, the network is not specifically constrained by the severe attenuation of high-frequency harmonics and nonlinear deformation of phonemes caused by the helium environment. This easily leads to the loss of high-frequency details and the collapse of highly deformed phoneme recognition, resulting in severe distortion of the final reconstructed speech, which cannot be used as a reliable solution for deep-sea operations.

[0006] This is insufficient to guarantee the safety of divers in the dangerous field of deep-sea saturation diving. Therefore, it is necessary to provide a novel helium speech reconstruction method based on deep learning to solve the above problems and meet the high requirements of deep-sea saturation diving operations. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a helium-based speech processing method that integrates cross-depth acoustic feature fusion and dynamic phoneme error correction. Employing speech recognition and high-fidelity text-to-speech technology, it achieves rapid and accurate reconstruction of distorted speech in a helium-filled environment.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A helium speech processing method with cross-depth acoustic feature fusion and dynamic phoneme error correction, comprising the following steps: S1. Calculate a dedicated noise masking threshold for the collected helium speech to remove deep-sea noise. Use a first-order finite impulse response pre-emphasis filter to achieve nonlinear compensation of the formant peak. After zero-padding, perform a 512-point fast Fourier transform to obtain frequency domain features. Generate a helium speech distortion feature map based on the piecewise adaptive sensing weight function and the environmental sound speed variation ratio. Introduce a dynamic scaling factor driven by depth and helium-oxygen concentration to perform adaptive cepstral mean-variance normalization and output a two-dimensional time-frequency feature matrix. S2. Using the two-dimensional time-frequency feature matrix as input, a helium environment characterization matrix is ​​constructed by combining the diving depth and helium-oxygen concentration. The front-end feature extraction is completed by using helium speech neural inference and synthesis He-NIS operator. Phase misalignment and high-frequency dispersion are compensated by helium speech audio dispersion dynamic convolution attention He-DDCA. Cross-depth distortion compensation is achieved by using environmental parameter perception hybrid expert architecture and variational Bayesian adapter VB-Adapter. After connectionist temporal classification CTC and attention encoder-decoder AED joint loss training and physical rule-constrained bundle search phoneme correction, the text is output. S3. Based on the text, construct a helium high-risk phoneme degradation mapping dictionary and incorporate physical condition priors. Decouple the fundamental frequency through continuous wavelet transform at multiple scales. Complete the multi-scale reconstruction of the fundamental frequency and harmonics according to the sound speed variation ratio. Input the prosodic parameters and text features into a non-autoregressive decoder to obtain the Mel spectrum. S4. Input the Mel spectrum into the optimized high-fidelity generative adversarial network HiFi-GAN vocoder of the multi-receptive field fusion MRF module with built-in high-frequency harmonic enhancement branch, and output high-fidelity speech through adversarial optimization. Low-latency deployment at the edge is achieved through high-frequency preserved mixed precision quantization and adaptive streaming scheduling.

[0009] Compared with the prior art, the beneficial effects of the present invention are: This invention addresses the shortcomings of existing methods, such as nonlinear distortion of helium speech, severe attenuation of high-frequency harmonics, interference from complex noise in the deep sea, poor sound quality, low recognition accuracy, and insufficient real-time performance, through a comprehensive technical solution of cross-depth acoustic feature fusion and dynamic phoneme error correction. It achieves accurate and rapid reconstruction of distorted speech into high-fidelity normal speech in a helium-oxygen environment.

[0010] In terms of recognition performance, relying on the helium acoustic nonlinear inverse stretching operator, helium speech audio dispersion dynamic convolutional attention, environmental parameter perception hybrid expert architecture and physical rule-driven cluster search phoneme correction, it significantly corrects formant shift and phoneme deformation. The actual test word error rate (WER) is reduced to 8.2%, which is far better than traditional signal processing methods and conventional deep learning models. It still maintains stable recognition under extreme depth and strong noise, and completely solves the problems of high-frequency detail loss and deformed phoneme recognition collapse.

[0011] In terms of speech reconstruction quality, the fundamental frequency and harmonics are accurately decoupled and reconstructed through multi-scale wavelet transform. A dedicated high-frequency harmonic enhancement branch is added to the HiFi-GAN vocoder to effectively compensate for the high-frequency energy loss caused by helium. The naturalness score of the synthesized speech reaches 4.1 / 5, and the high-frequency harmonic integrity is improved by 40% compared with the traditional method. Mechanical distortion and semantic breaks are eliminated, and the diver's real timbre and pronunciation rhythm are fully preserved.

[0012] In terms of engineering practicality, it adopts non-autoregressive decoding, adaptive streaming scheduling and multi-threaded ring buffer design, with end-to-end processing latency of less than 150ms, meeting the requirements of full-duplex real-time communication; at the same time, it achieves model lightweighting through ONNX conversion, operator fusion and high-frequency preserved hybrid precision quantization, reducing memory usage by 48%, and can be stably deployed on edge terminals such as deep-sea mother ships and diving bells. It has the comprehensive advantages of full-depth adaptability, strong anti-interference and low computing power consumption, and is fully adapted to the harsh communication requirements of deep-sea saturation diving operations.

[0013] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0014] Figure 1 This is a flowchart of the helium speech processing method for cross-depth acoustic feature fusion and dynamic phoneme error correction provided by the present invention; Figure 2 This is a schematic diagram of the speech recognition process provided by the present invention; Figure 3 This is a schematic diagram of the speech synthesis process provided by the present invention; Figure 4 This is the short-time window spectrum diagram provided by the present invention. Detailed Implementation

[0015] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings, so as to more clearly understand the purpose, features and advantages of this invention. It should be understood that the embodiments shown in the drawings are not intended to limit the scope of this invention, but are only for illustrating the essential spirit of the technical solutions of this invention. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.

[0016] Unless the context requires otherwise, throughout the specification and claims, the word “comprising” and its variations, such as “including” and “having”, shall be understood to have an open, inclusive meaning, that is, to be interpreted as “including, but not limited to”.

[0017] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.

[0018] The singular forms “a” and “the” used in this specification and the appended claims include plural references unless otherwise expressly stated herein. It should be noted that the term “or” is generally used to mean “and / or” unless otherwise expressly stated herein.

[0019] In the following description, in order to clearly demonstrate the structure and working method of the present invention, a number of directional terms will be used. However, terms such as "front", "back", "left", "right", "outside", "inside", "outward", "inward", "up", and "down" should be understood as convenient terms and not as limiting terms.

[0020] The following is combined with Figures 1 to 4 The implementation details of the embodiments of the present invention will be described in detail below. The following content is only for the convenience of understanding the implementation details and is not necessary for implementing this solution.

[0021] This invention provides a helium speech processing method with cross-depth acoustic feature fusion and dynamic phoneme error correction, such as... Figure 1 As shown, the main steps of this processing method are as follows: S1. Signal acquisition and preprocessing; To address the steady-state noise interference problem of deep-sea mothership equipment, this invention designs and calculates a helium-based noise masking threshold for speech: First, the energy spectrum of the preceding speech segment is extracted to calculate the initial noise floor. Then, combining the physical property of helium to concentrate speech energy to high frequencies, a dynamic frequency band weighting function is constructed to increase the judgment weight of high-frequency bands in endpoint detection. By dynamically calculating the masking threshold, the start and end points of effective speech segments are accurately determined, eliminating stable background noise in the deep-sea environment. The obtained effective speech segments are further preprocessed using the following steps: Step 1: Using a helium speech formant nonlinear compensation algorithm, a first-order finite impulse response (FIR) pre-emphasis filter is applied to the effective speech segment. The time-domain difference equation of this filter is: In the formula It is a dynamic compensation coefficient that is adaptively adjusted according to the helium-oxygen mixing ratio at different depths. The default reference value is 0.97, which is used to correct the high-frequency distortion caused by the nonlinear upward shift of the resonance peak at different depths.

[0022] Step 2: Framing and Windowing: This involves processing the pre-emphasis speech signal... Perform a framing operation, dividing the continuous signal into segments with a frame length of 25ms and a frame shift of 10ms. For the i-th frame of the segmented signal... A Hamming window is applied to suppress spectral leakage during frequency domain transformation. The windowing operation expression is as follows:

[0023] in, The length of a single frame signal.

[0024] Step 3: Helium Language Audio Domain Feature Mapping: This method uses a 16kHz sampling rate and a single frame signal length of... ,because For integer powers of 2, after padding each frame with 112 zeros, a 512-point Fast Fourier Transform is performed on the framed signal to convert the time-domain signal into a frequency-domain complex sequence. Subsequently, the discrete power spectral density was calculated. The construction length is One-dimensional perceptual weight vector ,in This represents the total number of filters. The weights of each filter Based on its center frequency The frequency band in which it is located is calculated using the following piecewise adaptive function:

[0025] in: The preset cutoff frequency threshold for deep-sea background mechanical noise (such as life support systems and breathing valves); and These are the low-frequency noise attenuation coefficient and the smoothing index, respectively. Helium speech High-frequency enhancement factor; is the environmental sound velocity variation ratio, defined as... ( Current diving depth Helium-oxygen mixing ratio The speed of sound in the mixed gas below (This refers to the speed of sound in air at normal pressure). Next, the power spectral density... Mapped to a scaled filter bank and combined with the generated perceptual weight matrix Feature weighting fusion is performed to generate a helium-specific distortion feature map. After taking the natural logarithm of the output energy of each filter, adaptive cepstral mean-variance normalization (CMVN) is executed. Here, a dynamic scaling factor jointly determined by the diving depth and helium-oxygen concentration is introduced to adaptively adjust the offset of the mean and variance, completely eliminating acoustic distortion caused by the variable channel in the deep-sea confined chamber. The final output is a two-dimensional time-frequency feature matrix for subsequent processing by the depth acoustic model.

[0026] S2, speech recognition; After the aforementioned signal acquisition and preprocessing, a two-dimensional time-frequency feature matrix specific to helium speech has been obtained, eliminating deep-sea channel distortion and compensating for high-frequency formant shift. Based on this preprocessed feature, speech recognition processing is performed, including acoustic modeling and text decoding, to achieve accurate mapping from helium speech to standard text, such as... Figure 2 This is a structural diagram of the speech recognition module. The specific steps are as follows: Step 1: Construction of the helium environment characterization matrix; The system first uses external sensors on the diving equipment to synchronously collect the diving depth at the moment of sound in real time. (Unit: meter) and the percentage of helium concentration in the gas mixture. Based on the derivation of the ideal gas sound speed equation, a real-time deep-sea sound speed ratio calculation model was constructed:

[0027] in This represents the molar mass of the corresponding gas. Subsequently, the nonlinear frequency shift parameter resulting from this ratio is extracted and combined with the diving depth. Based on the water pressure attenuation coefficient of the corresponding environment, a physical prior representation matrix specific to helium speech is generated. This matrix will not only serve as a routing guide for subsequent networks, but will also directly participate in the reconstruction of the underlying feature extraction operators.

[0028] Step 2: Front-end feature extraction based on the Helium Acoustic Nonlinear Inverse Stretching Operator (He-NIS Operator); To address the non-stationary distortion problem of severe high-frequency stretching and slight low-frequency drift in speech formants under helium conditions, conventional two-dimensional convolution sampling grids cannot capture this nonlinear frequency shift. Therefore, this invention proposes the He-NIS operator, a helium acoustic nonlinear inverse stretching operator, to replace the traditional convolution front-end.

[0029] Before performing the inner product operation, the He-NIS operator forcibly introduces an elastic coordinate offset based on the physical equations. The elastic coordinate offset Strictly constrained by the sound velocity ratio in step one:

[0030] in, This represents the frequency axis coordinate of the current sampling point. For hyperparameters, This is a learnable weight matrix. Through this physical constraint, the He-NIS operator maintains regular sampling in the low-frequency region, while its sampling grid changes according to frequency in the high-frequency region. The rate exhibits an extreme nonlinear expansion, and the eigenvalue at the current position after nonlinear stretching and rearrangement is... Its calculation formula is:

[0031] in, For the present The sampling point is located on the frequency axis after nonlinear stretching, i.e., the position of the rearranged high-frequency expanded grid. Represents the first frequency on the original frequency axis before stretching. Coordinates of each sampling point.

[0032] This is determined by the depth of the dive. and helium concentration Real-time driven adaptive elastic sampling mitigates the nonlinear stretching distortion caused by helium gas. Short-sequence modeling is then performed using a bidirectional gated cyclic unit (Bi-GRU) to output the fundamental acoustic features. .

[0033] Step 3: Helium-based Acoustic Dispersion Dynamic Convolutional Attention (He-DDCA). Conventional attention mechanisms cannot capture the acoustic phase misalignment and severe high-frequency dispersion caused by the inconsistency in high and low frequency wave velocities in helium media. This invention combines the fluid sound velocity equation with similarity calculation: Physical phase compensation mapping: As a constraint, dynamic deep convolution is used to map the input features, giving the severely damaged high-frequency channels a wider temporal receptive field, compensating for the phase lead of the sound waves in helium, and obtaining a time-aligned... and .

[0034] Reconstructing the attention scoring space: For the first time in and Interspersed by Generated medium dispersion modulation tensor The reconstructed attention scoring formula is defined as follows:

[0035] In the formula, The reconstructed attention scoring matrix; This is the query feature matrix aligned on the time axis after physical phase compensation mapping; This is the transpose of the key feature matrix aligned on the time axis after physical phase compensation mapping; The helium environment characterization matrix (physical prior features generated from diving depth and helium-oxygen concentration); For the reason Real-time generated medium dispersion modulation tensor; Key vector The feature dimension size is a common parameter in standard scaled dot product attention, used to prevent the softmax gradient from vanishing due to excessively large inner product values.

[0036] pass The scoring space is reconstructed into a physical Markov space that conforms to the acoustic attenuation law of helium, and physical penalties are imposed on the high-frequency channels that are prone to distortion, while high confidence is given to the stable low-frequency channels.

[0037] Broadband Harmonic Reconstruction Branch: Generation Matrix-time parallel dynamic depthwise convolution branches, through Real-time generation of convolution kernel dilation rate Capture the high-frequency micro-vibration characteristics caused by helium gas dispersion. The output is:

[0038] Solve the feature misalignment and collapse problem under fluid distortion environment, and output global acoustic features. .

[0039] Step 4: Hybrid expert architecture for nonlinear distortion of helium speech with water depth. In actual deep-sea saturation diving operations, with the increase of water depth gradient, the nonlinear displacement span of vocal tract formants, the energy attenuation rate of high-frequency harmonics, and the degree of phase misalignment caused by medium dispersion in helium speech all exhibit a step-like non-stationary evolution. Traditional acoustic networks, when dealing with such cross-depth variations, are prone to alignment collapse in extreme depth environments due to their overly narrow feature search space.

[0040] To address this, this invention innovatively introduces a hybrid expert architecture based on environmental parameter awareness, achieving full-depth adaptation through a "segmented expert compensation" strategy driven by physical parameters. This architecture consists of one shared expert and... It consists of a variational specific expert branch tailored to a specific water depth range.

[0041] Directly use the steps in step one Real-time diving depth extracted from matrix Calculate the physical activation weights of experts in each branch. The expression is:

[0042] in, The target water depth anchor point (e.g., 50m, 100m, 200m, 300m, etc.) corresponds to the i-th expert branch. Using this routing method, the system can automatically and smoothly switch between expert branches at different depths based on the diver's physical depth.

[0043] The activated expert branch employs a variational adapter (VB-Adapter) structure to separate the distorted signal from the damaged signal: Physical condition posterior path (right branch): combined The helium-oxygen ratio and water pressure parameters in the model were used to construct a conditional posterior distribution. This branch is specifically responsible for capturing deterministic frequency shift compensation in helium speech.

[0044] Acoustic variational distribution path (left branch): Generates the observational variational distribution The technique of reparameterization is used to capture unsteady residual fluctuations caused by individual differences in divers or complex echoes.

[0045] In the output phase of the expert branch, this invention designs a physical weight adaptive fusion mechanism. The final acoustic feature tensor... Generated using the following weighted logic:

[0046] This forms the basic semantic backbone for sharing expert outputs, serving as a global reference. The output of each VB-Adapter is determined by its corresponding physical activation weights. Scaling is applied. This means that when a diver is in medium to deep water, the corresponding expert branch will receive the highest weight. Within each expert branch, learnable parameters are used... Dynamic adjustment of deterministic compensation amount With stochastic adjustment The proportion. In environments with severe depth fluctuations, the system will automatically increase the learnable parameters of the physical condition path. This is to strengthen the constraints on feature reconstruction.

[0047] Step 5: Multi-task loss function. Projecting onto the vocabulary dimension, a joint training approach is adopted, combining connection-based temporal classification (CTC) and sequence-to-sequence attention decoding (AED). To overcome the alignment collapse in the decoder caused by the dispersion of helium gas in the deep sea, a multi-task loss function is designed:

[0048] in, To preset hyperparameters, maintain them for the first 10 epochs of training. By leveraging the strict timing hard alignment provided by CTC, attention divergence caused by initial helium dispersion is forcibly suppressed; subsequently, it is linearly decayed batch by batch to... Gradually, relying on the contextual reasoning ability of AEDs, errors in deformed phonemes can be corrected.

[0049] Step Six: Bundle Search Phoneme Correction Based on Acoustic Physics Rules. Building upon this, to thoroughly address the high misidentification rate of traditional ASR models in deep-sea environments, the decoder backend introduces helium speech phoneme correction rules constrained by both data and physical rules. Specifically, addressing the nonlinear distortion of high-frequency formants caused by changes in sound velocity in helium environments, the system first identifies the most severely affected fricatives (e.g., / s / , / f / ) and affricates (e.g., / z / , / c / ). Since the broadband noise features of these phonemes are highly susceptible to feature overlap due to frequency shifts in helium, leading to misalignment by traditional acoustic models, this invention establishes a priori transition matrices for these high-risk deformable phonemes by statistically analyzing misclassified phoneme pairs in actual helium speech datasets. The elements in this matrix Precise characterization of phonemes at specific helium-oxygen mixing ratios and depths. Misidentified as a phoneme due to acoustic distortion The prior probability.

[0050] Furthermore, in the final text decoding stage, this scheme reconstructs the scoring function for the bundle search. The original score for the search path is... When the candidate search path expands to a high-risk deformable phoneme and the acoustic confidence is low, a dynamic probability penalty mechanism is triggered. Specifically, the system determines the current candidate phoneme through set inclusion operations. Does it belong to the embedded high-risk deformed phoneme hash table? And combined with information entropy based on the probability distribution of the current time step Adaptive confidence threshold Make a judgment.

[0051] when And the back-delay probability output by the acoustic model At that time, calculate the corrected path score. The calculation formula is as follows:

[0052] in, The weighting coefficient for penalties for misjudgment; This represents the confidence score of the current acoustic model for the candidate phonemes; An adaptive confidence threshold; For indicator functions (when) The value is 1 if it is active, and 0 otherwise. This represents the compensation coefficient for the correct path. Through this equation, when the system encounters easily confused nonlinear deformed phonemes, the algorithm automatically applies a large negative penalty to the incorrect path based on the prior matrix, while simultaneously providing probabilistic compensation to the correct phoneme path that conforms to the physical laws of helium sound production. Through this error correction mechanism based on mathematical expectation, this method completely blocks the propagation of error decoding under extreme deformed acoustic features in traditional models, achieving highly robust text mapping.

[0053] S3 speech synthesis; After accurately decoding the distorted helium speech into standard text through the above speech recognition steps, the resulting standard text is used to enter the speech synthesis stage. High-fidelity speech reconstruction and high-frequency harmonic compensation are performed to output natural and clear normal speech, such as... Figure 3 This is a structural diagram of the speech synthesis module. The specific implementation steps are as follows: Step 1: Text feature extraction based on the prior knowledge of the physical degradation of high-risk phonemes in helium.

[0054] First, the input text undergoes front-end processing. Addressing the unique phenomenon of severe physical degradation of certain phonemes in the high-pressure helium-oxygen environment of the deep sea (e.g., high-frequency fricatives / s / and / f / experiencing sharp energy reduction due to turbulence caused by changes in gas density; high vowels / i / and / u / causing auditory confusion due to formant broadening), this step constructs a helium high-risk phoneme degradation mapping dictionary. During feature extraction, not only are basic phonemes transformed, but further, based on the aforementioned dictionary, a cascaded compensation mask is created for phoneme nodes identified as high-risk. This mask carries the current diving depth of the environment. With helium concentration As a priori physical condition, the phoneme sequence carrying the physical condition mask is then input into the Transformer text encoder, enabling the network to establish prior knowledge of which pronunciations will undergo what kind of distortion at the current depth during the text understanding stage, and output a context tensor with helium speech compensation features.

[0055] Step 2: CWT fundamental frequency multi-scale decoupling.

[0056] To address the involuntary fundamental frequency drift and high-frequency jitter observed in deep-sea divers due to abnormal auditory feedback from sealed helmets and changes in the helium medium, continuous wavelet transform is employed for frequency band separation, achieving decoupling across multiple physical dimensions. Based on the specialized mechanism that vocal cord vibration is less affected by gas but the vocal tract is greatly affected by helium, the Mexican cap wavelet is used as the generating function. Calculate the wavelet coefficients:

[0057] frequency trajectory Decomposed into wavelet coefficients of multiple scales The scaling factor in the equation It has a strict compressibility ratio with the acoustic wavelength of a helium-oxygen mixture.

[0058] Step 3: Fundamental-harmonic multiscale reconstruction controlled by the environmental sound speed variation ratio.

[0059] To eliminate distortion caused by helium and preserve the diver's individual voice, a dynamic scale threshold controlled by the environmental sound speed variation ratio was set. The threshold satisfies the function

[0060] This is directly determined dynamically by the ratio of the sound velocity in the mixed gas at the current depth to the sound velocity in air at normal pressure. During frequency band reconstruction: the scale factor... The low-frequency coefficients are preserved without loss as actual physiological vocal cord vibration components; The high-frequency coefficients are used as nonlinear resonance distortion components of the helium cavity, and a damping smoothing penalty that is positively correlated with the sound speed variation ratio is applied.

[0061] Step 4: Helium speech fundamental frequency-harmonic multiscale recovery.

[0062] The network uniformly adopts mean squared error as the optimization objective, and performs independent regression predictions on each scale component of duration, energy, and fundamental frequency to accurately fit the distortion distribution of speech prosody parameters under helium conditions, thereby improving the accuracy of duration alignment, energy compensation, and fundamental frequency reconstruction. In the fundamental frequency recovery stage, this invention employs a multi-scale weighted synthesis algorithm. Based on the dynamic scale weights determined by the diving depth and helium-oxygen mixing ratio, the fundamental frequency scale components obtained by continuous wavelet transform decomposition are weighted and calibrated. Subsequently, the calibrated scale components are fused and reconstructed through inverse wavelet transform to fully restore the inherent micro-vibration characteristics of the diver's real vocal cord vibration. This effectively eliminates the unconscious fundamental frequency drift and high-frequency distortion jitter caused by the helium medium and the high-pressure environment of the deep sea, restoring the natural fundamental frequency contour that conforms to the physiological pronunciation rules, and providing reliable prosodic support for subsequent Mel spectrum generation and high-fidelity speech synthesis.

[0063] Step 5: Non-autoregressive decoding and time-domain waveform generation.

[0064] During the decoding stage, the generated fundamental frequency, duration, and energy parameters, along with the text features, are input into a FastSpeech2-type non-autoregressive decoder. This efficiently maps the text and prosodic features into Mel spectra, avoiding the accumulation of timing errors and delay amplification caused by autoregressive decoding. Finally, the resulting Mel spectra are input into a high-frequency optimized neural vocoder (HiFi-GAN) to complete the conversion from frequency domain features to time domain waveforms, outputting high-definition, high-naturalness reconstructed speech that meets the high-fidelity and low-latency requirements for deep-sea saturation diving communication.

[0065] Step 6: High-fidelity waveform reconstruction based on an optimized HiFi-GAN vocoder.

[0066] The vocoder module adopts a HiFi-GAN architecture, including a generator and two types of discriminators: a multi-scale discriminator and a multi-period discriminator. To address the high-frequency harmonic loss problem that is prone to occur in helium speech reconstruction, the generator employs a multi-receptive field fusion module (MRF) specifically optimized for the high-frequency attenuation features of helium speech. This invention introduces an additional high-frequency harmonic enhancement branch on top of the three parallel sets of one-dimensional residual convolutional branches within the MRF (with dilation rates set to d∈{1,3,5}). Specifically, this enhancement branch embeds a helium speech high-frequency energy compensation algorithm: it uses a combination of small convolutional kernels and dense holes (dilation rate d∈{1,2,4}) to extract fine-grained high-frequency short-wavelength features, and connects a channel-level attention mechanism at the end of the branch to dynamically calculate the high-frequency weight mask of the feature channels, thereby achieving adaptive gain compensation for the lost high-frequency energy. The compensated high-frequency features are concatenated and fused with the regular branch features at the channel dimension. Combined with a leakage correction linear unit (negative half-axis slope α=0.1), the MRF module observes feature maps with different receptive field sizes in parallel. By setting one-dimensional convolutional layers with different dilation rates, it effectively captures the fundamental frequency and regular fine-grained harmonic features of the synthesized speech while ensuring computational efficiency. Without increasing the number of feature parameters excessively, the MRF module observes feature maps with different time spans in parallel through dilated convolutions with varying strides, multiplying the time-axis receptive field. Relying on the newly added high-frequency harmonic enhancement branch, it specifically and accurately captures and reconstructs the high-frequency fine-grained harmonic structure most severely affected by helium distortion. The generator output is jointly optimized using the adversarial loss of multi-scale and multi-period discriminators, ultimately outputting a high-fidelity, clear speech time-domain waveform with a sampling rate of 16kHz. Figure 4 The short-time-window spectrogram shown shows a distorted envelope in the helium speech spectrum before reconstruction, with energy densely compressed in the 30-60 Hz band. A large area of ​​vacuum exists in the low-frequency fundamental region, and the formant trajectories are blurred and severely adhered. After algorithm processing, the distorted energy is re-expanded, and the low-order formants in the 0-30 Hz band are filled and reconstructed with high quality. In the reconstructed spectrum, the vertical stripes representing the vocal cord vibration period become extremely sharp and clear, the boundaries between voiced and unvoiced sounds are distinct, and the formant trajectories are smooth and continuous.

[0067] S4, Edge Engineering Deployment.

[0068] In the implementation phase of the example project, considering the extremely stringent computing power and power consumption limitations of the deep-sea saturation diving mother ship compartment and the edge terminal of the underwater diving bell / umbilical cable diving suit, as well as the extremely high real-time requirements for life support communication in underwater high-pressure operations, this solution abandons conventional global indiscriminate acceleration methods and designs a dedicated edge deployment and acceleration architecture for deep-sea saturation diving. Specific implementation includes: dedicated hybrid precision quantization based on helium speech high-frequency detail preservation. During model export and lightweighting, addressing the physical pain point that traditional full-network FP16 or INT8 quantization causes irreversible loss of the already weak high-frequency harmonics in helium (such as easily attenuated friction sounds and high-frequency resonant peaks) during numerical truncation, this invention proposes a helium speech high-frequency detail-preserving hybrid precision quantization strategy. When processing the computation graph using the inference engine, shallow network structures that handle low-frequency macroscopic components, the speaker's basic timbre, and conventional text features undergo high-compression, low-precision quantization (such as INT8) to free up GPU memory. However, critical network layers that handle high-frequency harmonic features most severely affected by high-voltage distortion (such as the high-frequency feature extraction layer in the ASR front-end and the He-MRF compensation module in the TTS generator) are locked to FP32 single-precision or high-precision floating-point inference. This significantly reduces the GPU memory usage of the diving bell edge terminal while preserving the fine-grained high-frequency reconstruction accuracy of helium speech from the underlying computing power allocation level. Adapting to the diver's vocal rhythm, a dedicated low-latency scheduling algorithm for full-duplex helium speech communication is proposed. Addressing the cascading latency issue in underwater ASR recognition to TTS reconstruction, this solution abandons the traditional fixed-time-window-based ring buffer mechanism and innovatively proposes a low-latency scheduling algorithm for full-duplex communication in a helium-oxygen environment. This mechanism introduces an adaptive dynamic flow buffer based on the acoustic energy envelope of helium: considering the high-frequency, short breathing rhythms generated when deep-sea divers wear high-pressure breathing masks, the monitoring terminal captures the dynamic envelope containing mechanical sounds from the breathing valve and abrupt changes in speech energy in real time. The ASR module forces the trough of this energy envelope (i.e., the high-resistance breathing pause point) as the natural block boundary and outputs semantic slices; the TTS module synchronously monitors this dynamic buffer, and once the semantic boundary marked by the breathing rhythm is triggered, it immediately seizes the NPU / DSP computing resources of the mother ship control console or the edge terminal of the diving suit to preemptively start acoustic feature generation and waveform reconstruction.

[0069] Through the aforementioned hardware and software co-optimization of hardware binding, physical fidelity quantization, and vocal rhythm scheduling, this system achieves low-latency, high-speed communication that meets the requirements of combat-ready full-duplex natural interaction between surface support personnel on the mother ship and underwater divers operating at depths of up to 100 meters.

[0070] This invention relates to the field of artificial intelligence and intelligent speech signal processing technology. Addressing the problems of nonlinear distortion, high-frequency harmonic attenuation, and deep-sea noise interference caused by helium-oxygen mixtures in deep-sea saturated diving operations, as well as the low recognition accuracy, poor clarity and naturalness of reconstructed speech, and insufficient generalization of conventional deep learning models, this invention provides a helium speech processing method based on cross-depth acoustic feature fusion and dynamic phoneme error correction. The method achieves end-to-end correction and reconstruction of helium speech through signal acquisition and preprocessing, speech recognition, speech synthesis, and edge-end engineering deployment. The preprocessing stage employs helium speech formant nonlinear compensation, dynamic frequency domain feature mapping, and adaptive CMVN to eliminate channel distortion. The speech recognition stage introduces a helium acoustic nonlinear inverse stretching operator, dispersive dynamic convolutional attention, environmental parameter-aware hybrid expert architecture, and physical rules. Driven by cluster search phoneme correction, the recognition robustness is improved. In the speech synthesis stage, fundamental frequency harmonic reconstruction is achieved through multi-scale wavelet transform, and an optimized HiFi-GAN vocoder and high-frequency harmonic enhancement branch are used to restore high-frequency details. In the edge deployment stage, a hybrid precision quantization and low-latency scheduling strategy is adopted to adapt to the computing power limitations of deep-sea terminals. This invention can reduce the helium speech word error rate to 8.2%, the naturalness MOS score of synthesized speech to 4.1 / 5, improve the high-frequency harmonic integrity by 40%, the end-to-end processing latency to less than 150ms, and reduce the model memory usage by 48%. It has the characteristics of accurate recognition, high-fidelity reconstruction, strong real-time performance, lightweight deployment, full-depth adaptability, and strong anti-interference ability. It can effectively solve the problem of helium speech communication distortion in deep-sea saturation diving operations and significantly improve the safety, smoothness, and reliability of underwater operation communication.

[0071] Although the present invention has been described in detail with reference to the accompanying drawings and preferred embodiments, the invention is not limited thereto. Various equivalent modifications or substitutions can be made to the embodiments of the invention by those skilled in the art without departing from the spirit and essence of the invention. Such modifications or substitutions should all fall within the scope of the invention, or any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the invention should be covered within the protection scope of the invention. Therefore, the protection scope of the invention should be determined by the scope of the claims.

Claims

1. A helium speech processing method for cross-depth acoustic feature fusion and dynamic phoneme error correction, characterized in that, The method includes the following steps: S1. Calculate a dedicated noise masking threshold for the collected helium speech to remove deep-sea noise. Use a first-order finite impulse response pre-emphasis filter to achieve nonlinear compensation of the formant peak. After zero-padding, perform a 512-point fast Fourier transform to obtain frequency domain features. Generate a helium speech distortion feature map based on the piecewise adaptive sensing weight function and the environmental sound speed variation ratio. Introduce a dynamic scaling factor driven by depth and helium-oxygen concentration to perform adaptive cepstral mean-variance normalization and output a two-dimensional time-frequency feature matrix. S2. Using the aforementioned two-dimensional time-frequency feature matrix as input, a helium environment characterization matrix is ​​constructed by combining the diving depth and helium-oxygen concentration. Helium speech neural inference and synthesis He-NIS operators are used for front-end feature extraction. Phase misalignment and high-frequency dispersion are compensated through helium speech audio dispersion dynamic convolutional attention He-DDCA. Cross-depth distortion compensation is achieved using an environmental parameter-aware hybrid expert architecture and a variational Bayesian adapter (VB-Adapter). After connectionist temporal classification (CTC) and attention encoder-decoder (AED) joint loss training and physical rule-constrained bundle search phoneme correction, the text is output. The helium environment characterization matrix is ​​constructed by combining the diving depth and helium-oxygen concentration, using He... The NIS operator completes the front-end feature extraction, specifically including: Depth of descent at the moment of articulation The percentage of helium concentration in the gas mixture Based on the ideal gas sound speed equation, a real-time deep-sea sound speed ratio calculation model was constructed: ; in, This represents the molar mass of the corresponding gas; Subsequently, the nonlinear frequency shift parameter caused by this sound velocity ratio was extracted and combined with the diving depth. Based on the water pressure attenuation coefficient of the corresponding environment, a physical prior representation matrix specific to helium speech is generated. ; To address the non-stationary distortion problem of severe high-frequency stretching and slight low-frequency drift of speech formants in a helium environment, a helium acoustic nonlinear inverse stretching He-NIS operator is proposed. This He-NIS operator introduces an elastic coordinate offset based on physical equations before performing inner product operations. The elastic coordinate offset Based on the aforementioned sound speed ratio constraint, the formula is as follows: ; in, This represents the frequency axis coordinate of the current sampling point. For hyperparameters, The weight matrix is ​​learnable; through this constraint, the He-NIS operator maintains regular sampling in the low-frequency region, while its sampling grid changes according to frequency in the high-frequency region. The rate exhibits an extreme, non-linear expansion; the eigenvalue of the current position after non-linear stretching and rearrangement is... Its calculation formula is: ; in, This represents the frequency axis coordinates of the current sampling point after nonlinear stretching, i.e., the position of the rearranged high-frequency expanded grid. This represents the coordinates of the nth sampling point on the original frequency axis before stretching; Through He DDCA (Dispersion Dynamic Convolution Attention Compensation) addresses phase misalignment and high-frequency dispersion, specifically including: Physical phase compensation mapping: As a constraint, dynamic deep convolution is used to map the input features, endowing the severely damaged high-frequency channels with a wide temporal receptive field to compensate for the phase lead of the sound waves in helium, thus obtaining a time-aligned signal. and ; Reconstructing the attention scoring space: For the first time in and Interspersed by Generated medium dispersion modulation tensor The reconstructed attention scoring formula is defined as follows: ; In the formula, The reconstructed attention scoring matrix; This is the query feature matrix aligned on the time axis after physical phase compensation mapping; This is the transpose of the key feature matrix aligned on the time axis after physical phase compensation mapping; This is the helium environment characterization matrix; For the reason Real-time generated medium dispersion modulation tensor; Key vector The size of the feature dimension; pass The scoring space is reconstructed into a physical Markov space that conforms to the acoustic attenuation law of helium, and physical penalties are imposed on the high-frequency channels that are prone to distortion, while high confidence is given to the stable low-frequency channels. Broadband Harmonic Reconstruction Branch: Generation Matrix-time parallel dynamic depthwise convolution branches, through Real-time generation of convolution kernel dilation rate Capture the high-frequency micro-vibration characteristics caused by helium gas dispersion. The output is: ; Solve the feature misalignment and collapse problem under fluid distortion environment, and output global acoustic features. ; Utilizing an environment parameter-aware hybrid expert architecture and VB The Adapter implements cross-depth distortion compensation, specifically including: Introducing an environment parameter-aware hybrid expert architecture, this architecture consists of one shared expert and... It consists of a variational specific expert branch tailored to the water depth intervals defined by the target water depth anchor points; use Real-time diving depth extracted from matrix Calculate the physical activation weights of each branch expert. The expression is: ; in, Let i be the target water depth anchor point corresponding to the i-th expert branch; The activated expert branch uses a variational adapter (VB-Adapter) structure to separate the distorted signal from the damaged signal: Posterior path of physical conditions: Combining The helium-oxygen ratio and water pressure parameters in the model were used to construct a conditional posterior distribution. It is responsible for capturing deterministic frequency shift compensation in helium speech; Acoustic variational distribution path: generating observational variational distribution The technique of reparameterization is used to capture unsteady residual fluctuations caused by individual differences among divers or complex echoes. In the output phase of the expert branch, a physical weight adaptive fusion mechanism is adopted; the final acoustic feature tensor Generated using the following weighted logic: ; The basic semantic backbone is shared by experts and serves as a global reference; the output of each VB-Adapter is determined by its corresponding physical activation weights. Scaling is performed; within the expert group, this is achieved through learnable parameters. Dynamic adjustment of deterministic compensation amount With stochastic adjustment The proportion; in environments with severe depth fluctuations, the system will automatically increase the learnable parameters of the physical condition path. This is to strengthen the constraints on feature reconstruction; S3. Based on the text, construct a helium high-risk phoneme degradation mapping dictionary and incorporate prior knowledge of physical conditions. Decouple the fundamental frequency through continuous wavelet transform at multiple scales. Complete the multi-scale reconstruction of the fundamental frequency and harmonics according to the sound speed variation ratio. Input the prosodic parameters and text features into a non-autoregressive decoder to obtain the Mel spectrum. S4. Input the Mel spectrum into the optimized high-fidelity generative adversarial network HiFi-GAN vocoder of the multi-receptive field fusion MRF module with built-in high-frequency harmonic enhancement branch, and output high-fidelity speech through adversarial optimization. Low-latency deployment at the edge is achieved through high-frequency preserved mixed precision quantization and adaptive streaming scheduling.

2. The method according to claim 1, characterized in that, The S1 uses a first-order finite impulse response pre-emphasis filter to achieve nonlinear compensation of the resonance peak, including: A first-order finite impulse response pre-emphasis filter is applied to the effective speech segment after removing deep-sea noise. The time-domain difference equation of this filter is: In the formula It is a dynamic compensation coefficient that is adaptively adjusted according to the helium-oxygen mixing ratio at the current diving depth to correct the high-frequency distortion caused by the nonlinear upward shift of the resonance peak at different depths.

3. The method according to claim 1, characterized in that, In step S1, after zero-padding, a 512-point Fast Fourier Transform is performed to obtain frequency domain features. Based on a piecewise adaptive sensing weight function and the environmental sound speed variation ratio, a helium speech distortion feature map is generated, including: Perform a 512-point Fast Fourier Transform on each frame of the signal to convert the time-domain signal into a frequency-domain complex sequence. Subsequently, the discrete power spectral density was calculated. The construction length is One-dimensional perceptual weight vector ,in This represents the total number of filters. The length of a single frame signal; the first The weights of each filter Based on its center frequency The frequency band in which it is located is calculated using the following piecewise adaptive sensing weight function: ; in: The preset cutoff frequency threshold for deep-sea background mechanical noise; and These are the low-frequency noise attenuation coefficient and the smoothing index, respectively. The high-frequency enhancement coefficient for helium speech; The environmental sound speed variation ratio is defined as follows: , Current diving depth Helium-oxygen mixing ratio The speed of sound in the mixed gas below The speed of sound in air at normal pressure; then the power spectral density... Mapped to a scale filter bank and combined with the generated perceptual weight matrix Feature weighting fusion is performed to generate a helium speech distortion feature map.

4. The method according to claim 3, characterized in that, The S2 process, involving bundle search phoneme correction trained with joint CTC and AED loss and constrained by physical rules, outputs text, specifically including: Model training phase: Joint training is adopted using connection-to-sequence classification (CTC) and sequence-to-sequence attention decoding (AED), and a multi-task loss function is designed: ; in, The model's ability to recognize helium speech aberration phonemes is improved through dual-loss collaborative training, setting up hyperparameters. Phoneme correction and text decoding stage: Addressing the nonlinear distortion of high-frequency formants caused by changes in sound velocity under helium conditions, the most severely affected fricatives and affricates are identified. By statistically analyzing misclassified phoneme pairs in actual helium speech datasets, a priori transition matrix for high-risk deformed phonemes is established. Elements in this matrix represent the prior probability of phonemes being misidentified due to acoustic distortion at different helium-oxygen mixing ratios and diving depths. Based on this matrix, the candidate phoneme sequences output by the bundle search are corrected. Decoding is then performed. In the decoding stage, the original score of the search path is... When the candidate search path expands to a high-risk deformable phoneme and the acoustic confidence is low, a dynamic probability penalty mechanism is triggered; the current candidate phoneme is determined through set inclusion operations. Does it belong to the embedded high-risk deformed phoneme hash table? And combined with information entropy based on the probability distribution of the current time step Adaptive confidence threshold Make a judgment; when And the back-delay probability output by the acoustic model At that time, calculate the corrected path score. The calculation formula is as follows: ; in, The weighting coefficient for penalties for misjudgment; This represents the confidence score of the current acoustic model for the candidate phonemes; For indicator functions; This is the compensation coefficient for the correct path.

5. The method according to claim 4, characterized in that, The multi-scale decoupling of the fundamental frequency through continuous wavelet transform in S3 specifically includes: Using the Mexican hat wavelet as the generating function Calculate the wavelet coefficients: ; frequency trajectory Decomposed into wavelet coefficients of multiple scales , scale factor Corresponding to the acoustic wavelength compression ratio of a helium-oxygen mixture.

6. The method according to claim 5, characterized in that, The multi-scale reconstruction of fundamental frequency and harmonics based on the sound speed variation ratio in S3 specifically includes: To eliminate distortion caused by helium and preserve timbre, a dynamic scale determination threshold controlled by the variation ratio of ambient sound speed is set. The threshold satisfies the function ; During frequency band reconstruction: Scale factor The low-frequency coefficients are preserved without loss as actual physiological vocal cord vibration components; The high-frequency coefficients are used as nonlinear resonance distortion components of the helium cavity, and a damping smoothing penalty that is positively correlated with the sound speed variation ratio is applied.

7. The method according to claim 6, characterized in that, S4 specifically includes: inputting the Mel spectrum into an optimized HiFi-GAN vocoder with a built-in high-frequency harmonic enhancement branch MRF module; outputting high-fidelity speech through multi-receptive field fusion and adversarial optimization of the vocoder; simultaneously employing a high-frequency preserved hybrid precision quantization strategy to perform low-precision quantization on the shallow network structure that processes low-frequency macroscopic components, the speaker's basic timbre, and conventional text features to release video memory, while locking high-precision floating-point inference on the key network layer that processes high-frequency harmonic features to preserve high-frequency details; and introducing an adaptive dynamic streaming scheduling mechanism based on helium acoustic energy envelope to capture the diver's breathing rhythm as the semantic block boundary, achieving low-latency processing by preempting computing resources, and finally completing the low-latency deployment of the deep-sea edge terminal to meet the underwater full-duplex real-time communication requirements.

Citation Information

Patent Citations

  • Intelligent ultrasonic bird repelling device and method based on deep learning bird sound recognition

    CN118525835A

  • Audio depth forgery detection method fusing multi-source features and cross-scale modeling

    CN120783799A