A method of mixed voice processing, an electronic device, a computer readable medium
By combining biomimetic frequency domain analysis, time-frequency joint deconvolution, and dynamic frequency compensation filters, the channel distortion problem in speech processing under complex environments was solved, achieving high-quality speech signal separation and optimization.
Patent Information
- Application Number
- CN202510516957.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing speech processing technologies struggle to adapt to environmental changes in complex acoustic environments, resulting in low quality of mixed speech output, especially under channel distortion caused by multipath propagation.
By collecting mixed speech and environmental impact parameters for biomimetic frequency domain analysis, channel distortion data is generated. Direct and reflected sound components are separated using time-frequency joint deconvolution processing. A dynamic frequency compensation filter is constructed for energy redistribution and Doppler frequency shift correction is performed. Finally, real-time optimization is carried out in a lightweight edge computing architecture.
It significantly improves the clarity and intelligibility of speech signals, reduces signal distortion caused by environmental noise, and achieves efficient real-time speech optimization in complex environments.
Smart Images

Figure CN120236599B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing, and in particular to a hybrid speech processing method, an electronic device and a computer readable medium. BACKGROUND
[0002] Early speech processing relies on linear prediction analysis (LPC), mel frequency cepstral coefficient (MFCC) and other manual feature extraction techniques, combined with hidden Markov model (HMM) or Gaussian mixture model (GMM) to realize speech recognition and analysis. However, these methods have limited performance in processing complex speech environments such as noise and accent changes. With the improvement of computing power and data accumulation, deep neural networks (DNN), convolutional neural networks (CNN) and recurrent neural networks (RNN) have been gradually introduced into the field of speech processing, improving the robustness of speech recognition, synthesis and enhancement. In recent years, models based on the Transformer architecture (such as Wav2Vec and HuBERT) have made breakthroughs in speech processing tasks, further improving the accuracy of speech understanding and generation. However, in complex acoustic environments, speech signals often encounter channel distortion caused by multi-path propagation, and existing systems often rely on static filters, making it difficult to adapt to environmental changes, thus resulting in low quality of hybrid speech output. SUMMARY
[0003] Therefore, it is necessary to provide a hybrid speech processing method, an electronic device and a computer readable medium to solve at least one of the above technical problems.
[0004] To achieve the above-mentioned purpose, a hybrid speech processing method, the method comprising the following steps:
[0005] Step S1: collecting a hybrid speech and an environmental influence parameter; performing bionics frequency domain analysis on the hybrid speech to obtain low-frequency attenuation compensation feature data; performing multi-path effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental influence parameter to generate channel distortion data;
[0006] Step S2: performing time-frequency domain joint deconvolution processing on the hybrid speech using the channel distortion data to generate a direct sound component and a reflected sound component; performing adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data;
[0007] Step S3: constructing a dynamic frequency compensation filter based on a preset environmental acoustic feature; performing energy redistribution on the anti-multipath speech enhancement data using the dynamic frequency compensation filter to generate a hybrid compensation speech; performing Doppler frequency shift correction on the hybrid compensation speech to generate a pure speech fundamental signal;
[0008] Step S4: signal trend analysis is performed on the pure speech fundamental frequency signal to generate a fundamental frequency signal trend curve; and the dynamic frequency compensation filter is deployed in a lightweight edge computing architecture based on the fundamental frequency signal trend curve to perform real-time mixed speech optimization.
[0009] The present application can extract low-frequency attenuation characteristics through the collection of mixed speech and environmental impact parameters and the analysis of bionics frequency domain, and generate channel distortion data through multi-path effect propagation analysis, which lays a foundation for subsequent signal processing, especially the identification and compensation of multi-path effects, reducing signal distortion caused by environmental noise. Time-frequency joint deconvolution using channel distortion data can effectively separate direct sound and reflected sound components, providing clear signal sources for subsequent enhancement processing. This separation process is the basis of anti-multipath training, which can significantly improve the intelligibility of speech and enhance the intelligibility of environmental speech. By constructing a dynamic frequency compensation filter and redistributing the energy of anti-multipath speech enhancement data, it helps to optimize the frequency characteristics of the signal and reduce the frequency shift caused by the environment. Doppler shift correction further improves the accuracy of environmental speech signals, ensuring that the generated pure speech signal is closer to the actual speech source. The variation trend analysis of the pure speech fundamental frequency signal can generate a fundamental frequency signal trend curve, which provides a basis for the lightweight edge computing deployment of the dynamic frequency compensation filter. This process makes real-time mixed speech optimization more efficient and accurate, reduces the computational burden, and optimizes the ability of real-time processing. Therefore, the present application improves the output quality of mixed speech through multi-stage signal processing, frequency compensation and real-time optimization technology.
[0010] Preferably, step S1 comprises the following steps:
[0011] Step S11: collect mixed speech and environmental impact parameters;
[0012] Step S12: signal regularization is performed on the mixed speech to generate mixed frame-level speech;
[0013] Step S13: bionic filter analysis is performed on the mixed frame-level speech to generate bionic spectrum data; the frequency attenuation characteristics in the bionic spectrum data are analyzed, and the bionic spectrum data are inverse filtered and energy reconstructed to obtain low-frequency attenuation compensation characteristic data;
[0014] Step S14: propagation characteristics in the environmental impact parameters are extracted to obtain propagation environment characteristic data; beam tracking path simulation is performed on the propagation environment characteristic data to generate path simulation data;
[0015] Step S15: multi-path effect propagation modeling is performed on the low-frequency attenuation compensation characteristic data based on the path simulation data to generate multi-path coupling characteristic data; channel estimation is performed on the multi-path coupling characteristic data to generate channel distortion data.
[0016] The present application provides high-quality raw data for subsequent analysis and optimization by accurately collecting mixed speech signals and environmental impact parameters. This step ensures the accuracy of signal analysis, especially considering the complexity of special environments, and collecting comprehensive data is the basis for improving system performance. By signal regularization of mixed speech, generating mixed frame-level speech, the original speech signal can be converted into a standardized format, which helps subsequent analysis and processing. This step of regularization reduces the stray noise and irregularity in the signal, ensuring signal consistency and operability. Bionic filtering analysis of mixed speech and generation of bionic spectrum data can simulate the actual sound propagation characteristics in the environment. Through inverse filtering and energy reconstruction, low-frequency attenuation compensation feature data is effectively generated, which effectively restores the low-frequency loss caused by the environment and improves the intelligibility and quality of the speech signal. Extracting propagation characteristics (such as sound speed profile fitting and salinity-temperature-depth structure quantization) in environmental impact parameters and performing beam tracking path simulation can accurately understand the propagation characteristics of the environment, which has an important influence on the propagation behavior of the signal, helping subsequent processing to better adapt to the actual environment and further improve the transmission quality of the signal. By modeling the multipath effect propagation of path simulation data, the multipath effect in environmental signal transmission can be fully simulated and effectively compensated. At the same time, the channel estimation step can accurately generate channel distortion data, providing strong support for subsequent speech enhancement and recovery. This step can significantly reduce signal distortion caused by environmental multipath effects and improve speech intelligibility.
[0017] Preferably, the channel estimation of the multipath coupling feature data in step S15 comprises:
[0018] analyzing the path delay of the multipath coupling feature data;
[0019] performing gain inversion on the multipath coupling feature data according to the path delay to generate path gain data;
[0020] performing channel impulse response analysis on the path gain data to generate multipath impulse response data;
[0021] performing time-varying modeling on the multipath impulse response data to generate time-varying channel response data;
[0022] performing structure regularization and distortion measurement analysis on the time-varying channel response data to extract frequency response envelope, phase shift, and asymmetric distortion indicators, and generating channel distortion data.
[0023] The application can help to accurately understand the propagation time difference of different signal paths by analyzing the path delay. In the environment, different reflection paths and delays will be caused by multi-path signals. By accurately analyzing the path delay, the foundation for subsequent gain inversion and multi-path modeling can be provided to ensure the accuracy and clarity of signal recovery. According to the path delay, the gain inversion is performed on the multi-path coupling characteristic data to generate path gain data. Gain inversion can effectively compensate for signal attenuation and uneven gain caused by multi-path effects, ensuring that the signal intensity on different paths is properly adjusted and optimized, which provides more accurate signal characteristics for subsequent impulse response analysis and time-varying modeling, enhancing the reliability of the signal. Through channel impulse response analysis on path gain data, the impulse response characteristics in multi-path effects can be accurately obtained, and the generated multi-path impulse response data reflects the propagation characteristics of signals on each path, which helps to understand and compensate for signal distortion, especially the difference between reflection paths and direct paths. This step can effectively recover the characteristics of signals in complex environments, thereby improving the quality of speech. The time-varying characteristics of the multi-path channel will cause the signal to change over time. Time-varying modeling of the multi-path impulse response data can accurately reflect the dynamic changes of the channel at different times, generating time-varying channel response data. This modeling method can help to cope with the rapid changes of signals in the environment and improve the adaptability of the system, ensuring stable transmission of signals in dynamic environments. Through structure regularization and distortion measurement analysis of time-varying channel response data, frequency response envelope, phase offset and asymmetric distortion indicators can be extracted to further reveal the details of channel distortion. By accurately analyzing these distortion indicators, the distortion in the signal caused by multi-path effects, frequency offset and other factors can be effectively identified and compensated, ensuring that the finally generated channel distortion data has high accuracy and reliability. This step of optimization greatly improves the quality of speech signals and reduces noise and distortion in speech.
[0024] Preferably, step S2 comprises the following steps:
[0025] Step S21: performing time-domain deconvolution on the mixed speech and channel distortion data to generate initial direct sound waveform data;
[0026] Step S22: performing short-time Fourier transform on the mixed speech to generate frequency domain reflection sound estimation data;
[0027] Step S23: performing inter-frame feature alignment on the initial direct sound waveform data and the frequency domain reflection sound estimation data to generate joint acoustic separation feature data; performing time-frequency decomposition on the joint acoustic separation feature data to extract direct sound principal components and reflection sound multi-path interference components to generate direct sound components and reflection sound components;
[0028] Step S24: performing adversarial training based on the direct sound components and the reflection sound components to generate anti-multipath speech enhancement data.
[0029] The present application can effectively remove the influence of channel distortion by time domain deconvolution of mixed voice and channel distortion data, and restore the initial waveform data of direct sound. The deconvolution process reduces the influence of multipath effect and attenuation in the channel on the signal, providing purer direct sound data for subsequent acoustic analysis. By performing short-time Fourier transform on the mixed voice, the voice signal is converted into frequency domain representation, making the reflected sound components in the frequency domain more obvious. Frequency domain reflection sound estimation can reveal the reflection path in the environment and accurately estimate the reflected sound, providing reliable frequency domain feature data for subsequent acoustic separation. Aligning the initial direct sound waveform data with the frequency domain reflection sound estimation data can achieve more accurate acoustic separation. Through time-frequency decomposition, the principal component of direct sound and the multipath interference component of reflected sound can be extracted from the joint acoustic separation feature data. This decomposition method can analyze in both time and frequency domains, effectively reducing the interference of reflected sound and extracting clean direct sound and reflected sound components. Through the adversarial training based on direct sound and reflected sound components, the system can perform efficient voice enhancement. The adversarial training process makes the model more robust when dealing with multipath effect, effectively reducing the interference of reflected sound on direct sound, thereby generating multipath-resistant voice enhancement data. This process makes the final voice signal clearer and more understandable, and significantly improves the quality of voice communication even in complex environments.
[0030] Preferably, step S24 comprises the following steps:
[0031] Step S241: constructing an adversarial network for direct sound components and reflected sound components, building a separator-discriminator adversarial structure, and generating an adversarial training feature pair;
[0032] Step S242: performing feature residual guided training on the adversarial training feature pair to generate multipath-resistant discriminant feature data;
[0033] Step S243: performing multi-scale enhancement mapping on the multipath-resistant discriminant feature data to generate multipath-resistant voice enhancement data.
[0034] The application effectively separates the direct sound and reflected sound in the signal by constructing an adversarial network for the direct sound component and the reflected sound component, using the adversarial structure of the separator-discriminator. The key of the adversarial structure is to evaluate the output of the separator by the discriminator, forcing the separator to optimize its separation effect in order to generate clearer direct sound and reflected sound. Through this process, the generated adversarial training feature pair can promote self-learning of the model, thereby continuously improving the accuracy of signal separation. Feature residual guided training on the adversarial training feature pair can further optimize the separation effect of direct sound and reflected sound. Feature residual guided training enables the model to focus on processing residual features, thereby reducing the influence of multipath effect and enhancing the removal effect of reflected sound. The generated anti-multipath discrimination feature data can not only effectively distinguish direct sound and reflected sound, but also capture the detailed features that are difficult to handle in the multipath effect, thereby improving the clarity of the final speech. Multi-scale enhancement mapping of the anti-multipath discrimination feature data can process the multi-scale characteristics in the signal, further enhancing the quality of the signal. Multi-scale enhancement mapping can optimize the signal in different frequency and time scales, especially in complex environments, multi-scale mapping can effectively reduce the interference in different scales, especially the difference between low frequency and high frequency, enhancing the naturalness and audibility of the final speech. Through this enhancement, the generated anti-multipath speech enhancement data can significantly improve the clarity and understanding of environmental speech communication.
[0035] Preferably, the energy redistribution of the anti-multipath speech enhancement data by the dynamic frequency compensation filter in step S3 comprises:
[0036] Converting the anti-multipath speech enhancement data by the dynamic frequency compensation filter into a multi-channel time-frequency energy graph to generate a multi-channel speech energy graph;
[0037] Performing speech frequency response difference analysis on the multi-channel speech energy graph to generate frequency energy distribution data;
[0038] Calculating the frequency band weight of the frequency energy distribution data to generate a frequency compensation weight;
[0039] Performing frequency domain filtering reconstruction on the anti-multipath speech enhancement data according to the frequency compensation weight to generate frequency energy redistribution speech data;
[0040] Performing time-domain waveform reconstruction on the frequency energy redistribution speech data to generate hybrid compensation speech.
[0041] The application can convert the speech signal into a multi-channel energy representation by converting the multi-path speech enhancement data through a dynamic frequency compensation filter, which facilitates the analysis of the energy distribution of different frequencies and time domains, and provides rich frequency domain and time domain features for subsequent analysis, which helps to identify the influence of the multi-path effect in the environment on the speech signal. By analyzing the speech frequency response difference of the multi-channel speech energy graph, the energy difference between different frequency bands can be revealed, especially the frequency response difference between the reflected sound and the direct sound. This analysis can help identify which frequency bands are strongly affected by the environmental multi-path effect, providing an important basis for subsequent frequency compensation. The generated frequency energy distribution data can accurately describe the frequency response difference, providing necessary data support for filter design. According to the frequency band weight calculated based on the frequency energy distribution data, different compensation weights can be assigned to each frequency band. Through this process, it can more accurately identify which frequency bands need stronger compensation, thereby achieving more efficient frequency compensation. The frequency compensation weight generated by the filter design provides adaptive compensation parameters, making the filter more accurate and flexible when redistributing energy. By performing frequency domain filtering reconstruction on the multi-path speech enhancement data according to the frequency compensation weight, different frequency signals can be effectively compensated, reducing the frequency attenuation and interference caused by the multi-path effect. Frequency energy redistribution optimizes the speech signal in the frequency domain, restoring the attenuated or distorted parts and enhancing the clarity and quality of the speech signal. By reconstructing the time-domain waveform of the frequency energy redistribution speech data, the signal optimized in the frequency domain can be restored to a time-domain waveform, completing the final speech compensation process. The mixed compensation speech presents a more clear and coherent speech waveform in the time domain, greatly improving the audibility and naturalness of the speech. This process makes the final speech not only retain the clarity of the direct sound, but also effectively reduce the influence of reflected sound and multi-path effect.
[0042] Preferably, the step S3 includes:
[0043] The mixed compensation speech is subjected to biological frequency separation to separate the non-biological frequencies therefrom, generating biological separation speech;
[0044] Based on the biological separation speech, sound source positioning is performed to obtain biological motion trajectory data;
[0045] The relative speed and sound wave propagation direction of the biological motion trajectory data are analyzed, and the mixed compensation speech is subjected to Doppler frequency shift parameter calculation to obtain the Doppler frequency shift parameter;
[0046] The mixed compensation speech is subjected to nonlinear phase correction through the Doppler frequency shift parameter, generating a corrected frequency domain signal;
[0047] The corrected frequency domain signal is subjected to spectral analysis, and the fundamental frequency is extracted to generate a clean speech fundamental frequency signal. This invention effectively separates biological frequencies (e.g., whale or fish sounds) from non-biological frequencies (e.g., reflected sound or background noise) in the speech signal by performing biological frequency separation on the mixed compensated speech. This process helps reduce the interference of biological frequencies on the speech signal, making the remaining speech signal cleaner and clearer. The generated biologically separated speech provides a clean signal source for subsequent localization and correction processes. Sound source localization based on biologically separated speech can accurately obtain the motion trajectory data of environmental organisms, which is crucial for Doppler frequency shift correction because the Doppler frequency shift effect is closely related to the relative velocity of the organism and the direction of sound wave propagation. Accurate motion trajectory data allows for accurate calculation of the organism's motion state, providing necessary information for subsequent frequency shift parameter calculations. Analyzing the relative velocity and sound wave propagation direction in the biological motion trajectory data helps calculate Doppler frequency shift parameters, which directly affect the frequency and phase of the environmental speech signal, especially in the presence of biological motion. By accurately calculating the Doppler frequency shift parameter, a precise reference can be provided for subsequent nonlinear phase correction, thereby reducing distortion caused by the frequency shift and restoring the original frequency characteristics of the signal. Using the Doppler frequency shift parameter for nonlinear phase correction of hybrid compensated speech can eliminate phase shifts caused by biological motion and restore the frequency domain structure of the signal. Nonlinear phase correction can effectively address the Doppler effect in complex environments, reducing speech signal blurring or distortion caused by phase distortion, thus improving speech clarity and intelligibility. By performing spectral analysis on the corrected frequency domain signal and extracting the fundamental frequency, a clean speech fundamental frequency signal can be obtained. The fundamental frequency is the most critical part of the speech signal, representing the main pitch and tone of the speech. Extracting the clean fundamental frequency signal helps restore the natural timbre of the speech and ensures the audibility and accuracy of speech in the environment. The generated clean speech fundamental frequency signal provides a high-quality audio foundation for subsequent speech reconstruction and optimization.
[0048] Preferably, step S4 includes the following steps:
[0049] Step S41: Analyze the local frequencies of the pure speech fundamental frequency signal, extract the trend features of local frequency changes, and generate fundamental frequency signal change trend data;
[0050] Step S42: Perform sliding window analysis on the fundamental frequency signal variation trend data to smooth out short-term fluctuations in the signal and generate a stable fundamental frequency signal variation trend curve;
[0051] Step S43: Based on the light edge computing architecture, the smooth base frequency signal change trend curve is used to deploy the dynamic frequency compensation filter into the preset edge computing node for real-time calculation and frequency compensation, to generate an edge computing optimization filter, so as to perform the quality optimization work of the mixed voice.
[0052] The application can reveal the law of frequency change in the environmental voice signal by analyzing the local frequency of the pure voice base frequency signal and extracting the change trend characteristics, which provides an important basis for subsequent filtering and optimization, helps to identify the periodic change, frequency drift or other instability factors in the signal, and the generated base frequency signal change trend data provides strong support for subsequent smoothing processing and dynamic frequency compensation. The sliding window analysis can smooth out short-term fluctuations in the signal and eliminate transient changes caused by noise or environmental factors, so that the base frequency signal change trend is more stable, and this process can remove unnecessary noise and transient anomalies, so that the subsequent filtering and optimization work is more accurate and reliable, and the generated smooth base frequency signal change trend curve provides stable reference data for the deployment of the dynamic frequency compensation filter, thereby improving the accuracy and effect of optimization. The smooth base frequency signal change trend curve is transmitted to the preset edge computing node for real-time calculation and frequency compensation by using the light edge computing architecture, which can realize efficient real-time processing, and this distributed computing method can reduce the burden of the central server, reduce the delay, and ensure that the voice optimization process can quickly respond in the real-time environment, and the generated edge computing optimization filter can dynamically adjust the frequency compensation parameters, so as to realize the real-time optimization of the mixed voice, and improve the intelligibility, sound quality and intelligibility of the voice.
[0053] The application provides an electronic device, which comprises:
[0054] at least one processor;
[0055] a memory connected in communication with the at least one processor;
[0056] wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method for processing mixed voice as described above.
[0057] The application further provides a computer readable medium having a computer program stored thereon, and the computer program is executed by a processor to implement the method for processing mixed voice as described above.
[0058] The beneficial effects of this invention lie in its ability to accurately identify the low-frequency attenuation characteristics of environmental signals by collecting mixed speech and environmental influence parameters and performing biomimetic frequency domain analysis, providing fundamental data for subsequent low-frequency compensation. The generation of low-frequency attenuation compensation characteristic data enables targeted attenuation compensation of environmental speech signals during signal processing, improving speech clarity and intelligibility. Through joint time-frequency domain deconvolution processing, direct and reflected sounds in mixed speech can be effectively separated, providing more accurate sound components. This provides a clean signal source for subsequent speech enhancement and effectively suppresses multipath interference caused by reflected sounds, improving the quality of the speech signal. After constructing a dynamic frequency compensation filter based on environmental acoustic characteristics, it can achieve energy redistribution against multipath speech enhancement data, optimizing the frequency response of environmental speech and making the speech signal clearer and more stable in the environment. Furthermore, Doppler frequency shift correction ensures that the fundamental frequency of the signal is not affected by frequency shifts caused by environmental movement, thereby generating a purer fundamental frequency speech signal. By analyzing the changing trends of the pure speech fundamental frequency signal, valuable feature data can be extracted. This data provides a basis for stabilizing the fundamental frequency signal, thereby reducing instability caused by noise and fluctuations. By deploying a dynamic frequency compensation filter in a lightweight edge computing architecture, real-time processing and frequency compensation are achieved, significantly reducing latency and enhancing the real-time performance and efficiency of environmental speech optimization, ensuring rapid response in speech optimization operations. Therefore, this invention improves the output quality of mixed speech through multi-stage signal processing, frequency compensation, and real-time optimization techniques. Attached Figure Description
[0059] Figure 1 A flowchart illustrating the steps of a hybrid speech processing method;
[0060] Figure 2 for Figure 1 A detailed flowchart illustrating the implementation steps of step S2.
[0061] Figure 3 for Figure 1 A detailed flowchart illustrating the implementation steps of step S4.
[0062] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0063] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0064] Further, the accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application. In the drawings:
[0065] It is to be understood that, although terms such as "first", "second", and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0066] To achieve the above object, there is provided Figures 1 to 3 A method of mixed speech processing, the method comprising the steps of:
[0067] Step S1: collecting mixed speech and environmental influence parameters; performing bionics frequency domain analysis on the mixed speech to obtain low-frequency attenuation compensation feature data; performing multi-path effect propagation analysis on the low-frequency attenuation compensation feature data by the environmental influence parameters to generate channel distortion data;
[0068] Step S2: performing time-frequency domain joint deconvolution processing on the mixed speech by using the channel distortion data to generate direct sound components and reflected sound components; performing adversarial training based on the direct sound components and the reflected sound components to generate anti-multipath speech enhancement data;
[0069] Step S3: constructing a dynamic frequency compensation filter based on preset environmental acoustic features; performing energy redistribution on the anti-multipath speech enhancement data by using the dynamic frequency compensation filter to generate mixed compensation speech; performing Doppler frequency shift correction on the mixed compensation speech to generate a pure speech fundamental signal;
[0070] Step S4: performing signal change trend analysis on the pure speech fundamental signal to generate a fundamental signal change trend curve; deploying a lightweight edge computing architecture of the dynamic frequency compensation filter by using the fundamental signal change trend curve to perform real-time mixed speech optimization tasks.
[0071] The application can extract low-frequency attenuation characteristics by collecting mixed voice and environmental influence parameters and carrying out bionics frequency domain analysis, and generate channel distortion data through multipath effect propagation analysis, which lays a foundation for subsequent signal processing, especially the identification and compensation of multipath effect, reduces signal distortion caused by environmental noise. Time-frequency joint deconvolution using channel distortion data can effectively separate direct sound and reflected sound components, providing clear signal sources for subsequent enhancement processing. This separation process is the basis of anti-multipath training, which can significantly improve the intelligibility of the voice and enhance the intelligibility of the environmental voice. By constructing a dynamic frequency compensation filter and redistributing the energy of the anti-multipath voice enhancement data, it helps to optimize the frequency characteristics of the signal and reduce the frequency shift caused by the environment. Doppler shift correction further improves the accuracy of the environmental voice signal, ensuring that the generated pure voice signal is closer to the actual voice source. Trend analysis of the pure voice fundamental frequency signal can generate a fundamental frequency signal trend curve, which provides a basis for the lightweight edge computing deployment of the dynamic frequency compensation filter. This process makes real-time mixed voice optimization more efficient and accurate, reduces the computational burden, and optimizes the ability of real-time processing. Therefore, the application improves the output quality of mixed voice through multi-stage signal processing, frequency compensation and real-time optimization technology.
[0072] In the embodiment of the application, as shown in the reference Figure 1 The method for processing mixed voice includes the following steps:
[0073] Step S1: Collect mixed voice and environmental influence parameters; carry out bionics frequency domain analysis on the mixed voice to obtain low-frequency attenuation compensation feature data; and carry out multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental influence parameters to generate channel distortion data;
[0074] In the embodiment of the present application, an environmental high-sensitivity acoustic sensor array (such as a hydrophone array) is used to collect audio data containing mixed signals of speech, background noise, echo, etc. in the target area. The sampling frequency of the mixed speech signal is generally set to 48 kHz to ensure complete coverage of human speech and low-frequency signals. The environmental detector is used to obtain key physical parameters such as temperature and humidity of the environment, which will be used to model the environmental sound propagation characteristics, including sound speed distribution, refractive index and multipath propagation structure. The processing mechanism of the human auditory system is simulated to perform bionic filtering and frequency domain envelope extraction on the mixed speech signal. The Gammatone filter bank can be used to simulate the cochlea filtering process and extract multi-band envelope features. The low-frequency region (generally <300Hz) is enhanced to compensate for the loss of low-frequency signal energy caused by environmental propagation. In each filter channel, the short-time energy, signal-to-noise ratio, frequency attenuation trend and other features are calculated, and the low-frequency components are dynamically compensated using an adaptive gain function. The output compensation result is the low-frequency attenuation compensation feature data, which is used for subsequent channel model adjustment. Based on the physical parameters in the environment, a sound speed profile (SSP) model is constructed. The Bellhop model or wave equation simulation method is used to model the propagation of sound waves in different paths, considering reflection, refraction, diffraction and other phenomena. The low-frequency attenuation compensation feature data is input into the above multi-path propagation model to simulate its propagation changes in the current environment, and time delay spread, frequency drift, amplitude attenuation and other distortion phenomena are obtained. The amplitude-frequency response function of the above propagation path is convolved with the initial signal to simulate the actual propagation result of the signal in the environment, and the output is the channel distortion data. The channel distortion data contains: multi-path response function, path time delay distribution, frequency response distortion function, etc., which is used for environmental sound channel equalization or inversion reconstruction.
[0075] Step S2: performing time-frequency joint deconvolution processing on the mixed speech using the channel distortion data to generate a direct sound component and a reflected sound component; performing adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data;
[0076] In the embodiment of the application, the channel distortion data obtained in step S1 is represented as a multipath impulse response function h(t, f), which contains path time delay, amplitude attenuation and frequency distortion components. The mixed speech signal x(t) is modeled as: x(t) = s(t) h(t) + n(t); where s(t) is the original speech signal at time point t, h(t) is the channel response at time point t, and n(t) is the background noise at time point t. The mixed speech signal is deconvolved in the time domain using a deconvolution filter (such as a Wiener filter or an LMS deconvolutioner), which preliminarily restores the speech signal on the direct path. The mixed speech is subjected to short-time Fourier transform (STFT) to extract a time-frequency spectrogram. In the frequency domain, an inverse filtering method is applied to eliminate the channel frequency response H(f) to recover high-frequency details and suppress distortion components. In the deconvolution output, based on the path delay and channel energy distribution, a time window function is used to divide the deconvolution result into: a direct path signal component (Direct Path Signal): the component with the shortest propagation delay, the strongest energy and the most concentrated spectrum. A reflected path signal component (Reflected Path Signal): a multi-path component with longer propagation delay, spectrum spreading and overlapping interference. A generative adversarial network (GAN) structure is constructed: the generator takes the direct path signal component and the reflected path signal component as input, learns the speech enhancement mapping, and outputs the multi-path resistant speech enhancement result. The discriminator discriminates whether the generated enhanced speech is close to the direct path original feature, to ensure that the enhanced speech has clarity and time-frequency structure integrity. The training data set includes real direct speech samples, artificially simulated reflected signals, and component pairs after joint deconvolution. The loss function is designed: adversarial loss: to constrain the authenticity of the generated speech. Perception loss: to ensure speech clarity and naturalness. Multi-path residual constraint: to minimize the interference of reflected residual on speech quality. The comprehensive loss function is as follows: L total = λ1L adv + λ2L perceptual + λ3L residual ; where L adv is the adversarial loss coefficient, L perceptual is the perception loss coefficient, and L residual is the multi-path residual constraint coefficient, λ1, λ2, and λ3 are loss weights. The trained generator can receive separated direct sound and reflected sound in real time and output high-quality speech that has removed the effects of multi-path, i.e., multi-path resistant speech enhancement data.
[0077] Step S3: constructing a dynamic frequency compensation filter based on the preset environmental acoustic characteristics; using the dynamic frequency compensation filter to perform energy redistribution on the multi-path resistant speech enhancement data to generate mixed compensation speech; performing Doppler frequency shift correction on the mixed compensation speech to generate a pure speech fundamental signal;
[0078] In the embodiment of the present application, an environmental acoustic feature database is established, and the features include the basic frequency range of whale sound, frequency modulation mode, spectral energy distribution, harmonic structure, and sound duration parameters. By analyzing various acoustic samples of whales, the above-mentioned features are extracted and standardized to form a unified format of whale acoustic model set. The spectral characteristics of the current speech signal are extracted from the multipath speech enhancement data to be processed, and matched with the whale acoustic model. Based on the matching result, the closest whale acoustic template is selected, and a dynamic frequency compensation filter is constructed accordingly. The filter can be implemented by using an adaptive filtering design method, and its frequency response function includes a gain control term and a phase compensation term, which can dynamically adjust the frequency response according to the current spectral characteristics, so as to adapt to the complex environmental propagation environment and the sound characteristics of whales. The constructed dynamic frequency compensation filter is applied to the multipath speech enhancement data, and the speech signal is filtered by frequency domain transformation. In the filtering process, the signal energy of different frequency bands is weighted and allocated according to the frequency energy distribution characteristics in the whale acoustic model. This process effectively strengthens the main components of the signal and suppresses background interference and frequency drift noise. After completing the frequency energy redistribution, the signal is restored to the time domain by inverse transformation to obtain the environmental speech signal fused with the compensation characteristics, i.e. the mixed compensation speech. The signal enhances the stability and recognizability of the original speech content in the environment. Since the relative speed difference in the environmental propagation process causes the speech signal to produce Doppler shift, the mixed compensation speech needs to be frequency-shifted and corrected. First, the Doppler shift of the speech signal is estimated by combining the actual environmental parameters (such as the relative motion speed between the sound source and the receiving device, the sound speed, etc.). After obtaining the frequency shift estimation value, the spectrum of the mixed compensation speech is processed in the reverse direction to restore its true frequency position. At the same time, the phase alignment technology is used to ensure the coherence and stability of the corrected signal in the time-frequency domain, and to avoid harmonic structure distortion. The final output signal is a pure speech fundamental signal that has been compensated for multipath, energy reconstruction, and frequency shift correction, and has good spectral stability, structure clarity, and subsequent analysis adaptability, which can be used for environmental speech recognition, target classification, and other tasks.
[0079] Step S4: analyzing the signal change trend of the pure speech fundamental signal to generate a fundamental signal change trend curve; deploying the dynamic frequency compensation filter through the lightweight edge computing architecture of the fundamental signal change trend curve to perform real-time mixed speech optimization tasks.
[0080] In the embodiment of the present application, the pure voice fundamental frequency signal obtained in step S3 is used for time sequence data extraction and multi-scale feature modeling. The extracted features include but are not limited to fundamental frequency variation rate, frequency band jump amplitude, amplitude stability, periodic variation, etc. The sliding window method and the first derivative analysis method are used to dynamically track the changes of the fundamental frequency signal on the time axis, and the adaptive smoothing filter algorithm (such as Savitzky-Golay filter) is used to denoise the frequency jitter. A continuous and smooth fundamental frequency signal trend curve is generated. Polynomial fitting, spline interpolation or wavelet transform method is used to model and segment the trend of the fundamental frequency, so that it has strong predictability and adaptability. Finally, the trend curve can be expressed as a function: Where f0 is the initial value of the fundamental frequency, φ i (t) is the trend base function, a i is the fitting coefficient. The frequency adjustment logic of the dynamic frequency compensation filter is converted into a function mapping relationship with the trend curve function F(t). The parameter sharing, interval prediction and incremental update mechanism are used to reduce the dependence on high-precision frequency domain operations, so as to realize the simplified expression of the filter in the resource-limited environment. The lightweight filter logic is integrated into the edge computing node, which can be an embedded processing chip, a low-power microprocessor unit or a buoy terminal platform on the environmental device. By embedding the trend prediction model in the edge node, the current fundamental frequency signal update value is received in real time, and the filter response parameters are quickly adjusted based on the trend curve to realize the approximate real-time dynamic compensation operation. In the running stage, the edge node dynamically adjusts the filter behavior according to the latest fundamental frequency signal value and the trend curve prediction result received each time, and performs fast filtering processing on the input original environmental voice data. The environmental voice output optimized by the edge has higher clarity, environmental adaptability and communication stability.
[0081] Preferably, step S1 comprises the following steps:
[0082] Step S11: Collecting mixed voice and environmental influence parameters;
[0083] Step S12: Signal conditioning of mixed voice to generate mixed frame-level voice;
[0084] Step S13: Bionic filter analysis of mixed frame-level voice to generate bionic spectrum data; analyzing the frequency attenuation characteristics in the bionic spectrum data, and performing inverse filtering and energy reconstruction on the bionic spectrum data to obtain low-frequency attenuation compensation feature data;
[0085] Step S14: Extracting propagation features in environmental influence parameters to obtain propagation environment feature data; performing beam tracking path simulation on the propagation environment feature data to generate path simulation data;
[0086] Step S15: according to the path simulation data, the low-frequency attenuation compensation characteristic data is modeled for multipath effect propagation, and multipath coupling characteristic data is generated; the multipath coupling characteristic data is channel estimated, and channel distortion data is generated.
[0087] In the embodiments of the present application, environmental microphones (such as hydrophones) are used to collect data in different environments. The frequency response of the sensor needs to be able to capture a wide band of signals from 20 Hz to 10 kHz. Through the integrated environmental sound sensor (for example, a CTD device), the sound velocity, salinity, and temperature data are obtained. Using such devices, temperature, salinity, and sound velocity information can be obtained simultaneously in different environments, and these data are saved as a time series. The collected environmental speech signals will first be subjected to noise suppression through a band-pass filter (for example, a low-pass filter to remove high-frequency noise, and a band-pass filter to remove low-frequency interference). Then, the signal is subjected to time-frequency analysis using short-time Fourier transform (STFT), which converts the entire signal into several frames, with the size of each frame adjusted according to the characteristics of the environmental signal. The common frame length is 20 ms, and the overlap is 50%. STFT is applied to each frame to convert the time-domain signal to a frequency spectrum. Then the frequency domain features (amplitude, phase, etc.) of each frame are extracted to form mixed frame-level speech data. Bionic filtering technology simulates the response of the biological ear to different frequencies, and a cochlea model filter (a filter based on biological hearing principles) is usually used. Here, the bionic filter processes the frequency spectrum of the mixed speech so that the frequency response is similar to that of the biological ear, with emphasis on compensating for the attenuation of low and medium frequencies. For the spectrum of the environmental signal, an inverse filtering algorithm is used to compensate for the low-frequency part. The goal of inverse filtering is to recover the attenuated low-frequency components using known environmental sound environment parameters (such as the attenuation curve). In this process, an inverse filtering technique based on minimum mean square error (MMSE) is used to compensate for the low-frequency attenuation in the signal. The sound velocity profile model is constructed using the collected environmental impact parameters (such as temperature, salinity, and depth). The sound velocity profile can be fitted using polynomial fitting or curve fitting methods based on the measured sound velocity values at different depths. Based on the sound velocity profile and the salinity-temperature-depth profile data, the propagation environment feature data is constructed, usually using linear interpolation or spline interpolation to fill in the missing environmental parameter data. These data are used as input to the propagation model to generate propagation environment feature data. On the basis of the environmental sound propagation model, a beam tracking algorithm (such as ray tracing) is applied to simulate the propagation path of the sound wave. Through this path simulation, the propagation phenomena of environmental sound waves such as reflection and refraction can be considered. Based on the path simulation data, a multipath propagation model is used for modeling. Here, the Rayleigh fading model is used, which is a standard model for signal attenuation in multipath propagation. The multipath propagation model can predict the attenuation effect of the signal caused by the coupling of multiple propagation paths (such as direct waves, reflected waves, etc.). After modeling the multipath effect, least squares estimation (LS) or Kalman filtering is used to estimate the channel. The purpose of channel estimation is to compensate for the channel distortion caused by the multipath effect. By estimating the delay, attenuation, and phase shift of each multipath signal, channel distortion data can be obtained and corrected to ultimately restore the quality of the original signal.
[0088] Preferably, the channel estimation on the multipath coupling characteristic data in step S15 comprises:
[0089] analyzing path delay of the multipath coupling characteristic data;
[0090] performing gain inversion on the multipath coupling characteristic data according to the path delay, to generate path gain data;
[0091] performing channel impulse response analysis on the path gain data, to generate multipath impulse response data;
[0092] performing time-varying modeling on the multipath impulse response data, to generate time-varying channel response data;
[0093] performing structure regularization and distortion metric analysis on the time-varying channel response data, to extract frequency response envelope, phase offset and asymmetric distortion index, to generate channel distortion data.
[0094] In the embodiment of the present application, cross-correlation technology is used to estimate the propagation delay of each multipath component in the environment signal. By comparing the received signal with the transmitted signal in the time domain, the time difference of different multipath signals arriving at the receiving end is calculated by using the cross-correlation function (CCF). Generally, time domain analysis can be performed by high-resolution FFT (Fast Fourier Transform) to accurately obtain the delay of each multipath. According to the delay value and signal strength, different multipath signals are screened out, and the propagation delay of each path is determined. In practical applications, the significant multipath paths can be extracted by the sparse representation method (such as matching pursuit). Once the delay information of the path is obtained, the gain of each path can be inverted by least squares estimation or maximum likelihood estimation. The specific method is to first compare the received signal with the template signal corresponding to the path delay, and estimate the attenuation degree of the signal. The result of inversion is the path gain data. By analyzing the signal power of each path, the gain of the corresponding path is calculated. Path gain G path Generally, it can be calculated by the following formula: where X received (t) is the signal received at time t, X ideal (t-τ) is the ideal transmitted signal received at time t, and τ is the path delay. The channel impulse response is a key indicator describing the propagation characteristics of the signal from the transmitting end to the receiving end. By using the path gain data and the path delay data, the channel impulse response can be calculated by inverse Fourier transform (IFFT): where h(t) is the channel impulse response, f i is the frequency of the path, G path is the path gain, and F-1 As an inverse Fourier transform, this impulse response describes the changing characteristics of the signal after passing through the channel, including the multipath interference effect. The channel response is dynamically changing, so it needs to be modeled as time-varying. In order to capture the changes of the time-varying channel, a time-varying convolution model can be used. Kalman filtering techniques are used to track the changes of the impulse response over time, or particle filtering is used to estimate the time-varying characteristics of the channel. By estimating the impulse response h(t) and gain G path (t) at each time point, the time-varying channel can be modeled to generate time-varying channel response data, and the goal of this time-varying modeling is to predict the changes of the channel response at future time, usually achieved by sliding window method or time domain filtering method. The time-varying channel response data is normalized to remove noise and ensure its stability. Wavelet transform or wavelet packet transform can be used to denoise the channel response data, and standard deviation analysis or mean square error (MSE) can be used to calculate the signal distortion. Channel distortion usually manifests as frequency response envelope, phase shift and asymmetric distortion. In order to extract these distortion features, envelope detection (such as Hilbert transform) can be performed on the frequency response of the channel to extract the envelope, phase shift can be calculated by phase demodulation technology to detect the change of phase, and nonlinear fitting technology can be used to analyze the asymmetry of the signal. By quantifying each distortion indicator (such as frequency response envelope, phase shift, asymmetric distortion), the final channel distortion data is generated, which can be used for subsequent channel optimization and compensation.
[0095] As an example of the present application, refer to FIG. 1, in which the step S2 in this example includes: Figure 2
[0096] Step S21: Time-domain deconvolution of mixed speech and channel distortion data to generate initial direct sound waveform data;
[0097] Step S22: Short-time Fourier transform of mixed speech to generate frequency domain reflected sound estimation data;
[0098] Step S23: Inter-frame feature alignment of initial direct sound waveform data and frequency domain reflected sound estimation data to generate joint acoustic separation feature data; time-frequency decomposition of joint acoustic separation feature data to extract direct sound principal component and reflected sound multipath interference component to generate direct sound component and reflected sound component;
[0099] Step S24: Adversarial training based on direct sound component and reflected sound component to generate anti-multipath speech enhancement data.
[0100] In the embodiments of the present application, the mixed speech signal and channel distortion data are obtained. In the environment, the signal is usually affected by multipath effect, so the direct sound and reflected sound (caused by channel distortion) are mixed together. The time domain deconvolution technology is used, which is a method of removing system response in time domain by convolution operation. Specifically, the goal of time domain deconvolution is to process the signal to minimize the distortion component in the signal by the known channel distortion characteristics. First, the channel response of the environment (such as the time delay and attenuation of the reflected sound) is estimated, which is usually estimated by modeling the sound wave propagation characteristics of the environment or using a known channel model. Then, by using the estimated channel response, the influence of the channel on the mixed signal is removed by the deconvolution algorithm, so that the component closest to the direct sound is extracted. After time domain deconvolution, the obtained signal is a direct sound waveform data that has preliminarily removed distortion. The mixed speech signal is input into the short-time Fourier transform (STFT) algorithm for frequency domain transformation. Fourier transform can convert the signal from time domain to frequency domain, so as to decompose the components of different frequencies. The short-time Fourier transform (STFT) is performed on the mixed speech signal to obtain the frequency domain representation of the signal in different time windows. By localizing in time and frequency, STFT can effectively extract the frequency domain characteristics of the signal. Through frequency domain analysis, the reflected sound component in the signal can be identified, because the reflected sound usually shows strong interference characteristics in certain frequency bands. Based on the frequency domain data, the reflected sound usually shows a specific spectral pattern, which is different from the direct sound. The reflected sound often has more frequency interference components due to the multipath effect. The reflected sound component is estimated by using frequency domain estimation methods, such as spectral decomposition and signal separation based techniques. The energy of the reflected sound is usually different from that of the direct sound, so the reflected sound component can be extracted by energy ratio, time delay characteristics or frequency domain analysis. Finally, frequency domain reflected sound estimation data are generated, which represent the reflected sound component in the environmental speech signal and usually contain multipath interference information. The direct sound waveform data and the reflected sound estimation data are divided into several time-frequency frames for feature alignment. Time-frequency frame refers to dividing the signal into small segments according to time window and performing frequency domain analysis on each small segment. Since the signal characteristics of the direct sound and the reflected sound are different, frame-to-frame feature alignment is needed. The goal of alignment is to ensure that the direct sound and reflected sound components in each frame are aligned in time for subsequent signal separation. The joint acoustic feature data in the frame are further decomposed to extract the multi-dimensional components of the direct sound and the reflected sound in the signal by using time-frequency decomposition technology. For example, wavelet transform or short-time Fourier transform is used to decompose the frequency spectrum component and time component of each frame signal. By these technologies, the direct sound main component and the multipath interference component of the reflected sound can be effectively separated. The main component of the direct sound is extracted by processing the time-frequency decomposition result, which is usually the part with the most concentrated energy.At the same time, the multipath interference component of the reflected sound is separated, which is usually the more dispersed and strong interference part in the signal. Through this process, two components are obtained: direct sound component and reflected sound component, which represent two important components in the environment signal, direct sound for speech recognition, and reflected sound for noise suppression. Using the generative adversarial network (GAN) framework, the generator and discriminator are constructed to optimize the environment speech enhancement process. The goal of the generator is to generate clear direct sound, while the goal of the discriminator is to determine whether the enhanced signal has the characteristics of real direct sound. The generator input is the direct sound component and the reflected sound component, and the trained generator learns how to enhance the direct sound component and suppress the influence of the reflected sound as much as possible. The discriminator evaluates the output of the enhanced speech signal to guide the generator to optimize, and finally makes the generator produce clearer direct sound component. After the adversarial training, the obtained anti-multipath speech enhancement data will significantly improve the clarity and quality of the environment speech signal, and reduce the interference of reflected sound on speech recognition.
[0101] Preferably, step S24 comprises the following steps:
[0102] Step S241: constructing an adversarial network for the direct sound component and the reflected sound component, constructing a separator-discriminator adversarial structure, and generating an adversarial training feature pair;
[0103] Step S242: conducting feature residual guidance training on the adversarial training feature pair to generate anti-multipath discriminant feature data;
[0104] Step S243: conducting multi-scale enhancement mapping on the anti-multipath discriminant feature data to generate anti-multipath speech enhancement data.
[0105] In the embodiments of the present application, the input data is the direct sound component and the reflected sound component that have been separated, which come from the time-frequency decomposition process in step S23. The direct sound component is the target signal (ideal speech), and the reflected sound component is the noise or interference signal (multipath interference). The task of the generator is to accept the input direct sound component and reflected sound component, generate an enhanced signal closer to the ideal direct sound through a neural network architecture. The generator will try to remove the interference from the reflected sound and enhance the quality of the direct sound. The task of the discriminator is to evaluate whether the generated output signal is real, that is, to judge whether it meets the characteristics of the direct sound. The discriminator not only evaluates whether the signal is real, but also scores according to the quality of the signal, thereby providing feedback to guide the optimization of the generator. The generator and the discriminator are optimized in the training process, the generator tries to deceive the discriminator and outputs signals closer and closer to the real direct sound, while the discriminator constantly improves its judgment ability to identify the false signals output by the generator. The direct sound component and the reflected sound component are used as input, and the adversarial training feature pairs are generated in the training process. The generator removes the interference of the reflected sound by learning, and improves the quality of the direct sound, while the discriminator provides feedback to generate feature pairs. These adversarial training feature pairs can be used as input for subsequent training to promote network learning to remove multipath interference. After adversarial training, a set of adversarial training feature pairs are generated as the data basis for subsequent training. In the generated adversarial training feature pairs, the feature residual is calculated. The feature residual refers to the difference between the generated signal and the real direct sound, which contains the components of reflected sound or multipath interference. By calculating the difference between the generated signal and the real signal (ideal direct sound), the multipath interference and noise components can be accurately found out for targeted optimization. In the training process, the residual information is used as a guide to design a residual-guided training mechanism. The feature residual is used as an auxiliary target to guide the model to focus on removing the reflected sound and improving the quality of the direct sound signal. By continuously minimizing the residual, the influence of the reflected sound can be effectively reduced, and the clarity and accuracy of the signal can be improved. After residual-guided training, the model generates anti-multipath discrimination feature data, which represents the signal after removing multipath interference. These feature data retain the main components of the direct sound and reduce the influence of the reflected sound. These feature data are used for subsequent enhancement mapping to further improve the quality of the environmental speech signal. The generated anti-multipath discrimination feature data can be used as input for the final speech enhancement process to help further optimize speech clarity. A multi-scale processing mechanism is designed to perform enhancement mapping on the anti-multipath discrimination feature data at multiple scales. This method can capture local details and global features of the signal at different scale levels, thereby more comprehensively optimizing the signal. The mapping network at each scale can perform fine-grained enhancement on the signal according to the features of different scales, such as using a multi-layer feature extraction mechanism of a **convolutional neural network (CNN)**.In the multi-scale network, the outputs of different scale levels are combined together to form a comprehensive signal enhancement effect. This process optimizes the clarity of the signal at different scales and removes the effects of reflected sound through training the network. During the training process, the model continuously adjusts the multi-scale mapping parameters to ensure that the signal quality at each scale is optimized. After the multi-scale enhancement mapping, the anti-multipath speech enhancement data obtained is an enhanced signal, whose direct sound component is clearer and the interference of reflected sound is greatly reduced.
[0106] Preferably, the energy redistribution of the anti-multipath speech enhancement data in step S3 includes:
[0107] Converting the anti-multipath speech enhancement data into a multi-channel time-frequency energy map using a dynamic frequency compensation filter;
[0108] Performing speech frequency response difference analysis on the multi-channel speech energy map to generate frequency energy distribution data;
[0109] Calculating the frequency band weight of the frequency energy distribution data to generate a frequency compensation weight;
[0110] Performing frequency domain filtering reconstruction on the anti-multipath speech enhancement data according to the frequency compensation weight to generate frequency energy redistribution speech data;
[0111] Performing time-domain waveform reconstruction on the frequency energy redistribution speech data to generate hybrid compensation speech.
[0112] In the embodiments of the present application, the speech signal is converted into a time-frequency domain representation by performing a short-time Fourier transform (STFT) on the anti-multipath speech enhancement data. On this basis, a multi-channel time-frequency energy map is generated. Each channel corresponds to a different frequency band or processing path, and different speech features can be extracted by selecting channels of different frequency ranges. The generated time-frequency energy map shows the energy changes of the speech signal in the time and frequency dimensions, providing rich time-frequency features. The energy value of each frequency point is calculated and visualized as a time-frequency energy map, where each row represents the frequency response at a certain time point, and the columns represent the frequency response changes at different times. The multi-channel speech energy map shows the frequency and time changes under multi-channel, providing a basis for subsequent analysis. The frequency response of each channel's time-frequency energy map is analyzed to evaluate the energy differences in different frequency ranges. The frequency response difference analysis aims to identify the fluctuations in the energy of the speech signal in the frequency region, especially those caused by the distortion of frequency components due to multipath effects. By analyzing the frequency response differences in the time-frequency energy map, the energy distribution in each frequency band is calculated, and the generated frequency energy distribution data reflects the degree to which different frequency bands are affected by multipath effects in the environment, helping to identify frequency distortion and providing a basis for subsequent compensation. Using a weighted average algorithm or an adaptive filter, a weight is assigned to each frequency band based on the frequency energy distribution data. The calculation of the weight is based on the energy distribution in the frequency band, and the frequency band with a larger frequency response difference will be given a larger weight to better compensate. Generally, lower frequency bands are more affected by multipath interference, so they will be assigned higher weights. Based on the above calculation results, a frequency compensation weight is generated, which can effectively enhance the energy of low and medium frequency bands while reducing the influence of frequency bands with smaller frequency response differences. The frequency compensation weight plays a key role in the compensation process, guiding how to process signals in different frequency bands to achieve the best speech enhancement effect. The generated frequency compensation weight is used to perform frequency domain filtering on the anti-multipath speech enhancement data. By applying the filter, the energy of different frequency bands is redistributed, especially by enhancing the energy of low and medium frequency bands and reducing the noise influence of high frequency bands. The process of frequency domain filtering can be realized through a convolutional neural network (CNN) or other frequency domain processing algorithms to ensure that the filtered signal is as close as possible to the real speech signal. After completing the frequency domain filtering, the generated frequency energy redistribution speech data contains the redistributed energy, improving the quality and intelligibility of the speech signal. These data are frequency-compensated speech signals suitable for time-domain waveform reconstruction. The frequency energy redistribution speech data is subjected to inverse short-time Fourier transform (ISTFT) to convert it from the frequency domain back to the time domain. In this process, the frequency domain signal is first converted into a time-frequency graph, and then the final time-domain waveform is generated through inverse transformation. The final output of the mixed compensation speech is the enhanced speech signal generated through the above processing steps, which has good clarity and reduces the influence of multipath interference.
[0113] Preferably, the step S3 includes:
[0114] The mixed compensation speech is subjected to biological frequency separation to separate the non-biological frequency therein to generate biological separation speech;
[0115] The sound source is positioned based on the biological separation speech to obtain biological motion trajectory data;
[0116] The relative speed and the sound wave propagation direction of the biological motion trajectory data are analyzed, and the Doppler frequency shift parameter is calculated for the mixed compensation speech to obtain the Doppler frequency shift parameter;
[0117] The mixed compensation speech is subjected to nonlinear phase correction through the Doppler frequency shift parameter to generate a corrected frequency domain signal;
[0118] The corrected frequency domain signal is subjected to frequency spectrum analysis, and the signal fundamental frequency is extracted to generate a pure speech fundamental frequency signal.
[0119] In the embodiment of the application, the frequency range related to the biological object is identified and extracted by using the frequency domain filtering or spectrum analysis method. The common biological frequency range is usually in a certain specific low frequency range. The non-biological frequency components (such as mechanical noise, external interference, etc.) in the environmental speech signal are removed by the band-pass filter or band-stop filter, and the biological related frequency is retained. The extracted biological related frequency is combined to generate a biological separation speech signal. The signal mainly contains the sound emitted by the biological object in the environment, and removes other irrelevant interference. Based on the environmental positioning system (such as sonar system, ultrasonic positioning technology, etc.), the space-time characteristics of the biological separation speech signal are used for environmental sound positioning. By analyzing the propagation time and propagation speed of the sound signal emitted by the biological object in the environment, and combining the received signals of multiple receivers, the three-dimensional position of the biological object is calculated. According to the positioning result, the motion trajectory of the biological object is tracked, and the position change thereof is recorded. This data is usually represented in the form of continuous time stamp and position coordinate, and the generated biological motion trajectory data will provide key information for subsequent Doppler frequency shift calculation, especially the motion speed and motion direction of the biological object. According to the biological motion trajectory data, the relative speed of the biological object relative to the receiver is calculated. The relative speed is calculated by the change rate of the current position of the biological object and the distance between the receiver position. At the same time, the relationship between the motion direction of the biological object and the sound wave propagation direction is analyzed. The sound wave propagation direction is usually determined by the relative position relationship between the receiver and the biological object. The Doppler effect formula is used: where f' is the observed frequency, f is the transmitted frequency, c is the speed of sound wave propagation, v is the speed of the organism relative to the receiver, and the positive sign indicates an approach, and the negative sign indicates a departure. Based on the relative speed of the organism and the direction of sound wave propagation, the Doppler shift parameter f'-f is calculated, which describes the frequency shift caused by the motion of the organism. In the frequency domain, the phase of each frequency component is adjusted to compensate for the phase shift caused by the Doppler effect. Based on the Doppler shift parameter, the phase adjustment is made to ensure that the phase change of different frequency components conforms to the motion trajectory of the organism. This nonlinear phase correction modifies the frequency shift of the frequency domain signal to restore the spectral characteristics of the original speech as much as possible. After nonlinear phase correction, a corrected frequency domain signal is generated, which has been adjusted to eliminate the Doppler shift effect caused by the motion of the organism. The Fourier transform is performed on the corrected frequency domain signal to obtain its spectral representation. By analyzing the spectrum, the strongest frequency component in the signal is extracted, which usually corresponds to the fundamental frequency of the speech signal. Based on the spectral analysis result, the fundamental frequency of the signal is extracted, which usually represents the low frequency part of the speech signal and is the basis of the tone of the speech. By removing high frequency noise and other irrelevant frequency components, a pure speech fundamental frequency signal is generated, which contains the core information of the speech.
[0120] As an example of the present application, refer to Figure 3 In this example, the step S4 includes:
[0121] Step S41: analyze the local frequency of the pure speech fundamental frequency signal, extract the trend characteristics of the local frequency, and generate the fundamental frequency signal trend data;
[0122] Step S42: sliding window analysis is performed on the fundamental frequency signal trend data to smooth out the short-term fluctuations of the signal and generate a smooth fundamental frequency signal trend curve;
[0123] Step S43: based on the lightweight edge computing architecture, the smooth fundamental frequency signal trend curve is used to deploy a dynamic frequency compensation filter to the preset edge computing node for real-time calculation and frequency compensation, and an edge computing optimized filter is generated to perform the quality optimization task of mixed speech.
[0124] In the embodiments of the present application, the frequency characteristics of the signal in the time domain are extracted by performing short-time Fourier transform (STFT) on the pure speech fundamental frequency signal. Through STFT, the spectral information of the signal in different time windows can be obtained, thereby analyzing the local frequency changes. The frequency envelope is calculated, which represents the frequency change trend of the signal over time. The frequency envelope can be extracted using Hilbert transform or envelope analysis method. Based on the frequency envelope, its rate of change is calculated, and its trend feature is extracted. Through time series analysis methods (such as autoregressive model AR, moving average model MA, etc.), the long-term and short-term trend of the fundamental frequency signal can be captured, and the fundamental frequency signal trend data is generated: through statistical analysis of the local frequency characteristics of the fundamental frequency signal, the fundamental frequency trend data is generated, which reflects the dynamic changes of the fundamental frequency, including the patterns of frequency rise, fall or fluctuation. The sliding window method is used to process the fundamental frequency signal trend data. The size of the sliding window is adjusted according to the fluctuation period of the signal and the target accuracy. Generally, the window size should be large enough to capture the long-period changes of the signal, but not too large to avoid losing important short-term information. In each sliding window, a weighted moving average or exponential weighted moving average (EWMA) method is used for smoothing. Through this method, short-term fluctuations in the signal can be effectively reduced while retaining the long-term trend of the signal. After sliding window processing, a smooth fundamental frequency signal trend curve is generated. This curve can reflect the stable change trend of the fundamental frequency, eliminating the interference caused by signal noise or short-term fluctuations. This smooth curve can provide accurate reference for subsequent frequency compensation and optimization processing. Utilizing the distributed computing power of the edge computing architecture, the dynamic frequency compensation filter is deployed to the preset edge computing nodes. Each edge node performs local calculation according to its geographical location and environmental speech environment, avoiding the bottleneck of transmitting all data to the central server. This deployment can use efficient algorithms such as convolutional neural network (CNN), support vector machine (SVM) or recurrent neural network (RNN) for real-time processing on edge computing nodes, ensuring fast response to environmental speech changes. By analyzing the smooth fundamental frequency signal trend curve, the edge computing node can calculate and apply the dynamic frequency compensation filter in real time. The filter compensates for the frequency according to the fundamental frequency trend, eliminating unnecessary frequency components in the signal while enhancing useful speech components. The frequency compensation filter can use an adaptive filtering algorithm to continuously update the compensation coefficients and adapt to environmental frequency changes in real time. Based on the results of real-time calculation, the edge node automatically adjusts the compensation strategy to optimize the parameters of the filter, thereby generating an edge computing optimization filter. The optimization filter can adaptively adjust in multiple environmental speech environments, providing stable frequency compensation effect, thereby significantly improving the quality of mixed speech. The quality optimization task of mixed speech is performed via the optimized filter.The optimization process includes improving the clarity of the voice, removing background noise, compensating for multipath effects, etc., ultimately achieving the improvement of voice quality. The optimized signal can be calibrated through a real-time feedback mechanism, constantly adapting to new environments and acoustic characteristics, ensuring continuous optimization of voice quality in dynamic environments.
[0125] Therefore, from any point of view, the embodiments should be considered as exemplary and non-limiting, the scope of the application being defined by the attached claims and not by the above description, therefore all variations falling within the meaning and scope of the equivalent elements of the application file are intended to be included in the present application.
[0126] The above description is merely illustrative of the application and not limiting thereof. Numerous modifications of the embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of hybrid speech processing, the method comprising: The method comprises the following steps: Step S1: collecting mixed speech and environmental influence parameters; performing bionics frequency domain analysis on the mixed speech to obtain low-frequency attenuation compensation feature data; performing multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental influence parameters to generate channel distortion data; Step S2: performing time domain-frequency domain joint deconvolution processing on the mixed speech using the channel distortion data to generate direct sound components and reflected sound components; Performing adversarial training based on the direct sound components and the reflected sound components to generate anti-multipath speech enhancement data; Step S3: constructing a dynamic frequency compensation filter based on preset environmental acoustic characteristics; Performing energy redistribution on the anti-multipath speech enhancement data using the dynamic frequency compensation filter to generate mixed compensation speech; performing Doppler frequency shift correction on the mixed compensation speech to generate a pure speech fundamental signal; Step S4: performing signal trend analysis on the pure speech fundamental signal to generate a fundamental signal trend curve; deploying a lightweight edge computing architecture for the dynamic frequency compensation filter through the fundamental signal trend curve to perform real-time mixed speech optimization tasks.
2. The method of hybrid voice processing of claim 1, wherein, Step S1 comprises the following steps: Step S11: collecting mixed speech and environmental influence parameters; Step S12: performing signal regularization on the mixed speech to generate mixed frame-level speech; Step S13: performing bionic filter analysis on the mixed frame-level speech to generate bionic spectrum data; analyzing the frequency attenuation characteristics in the bionic spectrum data, and performing inverse filtering and energy reconstruction on the bionic spectrum data to obtain low-frequency attenuation compensation feature data; Step S14: extracting propagation features from the environmental influence parameters to obtain propagation environment feature data; performing beam tracking path simulation on the propagation environment feature data to generate path simulation data; Step S15: modeling the multipath effect propagation of the low-frequency attenuation compensation feature data according to the path simulation data to generate multipath coupling feature data; performing channel estimation on the multipath coupling feature data to generate channel distortion data.
3. The method of hybrid voice processing of claim 2, wherein, The channel estimation on the multipath coupling feature data in step S15 comprises: analyzing the path delay of the multipath coupling feature data; performing gain inversion on the multipath coupling feature data according to the path delay to generate path gain data; performing channel impulse response analysis on the path gain data to generate multipath impulse response data; performing time-varying modeling on the multipath impulse response data to generate time-varying channel response data; performing structure regularization and distortion measurement analysis on the time-varying channel response data to extract frequency response envelope, phase shift and asymmetric distortion indicators to generate channel distortion data.
4. The method of hybrid voice processing of claim 1, wherein, Step S2 comprises the following steps: Step S21: performing time domain deconvolution on the mixed speech and the channel distortion data to generate initial direct sound waveform data; Step S22: performing short-time Fourier transform on the mixed speech to generate frequency domain reflected sound estimation data; Step S23: performing inter-frame feature alignment on the initial direct sound waveform data and the frequency domain reflected sound estimation data to generate joint acoustic separation feature data; performing time-frequency decomposition on the joint acoustic separation feature data to extract direct sound principal components and reflected sound multipath interference components to generate direct sound components and reflected sound components; Step S24: generating anti-multipath speech enhancement data based on the direct sound component and the reflected sound component through adversarial training.
5. The method of hybrid voice processing of claim 4, wherein, Step S24 includes the following steps: Step S241: constructing an adversarial network for the direct sound component and the reflected sound component, constructing a separator-discriminator adversarial structure, and generating an adversarial training feature pair; Step S242: performing feature residual guided training on the adversarial training feature pair to generate anti-multipath discriminant feature data; Step S243: performing multi-scale enhancement mapping on the anti-multipath discriminant feature data to generate anti-multipath speech enhancement data.
6. The method of hybrid voice processing of claim 1, wherein, The energy redistribution of the anti-multipath speech enhancement data using the dynamic frequency compensation filter in step S3 includes: Performing multi-channel time-frequency energy map conversion on the anti-multipath speech enhancement data using the dynamic frequency compensation filter to generate multi-channel speech energy map; Performing speech frequency response difference analysis on the multi-channel speech energy map to generate frequency energy distribution data; Calculating the frequency band weight of the frequency energy distribution data to generate a frequency compensation weight; Performing frequency domain filtering reconstruction on the anti-multipath speech enhancement data according to the frequency compensation weight to generate frequency energy redistribution speech data; Performing time-domain waveform reconstruction on the frequency energy redistribution speech data to generate hybrid compensation speech.
7. The method of hybrid voice processing of claim 1, wherein, The Doppler frequency shift correction of the hybrid compensation speech in step S3 includes: Performing biological frequency separation on the hybrid compensation speech to separate the non-biological frequency therein to generate biological separation speech; Performing sound source positioning based on the biological separation speech to obtain biological motion trajectory data; Analyzing the relative speed and sound wave propagation direction of the biological motion trajectory data, and performing Doppler frequency shift parameter calculation on the hybrid compensation speech to obtain the Doppler frequency shift parameter; Performing nonlinear phase correction on the hybrid compensation speech through the Doppler frequency shift parameter to generate corrected frequency domain signal; Performing frequency spectrum analysis on the corrected frequency domain signal and extracting the signal fundamental frequency to generate a pure speech fundamental frequency signal.
8. The method of hybrid voice processing of claim 1, wherein, Step S4 includes the following steps: Step S41: analyzing the local frequency of the pure speech fundamental frequency signal to extract the change trend feature of the local frequency, and generating fundamental frequency signal change trend data; Step S42: performing sliding window analysis on the fundamental frequency signal change trend data to smooth out the short-term fluctuations of the signal, and generating a smooth fundamental frequency signal change trend curve; Step S43: deploying the dynamic frequency compensation filter to the preset edge computing node based on the lightweight edge computing architecture for real-time calculation and frequency compensation, generating an edge computing optimized filter to perform quality optimization of the mixed speech.
9. An electronic device, comprising: The electronic device includes: at least one processor; a memory connected in communication with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for processing mixed speech according to any one of claims 1-8.
10. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method for processing mixed speech according to any one of claims 1-8.
Citation Information
Patent Citations
Multilayer carrier wave discrete multi-tone communication system and multilayer carrier wave discrete multi-tone communication method
CN102983893A
Underwater acoustic target recognition method based on adversarial residual network
CN113435276A
Cited By
Channel estimation with varying numbers of transmit layers
US12706780B2
Channel Estimation with Varying Numbers of Transmit Layers
US20260032020A1