Hybrid speech processing method, electronic equipment and computer readable medium

By collecting mixed speech and environment parameters, performing bionic frequency domain analysis and time domain-frequency domain combined deconvolution processing, a dynamic frequency compensation filter is constructed, which solves the problem of channel distortion in complex acoustic environments and improves the clarity and comprehensibility of speech signals.

CN120236599AActive Publication Date: 2025-07-01GUANGZHOU ZHIYU CLOUD NETWORK COMMUNICATIONS CO LTD

Patent Information

Application Number
CN202510516957.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-01
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Existing speech processing technologies are difficult to adapt to channel distortion caused by multipath propagation in complex acoustic environments, resulting in lower quality of mixed speech output.

Method used

By collecting mixed speech and environmental impact parameters, performing bionic frequency domain analysis, generating channel distortion data, and performing time-domain combined deconvolution processing, separating direct acoustic and reflected acoustic components, building a dynamic frequency compensation filter for energy redistribution, performing Doppler shift correction and lightweight edge computing optimization.

Benefits of technology

It significantly improves the clarity and comprehensibility of the voice signal, optimizes the frequency characteristics, reduces signal distortion caused by environmental noise, and achieves real-time and efficient voice optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236599A_ABST
    Figure CN120236599A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, in particular to a mixed voice processing method, electronic equipment and a computer readable medium. The method comprises the following steps: collecting mixed voice and environment influence parameters; carrying out bionic frequency domain analysis on the mixed voice to obtain low-frequency attenuation compensation characteristic data; performing multipath effect propagation analysis on the low-frequency attenuation compensation characteristic data through the environmental influence parameters to generate channel distortion data; performing time domain-frequency domain joint deconvolution processing on the mixed voice by using the channel distortion data to generate a direct sound component and a reflected sound component; performing adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data; and constructing a dynamic frequency compensation filter based on preset environmental acoustic characteristics. Through the multi-stage signal processing, frequency compensation and real-time optimization technology, the output quality of the mixed voice is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly to a method for hybrid speech processing, an electronic device, and a computer-readable medium. Background Art

[0002] Early speech processing relied on manual feature extraction techniques such as linear predictive analysis (LPC) and Mel-frequency cepstral coefficients (MFCC), combined with hidden Markov models (HMMs) or Gaussian mixture models (GMMs) to achieve speech recognition and analysis. However, these methods have limited performance in dealing with complex speech environments such as noise and accent changes. With the improvement of computing power and data accumulation, deep neural networks (DNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs) have been gradually introduced into the field of speech processing, improving the robustness of speech recognition, synthesis, and enhancement. In recent years, models based on the Transformer architecture (such as Wav2Vec and HuBERT) have made breakthroughs in speech processing tasks, further promoting the improvement of the accuracy of speech understanding and generation. However, in complex acoustic environments, speech signals often encounter channel distortion caused by multipath propagation, and existing systems often rely on static filters and are difficult to adapt to environmental changes, resulting in low-quality hybrid speech output. Summary of the Invention

[0003] Based on this, it is necessary to provide a method for hybrid speech processing, an electronic device, and a computer-readable medium to solve at least one of the above technical problems.

[0004] To achieve the above object, a method for hybrid speech processing includes the following steps:

[0005] Step S1: Collect hybrid speech and environmental impact parameters; perform bionic frequency-domain analysis on the hybrid speech to obtain low-frequency attenuation compensation feature data; perform multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental impact parameters to generate channel distortion data;

[0006] Step S2: Perform time-domain to frequency-domain joint deconvolution processing on the hybrid speech using the channel distortion data to generate a direct sound component and a reflected sound component; perform adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data;

[0007] Step S3: Construct a dynamic frequency compensation filter based on preset environmental acoustic features; perform energy redistribution on the anti-multipath speech enhancement data using the dynamic frequency compensation filter to generate hybrid compensation speech; perform Doppler frequency shift correction on the hybrid compensation speech to generate a pure speech fundamental frequency signal;

[0008] Step S4: Analyze the signal variation trend of the pure speech fundamental frequency signal to generate a fundamental frequency signal variation trend curve; deploy a lightweight edge computing architecture for the dynamic frequency compensation filter through the fundamental frequency signal variation trend curve to perform real-time hybrid speech optimization operations.

[0009] Through collecting hybrid speech and environmental impact parameters and performing bionic frequency domain analysis, the present invention can extract low-frequency attenuation characteristics and generate channel distortion data through multipath effect propagation analysis. This process lays a foundation for subsequent signal processing, especially the identification and compensation of multipath effects, reducing signal distortion caused by environmental noise. Using the channel distortion data for time-frequency joint deconvolution can effectively separate the direct sound and reflected sound components, providing a clear signal source for subsequent enhancement processing. This separation process is the basis of anti-multipath training, which can significantly improve the clarity of speech and enhance the intelligibility of environmental speech. By constructing a dynamic frequency compensation filter and performing energy redistribution on the anti-multipath speech enhancement data, it helps to optimize the frequency characteristics of the signal and reduce frequency offset caused by the environment. Doppler frequency shift correction further improves the accuracy of the environmental speech signal, ensuring that the generated pure speech signal is closer to the actual speech source. Analyzing the variation trend of the pure speech fundamental frequency signal can generate a fundamental frequency signal variation trend curve, which provides a basis for the lightweight edge computing deployment of the dynamic frequency compensation filter. This process makes the real-time hybrid speech optimization operation more efficient and accurate, reduces the computational burden, and optimizes the real-time processing ability. Therefore, through multi-stage signal processing, frequency compensation, and real-time optimization technologies, the present invention improves the output quality of hybrid speech.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: Collect hybrid speech and environmental impact parameters;

[0012] Step S12: Regularize the signal of the hybrid speech to generate hybrid frame-level speech;

[0013] Step S13: Perform bionic filtering analysis on the hybrid frame-level speech to generate bionic spectrum data; analyze the frequency attenuation characteristics in the bionic spectrum data, and perform inverse filtering and energy reconstruction on the bionic spectrum data to obtain low-frequency attenuation compensation feature data;

[0014] Step S14: Extract the propagation characteristics in the environmental impact parameters to obtain propagation environment feature data;; perform beam tracking path simulation on the propagation environment feature data to generate path simulation data;

[0015] Step S15: Perform multipath effect propagation modeling on the low-frequency attenuation compensation feature data according to the path simulation data to generate multipath coupling feature data; perform channel estimation on the multipath coupling feature data to generate channel distortion data.

[0016] By accurately collecting the mixed voice signal and environmental impact parameters, the present invention provides high-quality raw data for subsequent analysis and optimization. This step ensures the accuracy of signal analysis. Especially considering the complexity of special environments, collecting comprehensive data is the basis for improving system performance. By regularizing the mixed voice signal to generate mixed frame-level voice, the original voice signal can be converted into a standardized format, which helps subsequent analysis and processing. The regularization process of this step reduces the stray noise and irregularities in the signal, ensuring the consistency and operability of the signal. Conducting bionic filtering analysis on the mixed voice and generating bionic spectrum data can simulate the actual sound propagation characteristics in the environment. Through inverse filtering and energy reconstruction, the low-frequency attenuation compensation characteristic data is effectively generated. This compensation method effectively restores the low-frequency loss caused by the environment, improving the clarity and quality of the voice signal. Extracting the propagation characteristics (such as sound speed profile fitting and salinity-temperature-depth structure quantification) from the environmental impact parameters and conducting beam tracking path simulation can accurately understand the propagation characteristics of the environment. These environmental characteristics have an important impact on the propagation behavior of the signal, helping subsequent processing better adapt to the actual environment and further improving the transmission quality of the signal. By performing multipath effect propagation modeling on the path simulation data, the multipath effect in environmental signal transmission can be comprehensively simulated and effectively compensated. At the same time, the channel estimation step can accurately generate channel distortion data, thus providing strong support for subsequent voice enhancement and restoration. This step can significantly reduce the signal distortion caused by environmental multipath effects and improve the intelligibility of the voice.

[0017] Preferably, the channel estimation of the multipath coupling characteristic data in step S15 includes:

[0018] Analyzing the path delay of the multipath coupling characteristic data;

[0019] Performing gain inversion on the multipath coupling characteristic data according to the path delay to generate path gain data;

[0020] Conducting channel impulse response analysis on the path gain data to generate multipath impulse response data;

[0021] Performing time-varying modeling on the multipath impulse response data to generate time-varying channel response data;

[0022] Conducting structural regularization and distortion metric analysis on the time-varying channel response data, extracting frequency response envelope, phase offset, and asymmetric distortion indexes, and generating channel distortion data.

[0023] The present invention can help accurately understand the propagation time differences of different signal paths by analyzing the path delay. In an environment, multipath signals can cause different reflection paths and delays. By precisely analyzing the path delay, it can provide a basis for subsequent gain inversion and multipath modeling, ensuring the accuracy and clarity of signal recovery. Gain inversion is performed on the multipath coupling feature data according to the path delay to generate path gain data. Gain inversion can effectively compensate for signal attenuation and uneven gain caused by the multipath effect, ensuring that the signal intensities on different paths are appropriately adjusted and optimized. This provides more accurate signal characteristics for subsequent impulse response analysis and time-varying modeling, enhancing the reliability of the signal. By performing channel impulse response analysis on the path gain data, the impulse response characteristics in the multipath effect can be accurately obtained. The generated multipath impulse response data reflects the propagation characteristics of the signal on each path, helping to understand and compensate for signal distortion, especially the differences between the reflection path and the direct path. This step can effectively restore the characteristics of the signal in a complex environment, thereby improving the voice quality. The time-varying characteristics of the multipath channel can cause the signal to change over time. Performing time-varying modeling on the multipath impulse response data can accurately reflect the dynamic changes of the channel at different times, generating time-varying channel response data. This modeling method can help cope with the rapid changes of signals in the environment and improve the adaptability of the system, ensuring the stable transmission of signals in a dynamic environment. Performing structural regularization and distortion metric analysis on the time-varying channel response data can extract frequency response envelope, phase offset, and asymmetric distortion metrics, further revealing the details of channel distortion. By precisely analyzing these distortion metrics, it is possible to effectively identify and compensate for the distortion in the signal caused by factors such as the multipath effect and frequency offset, ensuring that the finally generated channel distortion data has high accuracy and reliability. The optimization of this step greatly improves the quality of the voice signal and reduces the noise and distortion in the voice.

[0024] Preferably, step S2 includes the following steps:

[0025] Step S21: Perform time-domain deconvolution on the mixed voice and the channel distortion data to generate initial direct sound waveform data;

[0026] Step S22: Perform short-time Fourier transform on the mixed voice to generate frequency-domain reflected sound estimation data;

[0027] Step S23: Perform inter-frame feature alignment on the initial direct sound waveform data and the frequency-domain reflected sound estimation data to generate joint acoustic separation feature data; perform time-frequency decomposition on the joint acoustic separation feature data to extract the direct sound principal component and the reflected sound multipath interference component, generating a direct sound component and a reflected sound component;

[0028] Step S24: Perform adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath voice enhancement data.

[0029] Through time-domain deconvolution of the mixed speech and channel distortion data, the present invention can effectively remove the influence of channel distortion and recover the initial waveform data of the direct sound. The deconvolution process reduces the influence of multipath effects and attenuation in the channel on the signal, providing purer direct sound data for subsequent acoustic analysis. By performing short-time Fourier transform on the mixed speech, the speech signal is transformed into a frequency-domain representation, making the reflected sound components in the frequency domain more obvious. Frequency-domain reflected sound estimation can reveal the reflection paths in the environment and accurately estimate the reflected sound, providing reliable frequency-domain feature data for subsequent acoustic separation. By performing frame-by-frame feature alignment between the initial direct sound waveform data and the frequency-domain reflected sound estimation data, more accurate acoustic separation can be achieved. Through time-frequency decomposition, the principal components of the direct sound and the multipath interference components of the reflected sound can be extracted from the joint acoustic separation feature data. This decomposition method can analyze simultaneously in the time domain and the frequency domain, thus effectively reducing the interference of the reflected sound and extracting clean direct sound components and reflected sound components. Through adversarial training based on the direct sound components and the reflected sound components, the system can perform efficient speech enhancement. The process of adversarial training makes the model more robust in dealing with multipath effects, effectively reducing the interference of the reflected sound on the direct sound, thereby generating anti-multipath speech enhancement data. This process makes the final speech signal clearer and more understandable, and the quality of speech communication is significantly improved even in complex environments.

[0030] Preferably, step S24 includes the following steps:

[0031] Step S241: Construct an adversarial network for the direct sound components and the reflected sound components, build a separator-discriminator adversarial structure, and generate adversarial training feature pairs;

[0032] Step S242: Perform feature residual-guided training on the adversarial training feature pairs to generate anti-multipath discriminative feature data;

[0033] Step S243: Perform multi-scale enhancement mapping on the anti-multipath discriminative feature data to generate anti-multipath speech enhancement data.

[0034] In the present invention, by constructing an adversarial network for the direct sound component and the reflected sound component, adopting the adversarial structure of a separator - discriminator, the direct sound and the reflected sound in the signal are effectively separated. The key of this adversarial structure lies in evaluating the output of the separator by the discriminator, forcing the separator to optimize its separation effect so as to generate clearer direct sound and reflected sound. Through this process, the generated pair of adversarial training features can promote the model's self - learning, thereby continuously improving the accuracy of signal separation. Conducting feature residual - guided training on the pair of adversarial training features can further optimize the separation effect of the direct sound and the reflected sound. Feature residual - guided training enables the model to focus on processing residual features, thereby reducing the influence of the multipath effect, enhancing the removal effect of the reflected sound, and generating anti - multipath discriminant feature data. This data can not only effectively distinguish the direct sound and the reflected sound but also capture the detailed features that are difficult to process in the multipath effect, thus improving the clarity of the final speech. Conducting multi - scale enhancement mapping on the anti - multipath discriminant feature data can handle the multi - scale characteristics in the signal and further enhance the signal quality. Multi - scale enhancement mapping can optimize the signal at different frequency and time scales. Especially in complex environments, multi - scale mapping can effectively reduce the interference at different scales, especially the difference between low frequencies and high frequencies, enhancing the naturalness and audibility of the final speech. Through this enhancement, the generated anti - multipath speech enhancement data can significantly improve the clarity and intelligibility of environmental voice communication.

[0035] Preferably, the energy redistribution of the anti - multipath speech enhancement data by using the dynamic frequency compensation filter in step S3 includes:

[0036] Converting the anti - multipath speech enhancement data into a multi - channel time - frequency energy map by using the dynamic frequency compensation filter to generate a multi - channel speech energy map;

[0037] Performing speech frequency response difference analysis on the multi - channel speech energy map to generate frequency energy distribution data;

[0038] Calculating the band weights of the frequency energy distribution data to generate frequency compensation weights;

[0039] Performing frequency - domain filtering reconstruction on the anti - multipath speech enhancement data according to the frequency compensation weights to generate frequency - domain energy - redistributed speech data;

[0040] Performing time - domain waveform reconstruction on the frequency - domain energy - redistributed speech data to generate hybrid compensation speech.

[0041] The present invention performs multi-channel time-frequency energy map conversion on multi-path voice enhancement data through a dynamic frequency compensation filter, which can convert voice signals into multi-channel energy representations, facilitating the analysis of energy distributions in different frequencies and time domains. The generated multi-channel voice energy map provides rich frequency-domain and time-domain features for subsequent analysis, helping to identify the impact of multi-path effects on voice signals in the environment. By analyzing the voice frequency response differences in the multi-channel voice energy map, the energy differences between different frequency bands can be revealed, especially the frequency response differences between reflected sounds and direct sounds. This analysis can help identify which frequency bands are strongly affected by environmental multi-path effects, providing an important basis for subsequent frequency compensation. The generated frequency energy distribution data can accurately describe the frequency response differences, providing the necessary data support for the design of the filter. Calculating the frequency band weights based on the frequency energy distribution data can assign different compensation weights to each frequency band. Through this process, it is possible to more precisely identify which frequency bands require stronger compensation, thereby achieving more efficient frequency compensation. The generated frequency compensation weights provide adaptive compensation parameters for the filter design, making the filter more accurate and flexible when performing energy redistribution. By performing frequency-domain filtering reconstruction on the multi-path voice enhancement data according to the frequency compensation weights, signals of different frequencies can be effectively compensated, reducing frequency attenuation and interference caused by multi-path effects. The frequency-domain energy redistribution optimizes the voice signal in the frequency domain, capable of restoring the attenuated or distorted parts, enhancing the clarity and quality of the voice signal. By performing time-domain waveform reconstruction on the frequency-domain energy redistributed voice data, the signal optimized in the frequency domain can be restored to the time-domain waveform, completing the final voice compensation process. The hybrid-compensated voice presents a clearer and more coherent voice waveform in the time domain, greatly improving the audibility and naturalness of the voice. This process enables the final voice to not only retain the clarity of the direct sound in the environment but also effectively reduce the impact of reflected sounds and multi-path effects.

[0042] Preferably, the Doppler frequency shift correction of the hybrid-compensated voice in step S3 includes:

[0043] Performing biological frequency separation on the hybrid-compensated voice to separate the non-biological frequencies therein, generating biologically separated voice;

[0044] Performing sound source localization based on the biologically separated voice to obtain biological motion trajectory data;

[0045] Analyzing the relative velocity and sound wave propagation direction of the biological motion trajectory data, and calculating the Doppler frequency shift parameters for the hybrid-compensated voice to obtain the Doppler frequency shift parameters;

[0046] Performing non-linear phase correction on the hybrid-compensated voice through the Doppler frequency shift parameters to generate a corrected frequency-domain signal;

[0047] Perform spectral analysis on the corrected frequency-domain signal and extract the fundamental frequency of the signal, thereby generating a pure speech fundamental frequency signal. Through bio-frequency separation of the hybrid-compensated speech, the present invention can effectively separate the bio-frequencies (such as the sounds of whales or fish) from the non-bio-frequencies (such as reflected sounds or background noises) in the speech signal. This process helps reduce the interference of bio-frequencies on the speech signal, making the remaining speech signal purer and clearer. The generated bio-separated speech provides a clean signal source for subsequent localization and correction processes. Based on the bio-separated speech, sound source localization can accurately obtain the movement trajectory data of environmental organisms, which is crucial for Doppler frequency shift correction because the Doppler frequency shift effect is closely related to the relative speed of the organism and the direction of sound wave propagation. With accurate movement trajectory data, the movement state of the organism can be accurately calculated, providing the necessary information for subsequent calculation of frequency shift parameters. Analyzing the relative speed and the direction of sound wave propagation in the biological movement trajectory data can help calculate the Doppler frequency shift parameters, which directly affect the frequency and phase of the environmental speech signal, especially when there is biological movement. By accurately calculating the Doppler frequency shift parameters, a precise reference can be provided for subsequent non-linear phase correction, thereby reducing the distortion caused by frequency shift and restoring the original frequency characteristics of the signal. Using the Doppler frequency shift parameters to perform non-linear phase correction on the hybrid-compensated speech can eliminate the phase shift caused by biological movement and restore the frequency-domain structure of the signal. Non-linear phase correction can effectively cope with the Doppler effect in complex environments, reducing the blurring or distortion of the speech signal caused by phase distortion, thereby improving the clarity and intelligibility of the speech. By performing spectral analysis on the corrected frequency-domain signal and extracting the fundamental frequency of the signal, a pure speech fundamental frequency signal can be obtained. The fundamental frequency is the most critical part of the speech signal, representing the main pitch and tone of the speech. Extracting the pure fundamental frequency signal helps restore the natural timbre of the speech and ensures the audibility and accuracy of the speech in the environment. The generated pure speech fundamental frequency signal provides a high-quality audio basis for subsequent speech reconstruction and optimization.

[0048] Preferably, step S4 includes the following steps:

[0049] Step S41: Analyze the local frequency of the pure speech fundamental frequency signal, extract the change trend characteristics of the local frequency, and generate fundamental frequency signal change trend data;

[0050] Step S42: Perform a sliding window analysis on the fundamental frequency signal change trend data to smooth out the short-term fluctuations of the signal and generate a stable fundamental frequency signal change trend curve;

[0051] Step S43: Based on the lightweight edge computing architecture, deploy the dynamic frequency compensation filter to a preset edge computing node for real-time calculation and frequency compensation by using the smooth fundamental frequency signal change trend curve, and generate an edge computing optimized filter to perform the quality optimization operation of the mixed voice.

[0052] By analyzing the local frequency of the clean voice fundamental frequency signal and extracting its change trend characteristics, the present invention can reveal the law of frequency change in the environmental voice signal. These change trend characteristics provide an important basis for subsequent filtering and optimization, and help to identify periodic changes, frequency drift or other instability factors in the signal. The generated fundamental frequency signal change trend data provides strong support for subsequent smoothing processing and dynamic frequency compensation. The sliding window analysis can smooth out the short-term fluctuations in the signal and eliminate the instantaneous changes caused by noise or environmental factors, making the fundamental frequency signal change trend smoother. This process can remove unnecessary noise and instantaneous anomalies, making the subsequent filtering and optimization work more accurate and reliable. The generated smooth fundamental frequency signal change trend curve provides stable reference data for the deployment of the dynamic frequency compensation filter, thereby improving the accuracy and effect of optimization. Using the lightweight edge computing architecture to transfer the smooth fundamental frequency signal change trend curve to a preset edge computing node for real-time calculation and frequency compensation can achieve efficient real-time processing. This distributed computing method can reduce the burden on the central server, reduce latency, and ensure that the voice optimization process can respond quickly in the real-time environment. The generated edge computing optimized filter can dynamically adjust the frequency compensation parameters, thereby realizing the real-time optimization of the mixed voice and improving the clarity, sound quality and intelligibility of the voice.

[0053] The present invention provides an electronic device, and the electronic device includes:

[0054] At least one processor;

[0055] A memory communicatively connected to the at least one processor;

[0056] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for processing mixed voice as described above.

[0057] The present invention also provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for processing mixed voice as described above is implemented.

[0058] The beneficial effects of the present invention are as follows: By collecting mixed speech and environmental impact parameters and conducting bionic frequency-domain analysis, it can accurately identify the low-frequency attenuation characteristics of environmental signals, providing basic data for subsequent low-frequency compensation. The generation of low-frequency attenuation compensation characteristic data enables targeted attenuation compensation for environmental speech signals during the signal processing process, improving the clarity and intelligibility of speech. Through time-domain to frequency-domain joint deconvolution processing, the direct sound and reflected sound in the mixed speech can be effectively separated, thereby providing more accurate sound components. This provides a clean signal source for subsequent speech enhancement, effectively suppressing the multipath interference caused by reflected sound and improving the quality of speech signals. After constructing a dynamic frequency compensation filter based on environmental acoustic characteristics, it can achieve energy redistribution for anti-multipath speech enhancement data, optimizing the frequency response of environmental speech and making the speech signal clearer and more stable in the environment. In addition, Doppler frequency shift correction is performed to ensure that the fundamental frequency of the signal is not affected by the frequency shift caused by environmental movement, thereby generating a purer speech fundamental frequency signal. By analyzing the changing trend of the pure speech fundamental frequency signal, valuable characteristic data can be extracted, which provides a basis for the stabilization of the fundamental frequency signal, thereby reducing the instability caused by noise and fluctuations. By deploying the dynamic frequency compensation filter in a lightweight edge computing architecture, real-time processing and frequency compensation are achieved, significantly reducing latency, enhancing the real-time performance and efficiency of environmental speech optimization operations, and ensuring a rapid response to speech optimization operations. Therefore, the present invention improves the output quality of mixed speech through multi-stage signal processing, frequency compensation, and real-time optimization technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a schematic diagram of the step flow of a method for processing mixed speech;

[0060] Figure 2 For Figure 1 It is a schematic diagram of the detailed implementation step flow of step S2 in

[0061] Figure 3 For Figure 1 It is a schematic diagram of the detailed implementation step flow of step S4 in

[0062] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0064] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0065] It should be understood that although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0066] To achieve the above object, please refer to Figures 1 to 3 , a method for hybrid voice processing, the method comprising the following steps:

[0067] Step S1: Collect hybrid voice and environmental impact parameters; perform bionic frequency-domain analysis on the hybrid voice to obtain low-frequency attenuation compensation feature data; perform multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental impact parameters to generate channel distortion data;

[0068] Step S2: Perform time-domain to frequency-domain joint deconvolution processing on the hybrid voice using the channel distortion data to generate a direct sound component and a reflected sound component; perform adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath voice enhancement data;

[0069] Step S3: Construct a dynamic frequency compensation filter based on a preset environmental acoustic feature; perform energy redistribution on the anti-multipath voice enhancement data using the dynamic frequency compensation filter to generate hybrid compensation voice; perform Doppler frequency shift correction on the hybrid compensation voice to generate a pure voice fundamental frequency signal;

[0070] Step S4: Perform signal change trend analysis on the pure voice fundamental frequency signal to generate a fundamental frequency signal change trend curve; deploy a lightweight edge computing architecture for the dynamic frequency compensation filter through the fundamental frequency signal change trend curve to perform real-time hybrid voice optimization operations.

[0071] By collecting mixed speech and environmental impact parameters and performing bionic frequency-domain analysis, the present invention can extract low-frequency attenuation characteristics and generate channel distortion data through multipath effect propagation analysis. This process lays the foundation for subsequent signal processing, especially the identification and compensation of multipath effects, reducing signal distortion caused by environmental noise. Using the channel distortion data for time-domain to frequency-domain joint deconvolution can effectively separate the direct sound and reflected sound components, providing a clear signal source for subsequent enhancement processing. This separation process is the basis of anti-multipath training, which can significantly improve the clarity of speech and enhance the intelligibility of environmental speech. By constructing a dynamic frequency compensation filter and redistributing the energy of the anti-multipath speech enhancement data, it helps to optimize the frequency characteristics of the signal and reduce the frequency offset caused by the environment. Doppler frequency shift correction further improves the accuracy of the environmental speech signal, ensuring that the generated pure speech signal is closer to the actual speech source. Analyzing the change trend of the pure speech fundamental frequency signal can generate a change trend curve of the fundamental frequency signal, which provides a basis for the lightweight edge computing deployment of the dynamic frequency compensation filter. This process makes the real-time mixed speech optimization operation more efficient and accurate, reducing the computational burden and optimizing the real-time processing ability. Therefore, the present invention improves the output quality of mixed speech through multi-stage signal processing, frequency compensation, and real-time optimization techniques.

[0072] In an embodiment of the present invention, with reference to Figure 1 shown in the figure, it is a schematic flowchart of the steps of a method for processing mixed speech according to the present invention. In this example, the method for processing mixed speech includes the following steps:

[0073] Step S1: Collect mixed speech and environmental impact parameters; perform bionic frequency-domain analysis on the mixed speech to obtain low-frequency attenuation compensation feature data; perform multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental impact parameters to generate channel distortion data;

[0074] In the embodiments of the present invention, audio data containing mixed signals such as speech, background noise, and echo are collected in a target area by using an environmentally highly sensitive acoustic sensor array (such as a hydrophone array, etc.). The sampling frequency of the mixed speech signal is generally set to 48 kHz to ensure complete coverage of human speech and low-frequency signals. Key physical parameters of the environment such as temperature and humidity are obtained by using environmental detection instruments, and these parameters will be used to model the environmental sound propagation characteristics, including sound speed distribution, refractive index, and multipath propagation structure. Simulating the processing mechanism of the human ear auditory system, bionic filtering and frequency-domain envelope extraction are performed on the mixed speech signal. A Gammatone filter bank can be used to simulate the cochlear filtering process and extract multi-band envelope features. The low-frequency region (generally <300 Hz) is enhanced emphatically to compensate for the energy loss of the low-frequency signal caused by environmental propagation. In each filtering channel, features such as short-time energy, signal-to-noise ratio, and frequency attenuation trend are calculated, and the low-frequency components are dynamically compensated by using an adaptive gain function. The output compensation result is the low-frequency attenuation compensation feature data, which is used for subsequent channel model adjustment. An environmental sound speed profile model (Sound Speed Profile, SSP) is constructed based on the physical parameters in the environment. The Bellhop Model or the wave equation simulation method is used to model the propagation of sound waves on different paths, considering phenomena such as reflection, refraction, and diffraction. The low-frequency attenuation compensation feature data is input into the above multipath propagation model to simulate its propagation changes in the current environment, and distortion phenomena such as delay spread, frequency drift, and amplitude attenuation are obtained. The amplitude-frequency response function on the above propagation path is convolved with the initial signal to simulate the actual propagation result of the signal in the environment, and the output is the channel distortion data. The channel distortion data includes: multipath response function, path delay distribution, frequency response distortion function, etc., which are used for environmental sound channel equalization or inverse reconstruction.

[0075] Step S2: Perform time-domain to frequency-domain joint deconvolution processing on the mixed speech by using the channel distortion data to generate a direct sound component and a reflected sound component; perform adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data;

[0076] In the embodiments of the present invention, the channel distortion data obtained in step S1 is represented as a multipath impulse response function h(t, f), which includes path delay, amplitude attenuation, and frequency distortion components. The mixed speech signal x(t) is modeled as: x(t) = s(t)·h(t) + n(t); where s(t) is the original speech signal at time point t, h(t) is the channel response at time point t, and n(t) is the background noise at time point t. Use a deconvolution filter (such as a Wiener filter, LMS deconvolver) to perform deconvolution on the mixed speech signal in the time domain to initially restore the speech signal on the direct path. Perform a short-time Fourier transform (STFT) on the mixed speech to extract the time-frequency spectrogram. Apply an inverse filtering method in the frequency domain to eliminate the channel frequency response H(f), restore high-frequency details, and suppress distortion components. In the deconvolution output, based on the path delay and channel energy distribution, use a time window function to divide the deconvolution result into: Direct Path Signal: The component with the shortest propagation delay, the strongest energy, and the most concentrated spectrum. Reflected Path Signal: The multipath component with a longer propagation delay, a spread spectrum, and overlapping interference. Construct a generative adversarial network (GAN) structure: The generator takes the direct path component and the reflected path component as inputs, learns the speech enhancement mapping, and outputs the anti-multipath speech enhancement result. The discriminator discriminates whether the generated enhanced speech is truly close to the original characteristics of the direct path sound, ensuring that the enhanced speech has clarity and time-frequency structure integrity. The training dataset includes real direct speech samples, artificially simulated reflected signals, and component pairs after joint deconvolution. Design the loss function: Adversarial loss: Constrains the authenticity of the generated speech. Perceptual loss: Ensures speech clarity and naturalness. Multipath residual constraint: Minimizes the interference of the reflected residual on the speech quality. The comprehensive loss function is as follows: L total = λ1L adv + λ2L perceptual + λ3L residual ; where L adv is the adversarial loss coefficient, L perceptual is the perceptual loss coefficient, L residual is the multipath residual constraint coefficient, λ1, λ2, and λ3 are loss weights. The trained generator can receive the separated direct path sound and reflected sound in real time and output high-quality speech with the multipath effect removed, that is, anti-multipath speech enhancement data.

[0077] Step S3: Construct a dynamic frequency compensation filter based on the preset environmental acoustic characteristics; use the dynamic frequency compensation filter to perform energy redistribution on the anti-multipath speech enhancement data to generate mixed compensation speech; perform Doppler frequency shift correction on the mixed compensation speech to generate a pure speech fundamental frequency signal;

[0078] In the embodiments of the present invention, by establishing an environmental acoustic feature database, the features include parameters such as the fundamental frequency range of cetacean vocalizations, frequency modulation patterns, spectral energy distribution, harmonic structure, and vocalization duration. By analyzing a variety of cetacean acoustic samples, the above features are extracted and standardized to form a set of cetacean acoustic models in a unified format. The spectral characteristics of the current speech signal are extracted from the multi-path speech enhancement data to be processed and matched with the cetacean acoustic models. Based on the matching results, the most similar cetacean acoustic template is selected, and a dynamic frequency compensation filter is constructed accordingly. The filter can be implemented using an adaptive filtering design method, and its frequency response function includes a gain control term and a phase compensation term, which can dynamically adjust the frequency response according to the current spectral characteristics to adapt to complex environmental propagation environments and cetacean vocalization characteristics. The constructed dynamic frequency compensation filter is applied to the multi-path speech enhancement data, and the speech signal is filtered through a frequency domain transformation method. During the filtering process, according to the frequency energy distribution characteristics in the cetacean acoustic model, weighted distribution of the signal energy in different frequency bands is performed. This process effectively strengthens the main components in the signal and suppresses background interference and frequency drift noise. After completing the frequency energy re-distribution, it is restored to the time domain through an inverse transformation to obtain an environmental speech signal integrated with compensation characteristics, that is, a hybrid compensation speech. This signal enhances its stability and recognizability in the environment while maintaining the original speech content. Due to the Doppler frequency shift of the speech signal caused by the relative speed difference during environmental propagation, it is necessary to correct the frequency shift of the hybrid compensation speech. First, combined with actual environmental parameters (such as the relative movement speed between the sound source and the receiving device, the speed of sound, etc.), the Doppler shift amount of the speech signal is estimated. After obtaining the frequency shift estimate, the spectrum of the hybrid compensation speech is subjected to reverse frequency shift processing to restore its true frequency position. At the same time, a phase alignment technique is adopted to ensure the coherence and stability of the corrected signal in the time-frequency domain and avoid harmonic structure distortion. The finally output signal is a pure speech fundamental frequency signal that has undergone multi-path compensation, energy reconstruction, and frequency shift correction, with good spectral stability, structural clarity, and adaptability for subsequent analysis, and can be used for tasks such as environmental speech recognition and target classification.

[0079] Step S4: Analyze the signal change trend of the pure speech fundamental frequency signal to generate a fundamental frequency signal change trend curve; deploy a lightweight edge computing architecture for the dynamic frequency compensation filter through the fundamental frequency signal change trend curve to perform real-time hybrid speech optimization operations.

[0080] In the embodiments of the present invention, timing data extraction and multi-scale feature modeling are performed on the pure speech fundamental frequency signal obtained in step S3. The extracted features include, but are not limited to, the fundamental frequency change rate, the frequency band jump amplitude, the amplitude stability, the periodic change, etc. The sliding window method and the first derivative analysis method are used to dynamically track the change of the fundamental frequency signal on the time axis, and the adaptive smoothing filtering algorithm (such as Savitzky-Golay filtering) is combined to denoise the frequency jitter. In this way, a continuous and smooth fundamental frequency signal change trend curve is generated. Methods such as polynomial fitting, spline interpolation or wavelet transform are used to model and segmentally reconstruct the fundamental frequency trend to make it have strong predictability and adaptability. Finally, the trend curve can be expressed in the form of a function: where f0 is the initial value of the fundamental frequency, and φ i (t) is the trend basis function, and a i is the fitting coefficient. The frequency adjustment logic of the dynamic frequency compensation filter is transformed into a function mapping relationship with the trend curve function F(t). The parameter sharing, interval prediction and incremental update mechanisms are adopted to reduce the dependence on high-precision frequency domain operations, so as to realize the concise expression of the filter in a resource-constrained environment. The lightweight filter logic is integrated into the edge computing node, which can be an embedded processing chip, a low-power microprocessing unit or a buoy terminal platform on the environmental device. By building a trend prediction model in the edge node, receiving the updated value of the current fundamental frequency signal in real time, and quickly adjusting the filter response parameters based on the trend curve, an approximate real-time dynamic compensation operation is realized. During the operation stage, the edge node dynamically adjusts the filter behavior according to the latest fundamental frequency signal value received each time and the trend curve prediction result, and performs fast filtering processing on the input original environmental voice data. The environmentally optimized voice output has higher clarity, environmental adaptability and communication stability.

[0081] Preferably, step S1 includes the following steps:

[0082] Step S11: Collect the mixed voice and environmental impact parameters;

[0083] Step S12: Regularize the mixed voice signal to generate mixed frame-level voice;

[0084] Step S13: Perform bionic filtering analysis on the mixed frame-level voice to generate bionic spectrum data; analyze the frequency attenuation characteristics in the bionic spectrum data, and perform inverse filtering and energy reconstruction on the bionic spectrum data to obtain low-frequency attenuation compensation feature data;

[0085] Step S14: Extract the propagation characteristics in the environmental impact parameters to obtain propagation environment feature data;; perform beam tracking path simulation on the propagation environment feature data to generate path simulation data;

[0086] Step S15: Perform multipath effect propagation modeling on the low-frequency attenuation compensation feature data according to the path simulation data to generate multipath coupling feature data; perform channel estimation on the multipath coupling feature data to generate channel distortion data.

[0087] In the embodiments of the present invention, an environmental microphone (such as a hydrophone) is used to collect data in different environments. The frequency response of the sensor needs to be able to capture broadband signals in the range of 20 Hz to 10 kHz. Through an integrated environmental acoustic environment sensor (such as a CTD device), sound speed, salinity, and temperature data are obtained. Using such devices, information such as temperature, salinity, and sound speed can be obtained simultaneously in different environments, and these data are saved as time series. The collected environmental voice signals will first be subjected to noise suppression through a band-pass filter (for example, a low-pass filter removes high-frequency noise, and a band-pass filter removes low-frequency interference). Then, the short-time Fourier transform (STFT) is used to perform time-frequency analysis on the signal, converting the entire signal into several frames, and the size of each frame is adjusted according to the characteristics of the environmental signal. A common frame length is 20 ms, and the overlap is 50%. STFT is applied to each frame to convert the time-domain signal into a spectrum. Then, the frequency-domain characteristics (amplitude, phase, etc.) of each frame are extracted to form mixed frame-level voice data. Bionic filtering technology simulates the response of the biological ear to different frequencies and usually uses a cochlear model filter (a filter based on the principle of biological hearing). Here, the bionic filter processes the spectrum of the mixed voice, making the frequency response similar to that of the biological ear, with a focus on compensating for the attenuation of low and middle frequencies. For the spectrum of the environmental signal, an inverse filtering algorithm is used to compensate for the low-frequency part. The goal of inverse filtering is to restore the attenuated low-frequency components through known environmental acoustic environment parameters (such as the attenuation curve). In this process, an inverse filtering technology based on the minimum mean square error (MMSE) is used to compensate for the low-frequency attenuation in the signal. The collected environmental impact parameters (such as temperature, salinity, depth) are used to construct a sound speed profile model. The sound speed profile can be fitted by polynomial fitting or curve fitting methods according to the measured sound speed values at different depths. Based on the sound speed profile and the CTD profile data, propagation environment characteristic data are constructed, and linear interpolation or spline interpolation is usually used to fill in the missing environmental parameter data. These data are used as the input of the propagation model to generate propagation environment characteristic data. Based on the environmental sound propagation model, a beam tracing algorithm (such as ray tracing) is applied to simulate the propagation path of sound waves. Through this path simulation, propagation phenomena of environmental sound waves such as reflection and refraction can be considered. Based on the path simulation data, a multipath propagation model is used for modeling. Here, the Rayleigh fading model is adopted, which is a standard model for signal attenuation in multipath propagation. The multipath propagation model can predict the attenuation effect caused by the coupling of signals through multiple propagation paths (such as direct waves, reflected waves, etc.). After modeling the multipath effect, the least squares estimation (LS) or Kalman filtering is used to estimate the channel. The purpose of channel estimation is to compensate for the channel distortion caused by the multipath effect. By estimating the delay, attenuation, and phase shift of each multipath signal, channel distortion data can be obtained and corrected, ultimately restoring the quality of the original signal.

[0088] Preferably, the channel estimation of the multipath coupling characteristic data in step S15 includes:

[0089] Analyze the path delay of the multipath coupling characteristic data;

[0090] Perform gain inversion on the multipath coupling characteristic data according to the path delay to generate path gain data;

[0091] Perform channel impulse response analysis on the path gain data to generate multipath impulse response data;

[0092] Perform time-varying modeling on the multipath impulse response data to generate time-varying channel response data;

[0093] Perform structural regularization and distortion metric analysis on the time-varying channel response data, extract frequency response envelope, phase offset and asymmetric distortion index, and generate channel distortion data.

[0094] In the embodiment of the present invention, the cross-correlation technique is used to estimate the propagation delay of each multipath component in the environmental signal. By comparing the received signal with the transmitted signal in the time domain, the cross-correlation function (CCF) is used to calculate the time difference between different multipath signals arriving at the receiving end. Usually, high-resolution FFT (Fast Fourier Transform) can be used for time-domain analysis to accurately obtain the delay of each multipath. According to the delay value and signal strength, different multipath signals are screened out, and the propagation delay of each path is determined. In practical applications, the sparse representation method (such as matching pursuit) can be used to extract significant multipath paths. Once the delay information of the path is obtained, the gain of each path can be inverted by least squares estimation (Least Squares Estimation) or maximum likelihood estimation (MLE). Specifically, first, the received signal is compared with the template signal corresponding to the path delay to estimate the attenuation degree of the signal. The result of the inversion is the path gain data. By analyzing the signal power of each path, the gain of the corresponding path is calculated. The path gain G path Usually can be calculated by the following formula: Where X received (t) is the signal received at time t, X ideal (t - τ) is the ideal transmitted signal received at time t, and τ is the path delay. The channel impulse response is a key index to describe the propagation characteristics of the signal from the transmitting end to the receiving end. Using the path gain data and path delay data, the impulse response of the channel can be calculated by inverse Fourier transform (IFFT): Where h(t) is the impulse response of the channel, f i is the frequency of the path, G path is the path gain, F-1 is the inverse Fourier transform. This impulse response describes the change characteristics of the signal after passing through the channel, including the multipath interference effect. The channel response is dynamically changing, so it is necessary to perform time-varying modeling on it. To capture the changes in the time-varying channel, a Time-Varying Convolution Model can be used. The Kalman Filtering technology is used to track the change of the impulse response over time, or the Particle Filtering is used to estimate the time-varying characteristics of the channel. By estimating the impulse response h(t) and the gain G path (t) at each time point, the time-varying channel can be modeled to generate time-varying channel response data. The goal of this time-varying modeling is to predict the change of the channel response at future moments, usually achieved by the sliding window method or the time-domain filtering method. Regularize the time-varying channel response data to remove noise and ensure its stability. Wavelet transform or wavelet packet transform can be used to denoise the channel response data, and the signal distortion degree can be calculated through standard deviation analysis or mean square error (MSE). Channel distortion usually manifests as frequency response envelope, phase shift, and asymmetric distortion. To extract these distortion features, envelope detection (such as Hilbert transform) can be performed on the frequency response of the channel to extract the envelope, the phase shift can be calculated by phase demodulation technology to detect the change of the phase, and non-linear fitting technology can be used to analyze the asymmetry of the signal. By quantifying each distortion index (such as frequency response envelope, phase shift, asymmetric distortion), the final channel distortion data can be generated, and these data can be used for subsequent channel optimization and compensation.

[0095] As an example of the present invention, refer to Figure 2 shown, in this example, the step S2 includes:

[0096] Step S21: Perform time-domain deconvolution on the mixed speech and the channel distortion data to generate initial direct sound waveform data;

[0097] Step S22: Perform short-time Fourier transform on the mixed speech to generate frequency-domain reflected sound estimation data;

[0098] Step S23: Perform inter-frame feature alignment on the initial direct sound waveform data and the frequency-domain reflected sound estimation data to generate joint acoustic separation feature data; perform time-frequency decomposition on the joint acoustic separation feature data, extract the direct sound principal component and the reflected sound multipath interference component, and generate the direct sound component and the reflected sound component;

[0099] Step S24: Perform adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data.

[0100] In the embodiments of the present invention, by obtaining a mixed speech signal and channel distortion data. In an environment, signals are usually affected by multipath effects, so the direct sound and the reflected sound (caused by channel distortion) will be mixed together. The time-domain deconvolution technique is used, which is a method of removing the system response in the time domain using convolution operations. Specifically, the goal of time-domain deconvolution is to process the signal through the known channel distortion characteristics to minimize the distorted components in the signal. First, estimate the channel response of the environment (such as the delay and attenuation of the reflected sound), which is usually estimated by modeling the characteristics of acoustic wave propagation in the environment or using a known channel model. Then, through the deconvolution algorithm, use the estimated channel response to remove the influence of the channel on the mixed signal, so as to extract the component as close to the direct sound as possible. After time-domain deconvolution, the obtained signal is a waveform data of the direct sound with preliminary distortion removal. Input the mixed speech signal into the short-time Fourier transform (STFT) algorithm for frequency-domain transformation. The Fourier transform can convert the signal from the time domain to the frequency domain, so as to decompose the components of different frequencies. Perform short-time Fourier transform (STFT) on the mixed speech signal to obtain the frequency-domain representation of the signal within different time windows. By localizing in time and frequency, STFT can effectively extract the frequency-domain characteristics of the signal. Through frequency-domain analysis, the reflected sound components in the signal can be identified, because the reflected sound usually shows strong interference characteristics in certain specific frequency bands. Based on the frequency-domain data, the reflected sound usually shows a specific spectral pattern. Different from the direct sound, the reflected sound often has more frequency interference components due to multipath effects. Adopt frequency-domain estimation methods, such as techniques based on spectral decomposition and signal separation, to estimate the reflected sound components. The energy of the reflected sound is usually different from that of the direct sound, so the reflected sound components can be extracted through the energy ratio, delay characteristics or frequency-domain analysis. Finally, generate frequency-domain reflected sound estimation data, which represent the reflected sound components in the environmental speech signal and usually contain information on multipath interference. Divide the direct sound waveform data and the reflected sound estimation data into several time-frequency frames for feature alignment. A time-frequency frame refers to dividing the signal into small segments according to a time window and performing frequency-domain analysis on each small segment. Since the signal characteristics of the direct sound and the reflected sound are different, frame-by-frame feature alignment is required. The goal of alignment is to ensure that the direct sound and reflected sound components in each frame can be aligned in time for subsequent signal separation. Using time-frequency decomposition techniques, further disassemble the joint acoustic feature data in the frame to extract the multi-dimensional components of the direct sound and the reflected sound in the signal. For example, use time-frequency analysis methods such as wavelet transform or short-time Fourier transform to disassemble the spectral components and time components of each frame of the signal. Through these techniques, the main component of the direct sound and the multipath interference component of the reflected sound can be effectively separated. Process the results of time-frequency decomposition to extract the main component of the direct sound, which is usually the part with the most concentrated energy.Meanwhile, the multipath interference components of the reflected sound are separated, which are usually the more dispersed and strongly interfering parts in the signal. Through this process, two components are obtained: the direct sound component and the reflected sound component, which represent two important components in the environmental signal. The direct sound is used for speech recognition, and the reflected sound is used for noise suppression. Using the generative adversarial network (GAN) framework, the environmental speech enhancement process is optimized by constructing a generator and a discriminator. The goal of the generator is to generate a clear direct sound, while the goal of the discriminator is to determine whether the enhanced signal has the characteristics of a real direct sound. The generator takes the direct sound component and the reflected sound component as inputs, and through training, the generator learns how to enhance the direct sound component and try to suppress the influence of the reflected sound. The discriminator guides the optimization of the generator by evaluating the output enhanced speech signal, and finally makes the generator generate a clearer direct sound component. After adversarial training, the obtained multipath-resistant speech enhancement data will significantly improve the clarity and quality of the environmental speech signal and reduce the interference of the reflected sound on speech recognition.

[0101] Preferably, step S24 includes the following steps:

[0102] Step S241: Construct an adversarial network for the direct sound component and the reflected sound component, build a separator-discriminator adversarial structure, and generate adversarial training feature pairs;

[0103] Step S242: Conduct feature residual-guided training on the adversarial training feature pairs to generate multipath-resistant discriminant feature data;

[0104] Step S243: Perform multi-scale enhancement mapping on the multipath-resistant discriminant feature data to generate multipath-resistant speech enhancement data.

[0105] In the embodiments of the present invention, the input data is the already separated direct sound component and reflected sound component, and these components come from the time-frequency decomposition process of step S23. The direct sound component is the target signal (ideal speech), and the reflected sound component is the noise or interference signal (multipath interference). The task of the generator is to receive the input direct sound component and reflected sound component and generate an enhanced signal that is closer to the ideal direct sound through a neural network architecture. The generator will try to remove the interference from the reflected sound and enhance the quality of the direct sound. The task of the discriminator is to evaluate whether the generated output signal is real, that is, to judge whether it conforms to the characteristics of the direct sound. The discriminator not only evaluates whether the signal is real, but also scores according to the quality of the signal, so as to provide feedback and guide the optimization of the generator. The generator and the discriminator perform adversarial optimization during the training process. The generator tries to deceive the discriminator by outputting signals that are increasingly close to the real direct sound, while the discriminator continuously improves its judgment ability to identify the fake signals output by the generator. Taking the direct sound component and the reflected sound component as inputs, adversarial training feature pairs are generated during the training process. The generator learns to remove the interference of the reflected sound and improve the quality of the direct sound, while the discriminator provides feedback to generate feature pairs. These adversarial training feature pairs can be used as inputs for subsequent training to promote the network to learn to remove multipath interference. After adversarial training, a set of adversarial training feature pairs are generated as the data basis for subsequent training. In the generated adversarial training feature pairs, the feature residual is calculated. The feature residual refers to the difference part between the generated signal and the real direct sound, and it contains the components of the reflected sound or multipath interference. By calculating the difference between the generated signal and the real signal (ideal direct sound), the multipath interference and noise components can be accurately found, so as to optimize them targeted. During the training process, using the residual information as a guide, a residual-guided training mechanism is designed. Taking the feature residual as an auxiliary target, guiding the model to focus on the removal of the reflected sound and improving the quality of the direct sound signal. By continuously minimizing the residual, the influence of the reflected sound can be effectively reduced, and the clarity and accuracy of the signal can be improved. After residual-guided training, the model will generate anti-multipath discriminant feature data. These feature data represent the signal after removing the multipath interference. They retain the main components of the direct sound and reduce the influence of the reflected sound. These feature data are used for subsequent enhancement mapping to further improve the quality of the environmental speech signal. The generated anti-multipath discriminant feature data can be used as the input for the final speech enhancement process to help further optimize the speech clarity. A multi-scale processing mechanism is designed to perform enhancement mapping on the anti-multipath discriminant feature data at multiple scales. This method can capture the local details and global features of the signal at different scale levels, so as to optimize the signal more comprehensively. The mapping network at each scale can perform fine-grained enhancement on the signal according to the features at different scales, such as through the multi-layer feature extraction mechanism of the **Convolutional Neural Network (CNN)**.In a multi-scale network, the outputs at different scale levels are combined to form a comprehensive signal enhancement effect. This process optimizes the clarity of the signal at different scales through training the network and removes the influence of reflected sound. During the training process, the model continuously adjusts the multi-scale mapping parameters to ensure that the signal quality at each scale is optimized. After multi-scale enhancement mapping, the anti-multipath speech enhancement data obtained presents as an enhanced signal, with a clearer direct sound component and significantly reduced interference from reflected sound.

[0106] Preferably, the energy reallocation of the anti-multipath speech enhancement data using the dynamic frequency compensation filter in step S3 includes:

[0107] Performing multi-channel time-frequency energy map conversion on the anti-multipath speech enhancement data using the dynamic frequency compensation filter to generate a multi-channel speech energy map;

[0108] Performing speech frequency response difference analysis on the multi-channel speech energy map to generate frequency energy distribution data;

[0109] Calculating the band weights of the frequency energy distribution data to generate frequency compensation weights;

[0110] Performing frequency-domain filtering reconstruction on the anti-multipath speech enhancement data according to the frequency compensation weights to generate frequency-domain energy reallocation speech data;

[0111] Performing time-domain waveform reconstruction on the frequency-domain energy reallocation speech data to generate hybrid compensation speech.

[0112] In the embodiments of the present invention, by performing short-time Fourier transform (STFT) on the anti-multipath speech enhancement data, the speech signal is converted into a time-frequency domain representation. On this basis, a multi-channel time-frequency energy map is generated. Each channel corresponds to a different frequency band or processing path, and different speech features can be extracted by selecting channels in different frequency ranges. The generated time-frequency energy map shows the energy change of the speech signal in the time and frequency dimensions, providing rich time-frequency features. Calculate the energy value for each frequency point and visualize it as a time-frequency energy map. Each row in the map represents the frequency response at a certain time point, and the columns represent the frequency response changes at different times. The multi-channel speech energy map shows the frequency and time changes under multiple channels, providing a basis for subsequent analysis. Perform frequency response analysis on the time-frequency energy map of each channel to evaluate the energy differences in different frequency ranges. The frequency response difference analysis aims to identify the fluctuations of the speech signal energy in the frequency region, especially those frequency components distorted due to the multipath effect. By analyzing the frequency response differences in the time-frequency energy map, calculate the energy distribution within each frequency band. The generated frequency energy distribution data reflects the degree to which different frequency bands are affected by the multipath effect in the environment, helping to identify frequency distortion and providing a basis for subsequent compensation. Use a weighted average algorithm or an adaptive filter to assign a weight to each frequency band according to the frequency energy distribution data. The calculation of the weight is based on the energy distribution within the frequency band. Frequency bands with larger frequency response differences will be assigned larger weights for better compensation. Usually, the lower frequency bands are more affected by multipath interference, so higher weights will be assigned. Based on the above calculation results, generate a frequency compensation weight, which can effectively enhance the energy of the low-frequency and mid-frequency bands while weakening the influence of frequency bands with smaller frequency response differences. The frequency compensation weight plays a key role in the compensation process, guiding how to process signals in different frequency bands to achieve the optimal speech enhancement effect. Use the generated frequency compensation weight to perform frequency domain filtering on the anti-multipath speech enhancement data. By applying the filter, redistribute the energy of different frequency bands, especially by enhancing the energy of the low-frequency and mid-frequency bands and reducing the noise influence of the high-frequency band. The process of frequency domain filtering can be implemented through a convolutional neural network (CNN) or other frequency domain processing algorithms to ensure that the filtered signal is as close as possible to the real speech signal. After completing the frequency domain filtering, the generated frequency domain energy redistributed speech data contains the redistributed energy, improving the quality and intelligibility of the speech signal. These data are frequency-compensated speech signals suitable for time-domain waveform reconstruction. Perform inverse short-time Fourier transform (ISTFT) on the frequency domain energy redistributed speech data to convert it back from the frequency domain to the time domain. In this process, first convert the frequency domain signal into a time-frequency map, and then generate the final time-domain waveform through the inverse transform. The finally output mixed compensation speech is the enhanced speech signal generated through the above processing steps, with better clarity and reduced influence of multipath interference.

[0113] Preferably, the Doppler frequency shift correction of the mixed compensated speech in step S3 includes:

[0114] Performing biological frequency separation on the mixed compensated speech, separating the non-biological frequencies therein, and generating biologically separated speech;

[0115] Performing sound source localization based on the biologically separated speech to obtain biological motion trajectory data;

[0116] Analyzing the relative velocity and the sound wave propagation direction of the biological motion trajectory data, and calculating Doppler frequency shift parameters for the mixed compensated speech to obtain Doppler frequency shift parameters;

[0117] Performing non-linear phase correction on the mixed compensated speech through the Doppler frequency shift parameters to generate a corrected frequency domain signal;

[0118] Performing spectral analysis on the corrected frequency domain signal and extracting the fundamental frequency of the signal, thereby generating a pure speech fundamental frequency signal.

[0119] In the embodiments of the present invention, by using frequency domain filtering or spectral analysis methods, the frequency range related to organisms is identified and extracted. The common biological frequency range is usually within certain specific low-frequency ranges. The non-biological frequency components (such as mechanical noise, external interference, etc.) in the environmental speech signal are removed through a band-pass filter or a band-stop filter, and the biological-related frequencies are retained. The extracted biological-related frequencies are combined to generate a biologically separated speech signal. This signal mainly contains the sounds emitted by organisms in the environment, removing other irrelevant interferences. Based on an environmental positioning system (such as a sonar system, ultrasonic positioning technology, etc.), the environmental sound localization is performed using the spatio-temporal characteristics of the biologically separated speech signal. By analyzing the propagation time and propagation speed of the sound signal emitted by organisms in the environment, and combining the received signals of multiple receivers, the three-dimensional position of the organisms is calculated. According to the positioning result, the motion trajectory of the organisms is tracked, and its position change is recorded. This data is usually represented in the form of continuous timestamps and position coordinates. The generated biological motion trajectory data will provide key information for subsequent Doppler frequency shift calculations, especially the motion speed and motion direction of the organisms. According to the biological motion trajectory data, the relative velocity of the organisms relative to the receiver is calculated. The relative velocity is calculated by the rate of change of the current position of the organisms and the distance between the receiver positions. At the same time, the relationship between the biological motion direction and the sound wave propagation direction is analyzed. The sound wave propagation direction is usually determined by the relative position relationship between the receiver and the organisms. Using the Doppler effect formula: Where f' is the observed frequency, f is the transmitted frequency, c is the sound wave propagation speed, and v is the relative speed of the organism with respect to the receiver. A positive sign indicates approaching, and a negative sign indicates moving away. Based on the relative speed of the organism and the direction of sound wave propagation, the Doppler shift parameter f' - f is calculated, which describes the frequency shift caused by the movement of the organism. In the frequency domain, the phase of each frequency component is adjusted to compensate for the phase shift caused by the Doppler effect. Based on the Doppler shift parameter, phase adjustment is performed to ensure that the phase changes of different frequency components conform to the movement trajectory of the organism. This non-linear phase correction corrects the frequency shift of the frequency domain signal, enabling the corrected signal to restore the spectral characteristics of the original speech as much as possible. After non-linear phase correction, a corrected frequency domain signal is generated, which has been adjusted to eliminate the Doppler shift effect generated by the movement of the organism. The Fourier transform is performed on the corrected frequency domain signal to obtain its spectral representation. By analyzing the spectrum, the strongest frequency component in the signal is extracted, which usually corresponds to the fundamental frequency of the speech signal. Based on the spectrum analysis results, the fundamental frequency of the signal is extracted. The fundamental frequency usually represents the low-frequency part of the speech signal and is the pitch basis of the speech. By removing high-frequency noise and other irrelevant frequency components, a pure speech fundamental frequency signal is generated, which contains the core information of the speech.

[0120] As an example of the present invention, refer to Figure 3 As shown, in this example, step S4 includes:

[0121] Step S41: Analyze the local frequency of the pure speech fundamental frequency signal, extract the change trend characteristics of the local frequency, and generate the fundamental frequency signal change trend data;

[0122] Step S42: Perform a sliding window analysis on the fundamental frequency signal change trend data to smooth out the short-term fluctuations of the signal and generate a stable fundamental frequency signal change trend curve;

[0123] Step S43: Based on the lightweight edge computing architecture, use the stable fundamental frequency signal change trend curve to deploy the dynamic frequency compensation filter to a preset edge computing node for real-time calculation and frequency compensation, generating an edge computing optimized filter to perform the quality optimization operation of the mixed speech.

[0124] In the embodiments of the present invention, by performing short-time Fourier transform (STFT) on the pure speech fundamental frequency signal, the frequency characteristics of the signal in the time domain are extracted. Through STFT, the spectral information of the signal within different time windows can be obtained, thereby analyzing the changes in local frequencies. Calculate the frequency envelope, which represents the frequency change trend of the signal over time. The frequency envelope can be extracted using the Hilbert transform or envelope analysis method. Based on the frequency envelope, calculate its change rate and extract its change trend characteristics. Through time series analysis methods (such as autoregressive model AR, moving average model MA, etc.), the long-term and short-term change trends of the fundamental frequency signal can be captured, generating fundamental frequency signal change trend data. By performing statistical analysis on the local frequency characteristics of the fundamental frequency signal, fundamental frequency change trend data is generated, which reflects the dynamic changes of the fundamental frequency, including patterns of frequency increase, decrease, or fluctuation. The sliding window method is used to process the fundamental frequency signal change trend data. The size of the sliding window is adjusted according to the fluctuation period of the signal and the target accuracy. Generally, the window size should be large enough to capture the long-period changes of the signal, but not too large to avoid losing important short-term information. Within each sliding window, weighted moving average or exponentially weighted moving average (EWMA) method is used for smoothing. Through this method, the short-term fluctuations in the signal can be effectively reduced while retaining the long-term trend of the signal. After processing by the sliding window, a stable fundamental frequency signal change trend curve is generated. This curve can reflect the stable change trend of the fundamental frequency, eliminating the interference caused by signal noise or short-term fluctuations. This stable curve can provide an accurate reference for subsequent frequency compensation and optimization processing. Utilizing the distributed computing power of the edge computing architecture, the dynamic frequency compensation filter is deployed to the preset edge computing nodes. Each edge node performs local calculations according to its geographical location and the ambient voice environment, avoiding the bottleneck of transmitting all data to the central server. This deployment can use efficient algorithms such as convolutional neural network (CNN), support vector machine (SVM), or recurrent neural network (RNN) for real-time processing on the edge computing nodes to ensure quick response to voice changes in the environment. By analyzing the stable fundamental frequency signal change trend curve, the edge computing node can calculate and apply the dynamic frequency compensation filter in real time. This filter performs frequency compensation according to the fundamental frequency change trend, eliminating unnecessary frequency components in the signal while enhancing useful speech components. The frequency compensation filter can adopt an adaptive filtering algorithm, and by continuously updating the compensation coefficients, it can adapt to the frequency changes in the environment in real time. Based on the results of real-time calculations, the edge node will automatically adjust the compensation strategy and optimize the parameters of the filter, thereby generating an edge computing optimized filter. This optimized filter can perform adaptive adjustment in multiple ambient voice environments, providing a stable frequency compensation effect, thus significantly improving the quality of the mixed speech. Via the optimized filter, the quality optimization operation of the mixed speech is performed.The optimization process includes improving the clarity of speech, removing background noise, compensating for multipath effects, etc., and ultimately achieving an improvement in speech quality. The optimized signal can be calibrated through a real-time feedback mechanism to continuously adapt to new environments and acoustic characteristics, ensuring the continuous optimization of speech quality in a dynamic environment.

[0125] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be embraced within the present invention.

[0126] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.

Claims

1. A method for mixed speech processing, characterized in that: The following steps are involved: Step S1: collecting mixed speech and environmental impact parameters; performing bionic frequency domain analysis on the mixed speech to obtain low-frequency attenuation compensation characteristic data; performing multipath effect propagation analysis on the low-frequency attenuation compensation characteristic data through environmental impact parameters to generate channel distortion data; Step S2: performing time-domain-frequency-domain joint deconvolution processing on the mixed speech using the channel distortion data to generate a direct sound component and a reflected sound component; Conduct adversarial training based on direct sound components and reflected sound components to generate anti-multipath speech enhancement data; Step S3: constructing a dynamic frequency compensation filter based on preset environmental acoustic characteristics; A dynamic frequency compensation filter is used to redistribute energy of the multipath speech enhancement data to generate mixed compensated speech; Doppler frequency shift correction is performed on the mixed compensated speech to generate a pure speech baseband signal; Step S4: Analyze the signal change trend of the pure voice baseband signal to generate a baseband signal change trend curve; deploy a lightweight edge computing architecture for the dynamic frequency compensation filter through the baseband signal change trend curve to perform real-time mixed voice optimization operations.

2. The method for mixed speech processing according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting mixed speech and environmental impact parameters; Step S12: performing signal regularization on the mixed speech to generate mixed frame-level speech; Step S13: performing bionic filtering analysis on the mixed frame-level speech to generate bionic spectrum data; analyzing the frequency attenuation characteristics in the bionic spectrum data, and performing inverse filtering and energy reconstruction on the bionic spectrum data to obtain low-frequency attenuation compensation feature data; Step S14: extracting propagation characteristics from the environmental impact parameters to obtain propagation environment characteristic data; performing beam tracking path simulation on the propagation environment characteristic data to generate path simulation data; Step S15: performing multipath effect propagation modeling on the low-frequency attenuation compensation characteristic data according to the path simulation data to generate multipath coupling characteristic data; performing channel estimation on the multipath coupling characteristic data to generate channel distortion data.

3. The method for mixed speech processing according to claim 2, characterized in that: The channel estimation of the multipath coupling characteristic data in step S15 includes: Analyze path delays of multipath coupling signature data; Performing gain inversion on the multipath coupling characteristic data according to the path delay to generate path gain data; Performing channel impulse response analysis on the path gain data to generate multipath impulse response data; Perform time-varying modeling on multipath impulse response data to generate time-varying channel response data; The time-varying channel response data is subjected to structural regularization and distortion metric analysis, the frequency response envelope, phase shift and asymmetric distortion indicators are extracted, and the channel distortion data is generated.

4. The method for mixed speech processing according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: performing time domain deconvolution on the mixed speech and channel distortion data to generate initial direct sound waveform data; Step S22: performing short-time Fourier transform on the mixed speech to generate frequency domain reflected sound estimation data; Step S23: performing inter-frame feature alignment on the initial direct sound waveform data and the frequency domain reflected sound estimation data to generate joint acoustic separation feature data; performing time-frequency decomposition on the joint acoustic separation feature data to extract the direct sound main component and the reflected sound multipath interference component to generate the direct sound component and the reflected sound component; Step S24: performing adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data.

5. The method for mixed speech processing according to claim 4, characterized in that: Step S24 includes the following steps: Step S241: constructing an adversarial network for the direct sound component and the reflected sound component, building a separator-discriminator adversarial structure, and generating adversarial training feature pairs; Step S242: performing feature residual guided training on the adversarial training feature pairs to generate anti-multipath discrimination feature data; Step S243: Perform multi-scale enhancement mapping on the anti-multipath discrimination feature data to generate anti-multipath speech enhancement data.

6. The method for mixed speech processing according to claim 1, characterized in that: The method of using a dynamic frequency compensation filter to perform energy redistribution on the anti-multipath speech enhancement data in step S3 includes: Using dynamic frequency compensation filter to perform multi-channel time-frequency energy map conversion on multi-path speech enhancement data to generate multi-channel speech energy map; Perform speech frequency response difference analysis on multi-channel speech energy graphs to generate frequency energy distribution data; Calculate frequency band weights of frequency energy distribution data to generate frequency compensation weights; Perform frequency domain filtering and reconstruction on the multipath-resistant speech enhancement data according to the frequency compensation weights to generate frequency domain energy redistribution speech data; The frequency domain energy redistribution speech data is reconstructed in the time domain to generate mixed compensation speech.

7. The method for mixed speech processing according to claim 1, characterized in that: The step S3 of performing Doppler frequency shift correction on the mixed compensation speech includes: Performing biological frequency separation on the mixed compensated speech, separating the non-biological frequencies therein, and generating biological separated speech; Based on the separated speech of organisms, the sound source is localized to obtain the biological motion trajectory data; Analyze the relative speed and sound wave propagation direction of the biological motion trajectory data, and calculate the Doppler frequency shift parameters of the mixed compensation speech to obtain the Doppler frequency shift parameters; Performing nonlinear phase correction on the mixed compensation speech through Doppler frequency shift parameters to generate a corrected frequency domain signal; The corrected frequency domain signal is subjected to spectrum analysis and the signal fundamental frequency is extracted to generate a pure speech fundamental frequency signal.

8. The method for mixed speech processing according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: Analyze the local frequency of the pure voice fundamental frequency signal, extract the change trend characteristics of the local frequency, and generate fundamental frequency signal change trend data; Step S42: Perform sliding window analysis on the baseband signal change trend data to smooth out short-term fluctuations of the signal and generate a stable baseband signal change trend curve; Step S43: Based on the lightweight edge computing architecture, the dynamic frequency compensation filter is deployed to the preset edge computing node using the smooth baseband signal change trend curve for real-time calculation and frequency compensation, and an edge computing optimization filter is generated to perform the quality optimization operation of the mixed voice.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for mixed speech processing as described in any one of claims 1-8.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for mixed speech processing as claimed in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Communication method of underwater digital voice

    CN102034480A

  • Multilayer carrier wave discrete multi-tone communication system and multilayer carrier wave discrete multi-tone communication method

    CN102983893A

  • Underwater acoustic target recognition method based on adversarial residual network

    CN113435276A

  • Echo cancellation method and apparatus, computer readable storage medium, and terminal device

    WO2025043993A1

Cited By

  • Voice recognition method and system based on neural network

    CN120452436A

  • A neural network-based sound recognition method and system

    CN120452436B

  • Multi-source sound field positioning and separation algorithm based on deep neural network

    CN120595237A

  • A multi-source sound field positioning and separation algorithm

    CN120595237B

  • Multi-object vibration fusion voice perception and reconstruction method based on millimeter wave radar

    CN121438801A