Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

29 results about "SPEECH DISTORTION" patented technology

Speech Sound Disorders. Substitutions- one sound is used in place of another. Distortions- a sound is changed slightly and made incorrectly. Omissions- certain sounds are completely left out.

Controllable causal double-path speech enhancement method and system

The invention belongs to the technical field of speech enhancement, and particularly relates to a controllable causal double-path speech enhancement method and system, and the method comprises the steps: taking noisy speech collected in an actual scene as input, carrying out the short-time Fourier transform, and obtaining a complex spectrum containing an amplitude spectrum and a phase spectrum; constructing a streaming encoder, taking the complex spectrum as input, and extracting local correlation of the noisy speech in a time dimension and a frequency dimension; a controllable causal double-path module is constructed, long-term dependency relationship modeling is carried out on the time dimension and the frequency dimension, and linear modulation is carried out on channel characteristics through control parameters; constructing a streaming amplitude / phase decoder, and reconstructing an amplitude spectrum and a phase spectrum of the enhanced speech; and performing short-time inverse Fourier transform by combining the reconstructed amplitude spectrum and phase spectrum to obtain the waveform of the enhanced speech. According to the invention, the real-time adjustment of the output characteristics of the voice enhancement network is realized, and the noise residue and the voice distortion level can be timely balanced according to the demands of a listener in the voice enhancement process.
Owner:SHANDONG UNIV OF TECH

Noise reduction method and device, mobile terminal and medium

The invention relates to a noise reduction method and device, a mobile terminal and a medium, and is applied to the mobile terminal, and the method comprises the steps: obtaining a sound signal with noise; obtaining the noise intensity of the steady-state noise signal based on the noise-containing sound signal and a noise reduction model; based on the noise intensity of the steady-state noise signal and the first estimator gain function, removing the steady-state noise signal from the noise-containing sound signal to obtain the noise intensity of the target sound signal; and obtaining the target sound signal based on the noise intensity of the target sound signal. The noise estimation of the steady-state noise signal is carried out by using the pre-stored noise reduction model, the speed and accuracy of noise estimation can be effectively improved, a better noise reduction effect is obtained, the problems of noise residual, voice distortion, unstable volume and the like caused by inaccurate noise estimation are avoided, the learning difficulty of the noise reduction model is low, and the noise reduction efficiency is improved. The calculation amount of the noise reduction process is reduced, and the applicability of the noise reduction method is improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Robust speech recognition method and system based on multi-stage feature fusion

The invention discloses a robust speech recognition method and system based on multi-stage feature fusion, and relates to the technical field of speech recognition. The invention provides a speech recognition model for robust speech recognition, which comprises the following steps: firstly, extracting a magnitude spectrum | Y | from noisy speech Y through a speech coding part, coding the magnitude spectrum | Y | into a preliminary feature Ybasic, then, processing the Ybasic through a speech enhancement part to obtain a hidden feature Yhidden, a masking feature Ymask and a mapping feature Ymap, and finally, carrying out speech recognition on the hidden feature Yhidden, the masking feature Ymask and the mapping feature Ymap. Then, a fusion feature Ffuse is obtained through three-stage feature fusion, and finally, the Ffuse is decoded through a voice decoder to obtain Result. According to the model, feature complementation, feature interaction and high-low layer semantic alignment in the process are enhanced, the problem of voice distortion introduced by voice enhancement is systematically relieved, information loss in different stages is solved, and therefore the final voice recognition effect is guaranteed.
Owner:ANHUI UNIV

Ultra-low delay signal sequence conversion method and system

The invention provides an ultra-low time delay signal sequence conversion method and system, which can be widely applied to various time sequence conversion scenes, including hearing enhancement systems and equipment, such as hearing aid, denoising, simultaneous translation, sound conversion signal sequence, intelligent glasses, brain-computer systems and the like. According to the invention, the rapid high-fidelity noise elimination method is adopted, so that the limitation and difficulty in the traditional technology can be overcome. And for the problems of noise uncertainty and voice distortion in the sound related field, the high-quality voice can be recovered by recognizing the voice content as the intermediate language representation with high generalization, and meanwhile, the voice of the same speaker is synthesized by using a pre-trained artificial intelligence module. And the ultra-low time delay is realized by using a variable multi-partition block and a multi-head prediction mechanism in the process. Therefore, in the specific hearing enhancement field, the voice of the target speaker can be focused by using the voice separation technology, and the uncertainty of background noise is avoided.
Owner:SHANGHAI PEDAWISE INTELLIGENT TECH CO LTD

Voice hiding method and system based on Flow-Flow architecture, medium and equipment

The invention discloses a Flow-Flow architecture-based voice hiding method and system, a medium and equipment in the technical field of information hiding, and aims to solve the optimization problem of the existing information hiding technology. The method comprises the following steps: acquiring a Mel spectrogram of a to-be-hidden carrier text; mapping the secret voice to a hidden variable by using a FlowMap mapper; according to the hidden variable and the Mel spectrogram of the to-be-hidden carrier text, voice hiding is carried out through a FlowVocoder voice generator, and secret-carrying voice is obtained; performing text format conversion on the encrypted voice to obtain a Mel spectrogram of the encrypted voice in the text format; restoring the Mel spectrogram of the encrypted voice and the Mel spectrogram of the encrypted voice in the text format into hidden variables through a FlowVocoder voice inverter; and mapping the restored hidden variables into restored secret voice by using a FlowMap inverse mapper. According to the invention, the problem of secret voice extraction quality reduction caused by secret-carrying voice distortion is solved.
Owner:ARMY ENG UNIV OF PLA

Speech recognition method based on self-supervised pre-training and interactive fusion network

The application discloses a speech recognition method based on self-supervised pre-training and an interactive fusion network, constructs a speech recognition model, uses a self-supervised pre-training model as a feature extraction part after a speech enhancement module, effectively combines the speech enhancement module and the self-supervised pre-training method, and relieves speech distortion caused by speech enhancement; an interactive feature fusion method is used to fuse enhanced features and original audio features, so that information loss in the speech enhancement process is made up. By using the method, low-resource speech recognition results are more accurate, and the recognition accuracy of low resources in a complex environment is improved.
Owner:BEIJING TECH & BUSINESS UNIV

An automatic control system for audio output amplitude and method thereof

The present application relates to the technical field of short wave frequency band communication equipment, in particular to an automatic control system of audio output amplitude and a method thereof. The present application is convenient to build and easy to realize; by audio sampling, the present audio output amplitude and the change trend are actively identified, and the amplifier inside the DSP chip is controlled to dynamically compensate the audio digital signal sent to the speech compression circuit, so that the audio output amplitude meets the index requirement, the small short wave audio output amplitude change is ensured in the extreme environment, the speech distortion degree is reduced, the speech quality of the short wave communication is improved, and the communication requirement in the extreme environment is met.
Owner:SHAANXI FENGHUO ELECTRONICS

Voice noise reduction method, device and storage medium

The application relates to a voice noise reduction method, device and storage medium, wherein the voice noise reduction method comprises the following steps: acquiring a voice signal, extracting a frequency domain component of the voice signal, and dividing the frequency domain component into a high-frequency component and a low-frequency component according to the frequency size, wherein the frequency of the high-frequency component is greater than that of the low-frequency component; acquiring a sound source distance of the voice signal; and performing noise reduction on the voice signal, and in the case that the sound source distance is not lower than a preset threshold, the noise reduction intensity of the high-frequency component is not higher than that of the low-frequency component, thereby solving the problem of voice distortion when the sound source distance is far, and improving the quality of the voice signal.
Owner:ZHEJIANG HUACHUANG VISION TECH CO LTD

Language enhancement method and device, computer equipment and storage medium

The invention relates to the technical field of voice processing, can be applied to the fields of finance and medical treatment, and discloses a language enhancement method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining an input voice signal with noise; the input voice signal with noise is converted into noise embedded data through a pre-trained generative audio encoder; performing denoising processing on the noise embedded data through a denoising encoder to obtain clean embedded data; and converting the clean embedded data into an enhanced target voice signal through a pre-trained vocoder. According to the method, the naturalness of the enhanced voice and the consistency of the speaker are effectively improved, the modeling difficulty of complex noise distribution is reduced, the voice distortion is reduced, meanwhile, the model parameter quantity and the training complexity are greatly reduced, the reasoning speed is improved, and the real-time application can be realized in a low-resource environment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Live broadcast sound card noise reduction method based on environment feature recognition

The invention provides a live broadcast sound card noise reduction method based on environment feature recognition, relates to the technical field of noise reduction, and can realize dynamic adaptation and precise processing of multiple types of noise and reverberation factors in a live broadcast scene by constructing a noise reduction process oriented to environment feature recognition, thereby remarkably improving the definition and the reduction degree of voice signals. In particular, in a complex or violently changing live broadcast environment, intelligent identification can be carried out based on multiple parameters such as reverberation time, environmental noise energy and background spectrum entropy characteristics, and it is ensured that the noise reduction process has pertinence and real-time performance. Through a multi-stage processing strategy, reverberation tail sound is effectively suppressed, and residual noise is adaptively recognized and filtered, so that the common problem of voice distortion or detail loss in a traditional noise reduction method is avoided.
Owner:SHENZHEN MEIJIA PHOTOELECTRIC TECH CO LTD

A training method, device and equipment for speech enhancement model

The embodiments of the present application provide a method, apparatus, and device for training a speech enhancement model, which relates to the field of speech processing and recognition technology and is used to reduce speech distortion and suppress noise in audio through a speech enhancement model. In this method, a training sample set is first obtained; the first audio sample in the training sample set is preprocessed through the input layer of the speech enhancement model to obtain first audio sample data; the audio features of the first audio sample data are extracted through N hidden layers of the speech enhancement model; the audio features are respectively input into M output layers of the speech enhancement model to obtain M audio noise reduction results; the losses between the M audio noise reduction results and the audio masking results are determined through the loss functions corresponding to the M output layers to obtain M loss values; and the network parameters of the input layer, N hidden layers, and M output layers are adjusted according to the weighted results of the M loss values ​​to obtain a trained speech enhancement model.
Owner:QINGDAO HI-IMAGE TECH CO LTD

Voice instruction recognition method and system with noise robustness

The invention discloses a voice instruction recognition method and system with noise robustness. The method comprises the following steps: S1, constructing a training data set; s2, constructing a noise robustness voice instruction recognition model and training the model; and S3, processing actual noisy voice by using the noise robustness voice instruction recognition model trained in the step S2 so as to execute instruction recognition operation. Wherein the constructed noise robustness voice instruction recognition model comprises a voice enhancement model used for carrying out noise reduction processing on input voice with noise and outputting enhanced voice; the voice distortion sensor is used for outputting a distortion probability according to the voice enhancement model; the voice recognition model is used for processing the fused enhanced voice and outputting a voice instruction; according to the method, the voice distortion perceptron is arranged in the model to obtain the distortion probability in real time, difference mixing is carried out on the audio, voice distortion is reduced, and the robustness and accuracy of the whole system are further improved.
Owner:HANGZHOU DIANZI UNIV

Real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation

PendingCN121354580ASpeech analysisHigh level techniquesComputation complexitySpeech reconstruction
The invention discloses a real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation, and relates to a real-time single-microphone voice noise reduction algorithm. The invention aims to solve the problem that conversation voice is buried by noise due to noisy and diverse background noise of an interphone in a special communication scene. Noise power spectrum estimation is optimized by fusing a continuous minimum value tracking algorithm, a speech enhancement residual error is introduced as a real noise approximate value to participate in a recursive average process, and the response speed and accuracy of noise estimation are improved; a gain function is calculated in combination with an optimal correction logarithm MMSE estimator, and a closed expression approximates exponential integration to reduce the calculation complexity. The algorithm specifically comprises the steps of preprocessing and framing, noise power spectrum estimation, gain function calculation, voice reconstruction and the like, the segmentation signal-to-noise ratio is remarkably improved on the premise that low voice distortion is guaranteed, and the real-time processing requirement is met. The invention belongs to the technical field of voice signal processing.
Owner:HARBIN INST OF TECH

A doctor-patient communication information online synchronization system for a ward round vehicle

The application relates to the technical field of medical instruments, and discloses a ward round vehicle medical patient communication information online synchronization system, which comprises a data acquisition module, a data preprocessing module, a sound quality analysis module, a tone analysis module, a patient comprehensive analysis module, a report review module and a report generation module. Compared with traditional ward round technology, the system realizes comprehensive intelligentization and dataization of the ward round process through the integrated multi-module collaborative work. In terms of sound quality processing, the system solves the speech distortion problem in a high-noise environment through dynamic noise reduction and audio enhancement technology; in terms of dialect recognition, the system significantly improves the accuracy of dialect communication through tone analysis technology; in terms of report generation, the system greatly improves the efficiency and accuracy through intelligent decision-making and automatic generation mechanism. Overall, the system optimizes the ward round process, improves the medical patient communication quality, reduces the medical error rate, and provides more reliable support for clinical decision-making.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV

Compression method, system and equipment of speech synthesis model

The invention relates to a method, a system and equipment for compressing a speech synthesis model, and relates to the field of artificial intelligence. According to the technical scheme, the problem of traditional model deployment is solved, storage in an ONNX format occupies 112 MB originally, the hardware requirement is high, embedded equipment is difficult to bear, the model is only 30 MB after compression, reasoning is accelerated, and smooth operation can be achieved on edge equipment such as an ARM chip and a low-power-consumption MCU; moreover, the scheme is superior to a traditional compression means, solves the problem of proneness to voice distortion caused by static quantization and violent modification of a model structure, adopts quantization perception training to control errors, and ensures that voice details and tone quality are basically flush with those of an original model; in addition, according to the scheme, the training cost is reduced, 1000 epochs can reach the target on the basis of the original pre-training model, the GPU investment is reduced, the GPU cycle is shortened, and the threshold of enterprises and developers is reduced.
Owner:LOOTOM TELCOVIDEO NETWORK WUXI

A directional sound pickup and adaptive noise reduction system and method for a multi-microphone array

PendingCN122290623AReduce processing burdenavoid damageNoiseFeedback control
This invention relates to the field of audio signal processing technology, specifically to a directional sound pickup and adaptive noise reduction system and method using a multi-microphone array. The invention acquires signals through a multi-microphone array and performs time-frequency conversion. A beamforming unit forms a directional beam according to the target direction and constructs spatial nulls in the interference direction. A noise reduction unit adaptively suppresses noise in the beam output, simultaneously extracting residual noise energy and speech distortion feature values. A closed-loop feedback control unit executes a zero-point adaptive optimization algorithm based on the joint constraints of residual noise energy and speech distortion. Using the aforementioned feature values, a cost factor is calculated and a beam zero-point direction update is generated, which is fed back to the beamforming unit to achieve closed-loop adaptive adjustment of the null direction. This breaks the limitation that traditional beamforming and post-noise reduction are independent, and can simultaneously improve noise suppression capability and speech fidelity in complex acoustic environments.
Owner:SHENZHEN FUDEYUAN DIGITAL TECH CO LTD

Methods for synthesis-based clear hearing under noisy conditions

This invention provides a new and improved hearing aid system with high quality noise cancellation method and devices to overcome the limitations and difficulties encountered in conventional technologies. The technical limitations of the noise uncertainty and speech distortion in the hearing aid field are resolved by restoration of the high-quality speech by converting the speech content into an intermediate linguistic representation and by synthesizing the speech of the same speaker with pre-trained using artificial intelligence (AI) modules. In this invention, the noise uncertainties are circumvented by focusing on the target speaker or picking up the dominant speech by choosing the corresponding setting assuming the speech from the target speaker is the dominant speech based on the Lombard effect.
Owner:WANG FULIANG

Voiceprint-driven voice noise reduction method and terminal equipment

The invention relates to the technical field of voice processing, and discloses a voiceprint-driven voice noise reduction method and terminal equipment, and the method comprises the steps: collecting a terminal user voice sample, and generating a user exclusive voiceprint packet; collecting an audio input by a user, analyzing the input audio through an AI algorithm, separating voiceprint feature information in the audio, and obtaining input voiceprint data; performing similarity comparison on the input voiceprint data and a user exclusive voiceprint packet pre-stored in the terminal, setting a similarity threshold value, retaining audio signals of which the similarity reaches the threshold value, and filtering noise signals of which the similarity does not reach the threshold value; and extracting syllable parameters in the reserved audio signal, combining rhythm information, optimizing voice integrity through a feature completion algorithm, and outputting the voice integrity. According to the method, the exclusive voiceprint of the user is taken as a core screening basis, the method is not limited by the noise type, frequency and intensity, all interference signals without target voiceprint can be filtered, the method is adaptive to a complex use scene, and voice distortion caused by traditional noise reduction is avoided.
Owner:WESTVALLEY DIGITAL TECH

Voice interaction method, apparatus, device, medium and program product

The application provides a voice interaction method, device, equipment, medium and program product. The voice interaction method comprises the following steps: determining a first confidence degree of a first original signal and a second confidence degree of a first voice enhanced signal; the first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal; determining a target signal based on the first confidence degree, the second confidence degree, a second original signal and a second voice enhanced signal; the second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal; and performing voice interaction with a target device based on the target signal. The application can reduce voice distortion caused by voice enhancement.
Owner:IFLYTEK CO LTD

An underwater sound acquisition device capable of avoiding voice distortion

ActiveCN224418931Uaccurate identificationAvoid voice distortionSound waveBass (sound)
The utility model discloses an underwater sound acquisition device capable of avoiding voice distortion, which comprises a shell, a sound wave filtering assembly, a sound wave receiving assembly and a signal transmission part, the shell has a waterproof sealed cavity inside, the sound wave filtering assembly is installed in the waterproof sealed cavity, is used for filtering bass parts formed due to resonance of the sealed space, the sound wave receiving assembly is installed in the waterproof sealed cavity, is used for receiving sound waves filtered through the sound wave filtering assembly, and the signal transmission part is used for realizing signal transmission between the sound wave receiving assembly and external equipment. The underwater sound acquisition device can be installed on a waterproof breathing mask for use to collect human voices of underwater personnel. The sound wave filtering assembly arranged can filter out bass parts formed due to resonance of the sealed space, can avoid voice distortion, the sound wave receiving assembly receives filtered sound waves, and then transmits the sound waves to external equipment through the signal transmission part, so that the sound wave receiving party can receive clear voice information and correctly identify voice content.
Owner:DIVEVOLK ZHUHAI INTELLIGENCE TECH CO LTD

Intelligent manufacturing workshop safety interaction control method and system based on large model

The invention relates to the field of industrial control, in particular to an intelligent manufacturing workshop safety interaction control method and system based on a large model, and the method comprises the steps: collecting a sound signal of a workshop environment, calculating the instantaneous energy of a sliding window, constructing an energy concentration ratio, and discriminating local impact noise; introducing adaptive Kalman filtering by taking an energy concentration ratio as a prior factor, tracking a smooth energy change rate, calculating a window length parameter, and adaptively determining a dynamic window length; and after short-time Fourier transform noise reduction, multiplexing or reconstructing window length according to frame similarity, and inputting pure voice into the large voice recognition model to realize safe interaction control. According to the invention, the problems of noise trailing, insufficient frequency resolution and voice distortion easily caused by a traditional fixed window length are solved, the equipment impact noise and the artificial voice instruction can be effectively distinguished, the signal processing precision and the anti-interference capability are improved, and the method is suitable for high-reliability voice safety interaction in a complex industrial environment.
Owner:SHANDONG BLUEBIRD IND INTERNET CO LTD

Dynamic feature enhancement method for noise robustness speech recognition

The invention relates to the technical field of voice signal processing, in particular to a dynamic feature enhancement method for noise robustness voice recognition. The method comprises the following steps: acquiring a multi-dimensional acoustic context information set, and based on the multi-dimensional acoustic context information set, analyzing a masking and interference relationship between a disaster site noise dynamic characteristic and a rescue worker voice key characteristic to obtain a voice characteristic dynamic damage assessment parameter set; based on the multi-dimensional acoustic context information set, analyzing a competition and masking relationship between the target human voice of the rescue personnel and interference human voice and noise in a disaster site in an acoustic feature space, and obtaining a target sound source dynamic separation degree evaluation parameter set; and based on the voice feature dynamic damage evaluation parameter set and in combination with the target sound source dynamic separation degree evaluation parameter set, analyzing the composite stress intensity of residual noise and voice distortion on final recognition performance, executing a dynamically optimized feature reconstruction and enhancement instruction, and generating and outputting a target human voice enhancement process log. And the reliability of rescue communication is improved.
Owner:SHENZHEN BOSHITE TECH CO LTD

Personalized voice noise reduction and enhancement method based on user specific time domain envelope reconstruction

The invention discloses a personalized voice noise reduction and enhancement method based on user specific time domain envelope reconstruction, which relates to the field of voice signal processing, and comprises the following steps: monitoring the environmental noise level in real time when a user uses an earphone to carry out daily conversation or voice input; when the environmental noise level is monitored to be lower than a preset noise threshold value, automatically collecting the voice signal of the user at the moment as a low-noise environmental voice sample; carrying out phoneme-level segmentation on the voice sample and extracting a time domain envelope to establish a target envelope database; during real-time processing, envelope differences are compared through a dynamic time warping algorithm, local gain correction parameters are generated for dynamic gain adjustment, and enhanced real-time voice signals are obtained. According to the method, the problems of voice distortion and insufficient individuation of a traditional noise reduction method are solved.
Owner:MINAMI ACOUSTICS LTD

A controllable causal two-way speech enhancement method and system

The present invention belongs to the field of speech enhancement technology, and specifically relates to a controllable causal two-way speech enhancement method and system. The method comprises: taking noisy speech collected in an actual scene as input, performing a short-time Fourier transform, and obtaining a complex spectrum including an amplitude spectrum and a phase spectrum; constructing a streaming encoder, taking the complex spectrum as input, and extracting local correlations of the noisy speech in the time and frequency dimensions; constructing a controllable causal two-way module, modeling long-term dependencies in the time and frequency dimensions, and linearly modulating channel features through control parameters; constructing a streaming amplitude / phase decoder, reconstructing the amplitude spectrum and phase spectrum of the enhanced speech; and combining the reconstructed amplitude spectrum and phase spectrum with an inverse short-time Fourier transform to obtain the waveform of the enhanced speech. The present invention achieves real-time adjustment of the output characteristics of the speech enhancement network, and can timely balance the noise residual and speech distortion levels according to the listener's needs during the speech enhancement process.
Owner:SHANDONG UNIV OF TECH

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium, relates to the technical field of artificial intelligence, and can be applied to a medical speech generation scene or a financial speech generation scene. Through a semantic token path, the modeling capability of voice on a deep semantic structure is enhanced, and the continuity and definition of semantic expression are ensured; and acoustic details related to rhythm, timbre and emotion are reserved through the text condition and the voice prompt path, so that the naturalness and expressive force of the voice can be improved. And an adjustable weight or a self-adaptive mechanism is introduced in the fusion process, so that the system can realize smooth adjustment between semantic controllability and voice naturalness according to different application requirements, thereby improving the flexibility and the application range of the whole system. According to the method, the problems of voice distortion, emotion weakening or single expression caused by information loss of a single path are effectively reduced, and the method has more stable and superior synthesis performance in complex scenes such as cross-speaker, cross-emotion and dialogue generation.
Owner:PING AN TECH (SHENZHEN) CO LTD

Non-invasive methods for enhancing speech distortion suppression for robust speech recognition

This invention discloses a non-intrusive method for enhancing speech distortion suppression for robust speech recognition. The method includes the following steps: S1: Input the original complex spectrum and the enhanced complex spectrum; S2: Obtain the distortion suppression coefficient based on the input in step S1; S3: Apply the distortion suppression coefficient to a distortion suppression interpolation algorithm to obtain the output corrected spectrum. This invention achieves low computational complexity and compatibility with existing streaming and non-streaming speech enhancement models by using a non-intrusive front-end and back-end bridging module; the enhancement model requires a small amount of training data and can quickly adapt to a small amount of labeled data; it does not change the output signal of the enhancement model, effectively maintaining the auditory gain of different enhancement algorithms for different aspects of enhanced speech.
Owner:SHANGHAI JIAOTONG UNIV

A deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling

PendingCN122314000ASaliency mapPESQ
This invention provides a deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling, comprising: preprocessing the acquired near-end microphone signal and far-end reference signal to obtain a complex spectrum with uniform time-frequency resolution; constructing an echo cancellation model based on the complex spectrum using a cross-modal gating attention mechanism, a conditional Transformer, and a multi-branch decoder; and performing echo cancellation on the newly acquired mixed speech signal based on the echo cancellation model to obtain enhanced near-end speech. This method can still learn echo saliency maps through pseudo-labels / weak supervision in unlabeled scenarios, and inference only requires a single forward computation, thereby improving PESQ / STOI and ERLE and reducing speech distortion.
Owner:ANHUI UNIV

Speech enhancement network generation method, speech enhancement method and speech enhancement device

The invention relates to a speech enhancement network generation method, a speech enhancement method and a speech enhancement device, and the method comprises the steps: inputting a sample mixed signal and a sample noise signal into an initial speech enhancement network for speech enhancement processing, and obtaining a first predicted pure speech signal corresponding to the sample mixed signal, and a second prediction pure voice signal corresponding to the sample noise signal. Performing mute fragment recognition on the first predicted pure voice signal to obtain a predicted mute signal; and according to the difference between the first predicted pure voice signal and the first pure voice tag, the difference between the predicted mute signal and the mute tag, and the difference between the second predicted pure voice signal and the second pure voice tag, performing voice enhancement training on the initial voice enhancement network to obtain a target voice enhancement network. According to the invention, in the speech enhancement processing process, speech distortion and noise suppression can be balanced, and speech damage during noise suppression can be effectively reduced.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Children inquiry method and device based on large language model, medium and program product

The invention discloses a children inquiry method and device based on a large language model, a medium and a program product, relates to the field of intelligent auxiliary medical treatment, and aims to solve the problems of voice distortion and inaccurate positioning caused by the fact that a single-channel noise reduction method cannot give consideration to voice quality and spatial information in an intelligent inquiry process. The method comprises the steps of performing short-time Fourier transform on inquiry voice data to obtain spatial information of the inquiry voice data to obtain a multi-channel time-frequency mask, obtaining a voice feature matrix based on the multi-channel time-frequency mask and a short-time Fourier transform coefficient, generating a generalized mutual information entropy value through kernel density estimation, quantifying direction consistency between channels, and obtaining a multi-channel mutual information entropy value. Noise reduction can be ensured, spatial information is effectively reserved, and the efficiency of intelligent inquiry is effectively improved.
Owner:XIAO ER FANG HEALTH TECH (BEIJING) CO LTD