Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

522 results about "Speech enhancement" patented technology

Speech enhancement aims to improve speech quality by using various algorithms. The objective of enhancement is improvement in intelligibility and/or overall perceptual quality of degraded speech signal using audio signal processing techniques.

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Voice enhancement method and device based on noise perception, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Hearing aid intelligent noise reduction and human voice enhancement technology based on electroencephalogram signals

The invention relates to a hearing aid intelligent noise reduction and human voice enhancement system based on electroencephalogram signals, and belongs to the field of biomedical engineering and acoustic signal processing. The system comprises an electroencephalogram signal acquisition module, a multi-channel acoustic sensor array, an embedded neural signal processor, an adaptive beam forming module, a dynamic speech enhancement engine and a dual-mode output device, and constructs electroencephalogram-acoustics joint features by extracting an alpha / theta wave power ratio, a P300 component and auditory cortical Gamma phase synchronism. A deep network is driven to separate target voice, a wave beam direction and a frequency response curve are dynamically adjusted based on neural feedback, a closed-loop calibration unit is innovatively adopted, gain is reversely adjusted according to N1-P2 wave amplitude, heart rate variability and eye movement data are fused to optimize decisions, and when the signal-to-noise ratio is-5dB, the voice recognition rate reaches 89%, the auditory fatigue is reduced by 37%, and the decision conflict rate is smaller than 6%. The defects of attention blind area, noise separation failure and physiological adaptation of a traditional hearing aid are overcome. The system is suitable for the fields of hearing impairment rehabilitation, special communication and intelligent cabins.
Owner:MAXSON GLOBAL GROUP INC

Pilot earphone hearing protection method and system based on voice recognition compensation

The invention relates to the field of aviation voice signal processing, and discloses a voice recognition compensation pilot earphone hearing protection method and system, and the method comprises the following steps: collecting multi-modal data, and separating a sound source through tensor decomposition; inferring a pilot state by using a dynamic network; performing context recognition and semantic evaluation on the attention target voice; and finally, dynamically modulating the sound field based on deep reinforcement learning, enhancing the voice in a personalized manner, and outputting after noise suppression. The system comprises a multi-mode perception data acquisition module, a sound source decoupling module, a pilot state inference module, a voice processing and semantic evaluation module and a sound field modulation and output module. According to the invention, high-fidelity speech extraction is realized through multi-modal perception and tensor decomposition; evaluating priority key information in combination with attention and semantics; and deep reinforcement learning and model prediction control are adopted to dynamically optimize the sound field, so that the voice recognition accuracy and the pilot information acquisition efficiency are improved.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

Implementation method and device of multi-channel voiceprint recognition system

The invention relates to the technical field of voice recognition, in particular to an implementation method and device of a multi-channel voiceprint recognition system, and the implementation method comprises the steps of multi-channel data acquisition and synchronization, signal preprocessing and enhancement, feature extraction and fusion, model training, real-time deployment and adaptive optimization. Compared with the problems that a traditional multichannel voiceprint recognition system depends on a fixed beam forming algorithm and an independent clock synchronization module, the synchronization error is large, manual parameter adjustment is needed for noise suppression, and generalization is poor, hardware-level clock synchronization is achieved through a PTP protocol, and the accuracy of noise suppression is improved. The method combines an end-to-end neural network to automatically learn noise distribution and a sound source space position, dynamically generates a beam forming weight, can improve the voice quality in a complex noise scene without manual intervention, remarkably reduces the interference of a synchronization error on sound source positioning, and enables the precision and stability of far-field voice enhancement to reach a new level.
Owner:MINAMI ACOUSTICS LTD

Speech enhancement method and system based on conditional stream matching and vocoder

The invention discloses a speech enhancement method and system based on conditional stream matching and a vocoder, and the method comprises the following steps: S1, constructing a Mel spectrum extraction module which is used for converting an input noisy speech into a noisy Mel spectrum; s2, constructing a condition flow matching noise reduction module which is used for processing the noisy Mel spectrum obtained in the step S1 and outputting an enhanced Mel spectrum; and S3, constructing a neural network vocoder module which is used for restoring the enhanced Mel spectrum obtained in the step S2 into a time domain voice waveform so as to obtain an enhanced voice signal. The speech enhancement method combining conditional stream matching and the vocoder is proposed for the first time, the conditional stream matching method is innovatively introduced into the Mel-frequency spectrum domain, an end-to-end speech enhancement system is constructed, and a complete processing flow from noise speech input to high-quality speech signal output is realized.
Owner:HANGZHOU DIANZI UNIV

Illegal voice detection method, device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the fields of finance, medical treatment and the like, and provides a violation voice detection method, device, equipment and medium, and the method comprises the following steps: carrying out suspicious fragment detection on an input voice, and outputting a low-definition voice fragment; performing voice enhancement processing on the low-definition voice segment to obtain an optimized audio; converting the optimized audio into text data through voice recognition; and based on the text data, judging whether violation guide content exists or not and generating a violation analysis report. Compared with the prior art, real-time monitoring and accurate recognition of the violation guiding behaviors of the salesman are achieved, a whole-process closed loop covering detection, enhancement, conversion and judgment is formed, and it is ensured that the business process is compliant.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method for training speech enhancement network, method for enhancing speech, and electronic device

PendingUS20250391419A1Speech analysisNoiseSpeech classification
A method for training a speech enhancement network, performed by an electronic device, includes: acquiring a first clean speech sample and a noise sample, and mixing them to generate a noisy speech sample; performing noise reduction on the noisy speech sample based on the speech enhancement network to obtain an enhanced speech sample; framing the enhanced speech sample into a plurality of enhanced speech frames, classifying speech effectiveness of the enhanced speech frames, and generating a first effectiveness distribution based on classification results of the enhanced speech frames; and determining a noise reduction accuracy based on the enhanced speech sample and the first clean speech sample, determining a speech classification accuracy based on the first effectiveness distribution, determining a speech enhancement accuracy based on the noise reduction accuracy and the speech classification accuracy, and training the speech enhancement network based on the speech enhancement accuracy.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Speech enhancement and high-precision recognition method and system in complex environment

PendingCN121641016ASpeech recognitionSpectral density estimationNerve network
The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

A method of mixed voice processing, an electronic device, a computer readable medium

The present application relates to the technical field of speech processing, and particularly relates to a mixed speech processing method, an electronic device and a computer readable medium. The method comprises the following steps: collecting mixed speech and environmental influence parameters; performing bionics frequency domain analysis on the mixed speech to obtain low-frequency attenuation compensation feature data; performing multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental influence parameters to generate channel distortion data; performing time domain-frequency domain joint deconvolution processing on the mixed speech by using the channel distortion data to generate direct sound components and reflected sound components; performing adversarial training based on the direct sound components and the reflected sound components to generate anti-multipath speech enhancement data; and constructing a dynamic frequency compensation filter based on preset environmental acoustic characteristics. The present application improves the output quality of mixed speech through multi-stage signal processing, frequency compensation and real-time optimization technology.
Owner:GUANGZHOU ZHIYU CLOUD NETWORK COMMUNICATIONS CO LTD

Voice separation method, electronic device, storage medium and computer program product

ActiveCN120581022ASpeech analysisField separationTesting Methods
The invention discloses a voice separation method, electronic equipment, a storage medium and a computer program product, and relates to the technical field of signal processing, and the method comprises the steps: carrying out the beam forming of a microphone array signal through employing a preset near-field speaker direction as a voice enhancement direction, and obtaining a first near-field speaker signal; performing beam forming processing on the microphone array signal by taking a preset far-field speaker direction as a voice enhancement direction to obtain a first far-field speaker signal; inputting the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal into a preset near-field separation model for processing to obtain a second near-field speaker signal; and inputting the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal into a preset far-field separation model for processing to obtain a second far-field speaker signal. According to the invention, the far and near field signals are preliminarily separated, so that the computing power resource consumption of a voice separation algorithm is reduced.
Owner:GOERTEK INC

Audio processing method and device, medium and electronic equipment

The invention belongs to the technical field of artificial intelligence, and particularly relates to an audio processing method, an audio processing device, a computer readable medium, electronic equipment and a computer program product. The audio processing method comprises the following steps: performing feature extraction on an audio signal to obtain frequency domain features of a plurality of audio frames; the audio signal carries a voice signal and a noise signal; serialized modeling is carried out on the frequency domain features of the audio frames, intra-frame sequence features are obtained, and the intra-frame sequence features are used for representing the frequency sequence relation among a plurality of frequency points in the single audio frame; serialization modeling is carried out on the intra-frame sequence features of the multiple audio frames to obtain inter-frame sequence features, and the inter-frame sequence features are used for representing the time sequence relation among the multiple audio frames; and performing signal modulation on the audio signal according to the inter-frame sequence features to obtain a voice signal after noise signal removal. According to the embodiment of the invention, the speech enhancement noise reduction effect can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio coding and decoding method and device based on stream matching, equipment and medium

The invention relates to the technical field of voice processing, and discloses an audio encoding and decoding method and device based on stream matching, equipment and a medium, which can be applied to financial service or digital medical service scenes, and can be used for obtaining an initial voice waveform corresponding to a target audio according to a target audio receiving instruction in response to the target audio receiving instruction; inputting the initial voice waveform into a preset encoder for encoding to obtain high-dimensional voice information; inputting the high-dimensional voice information into a preset quantizer for processing to obtain discrete voice information; inputting the discrete voice information into a preset decoder for decoding to obtain a first voice waveform; and performing voice enhancement on the first voice waveform by using the target stream matching model to generate an enhanced target voice, thereby realizing audio reconstruction with high perception quality at a low bit rate. And meanwhile, through the target stream matching model, the calculation overhead can be remarkably reduced, the reasoning speed is improved, and the method is suitable for full-band universal audio and is high in universality.
Owner:PING AN TECH (SHENZHEN) CO LTD

Time-frequency domain speech enhancement method based on KAN channel attention

The invention relates to the technical field of speech enhancement, in particular to a time-frequency domain speech enhancement method based on KAN channel attention, which comprises the following steps: processing a speech data set to obtain frequency domain representation; extracting local features in the frequency domain representation through an encoder to obtain output features; the output features are input into a TF-Transform block, noise components are recognized and suppressed, and output components are obtained; inputting the output component into a KAN-based channel attention module, sequentially obtaining a channel feature map and a spatial feature map, and successively multiplying the output component by the channel feature map and the spatial feature map and introducing jump connection to obtain refined features; and obtaining a time domain voice signal based on refined feature recovery, and processing the recovered time domain voice signal through an amplitude decoder and a phase decoder to correspondingly obtain an amplitude spectrum and a phase spectrum. And through a KAN-based channel attention module, the denoising and reconstruction performance of sparse and sensitive high-frequency components in the voice is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Single-microphone acoustic echo and noise suppression

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, the AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on the microphone signal and the reference audio signal and may suppress an echo component and a noise component of the microphone signal based on the determined set of masks.
Owner:SYNAPTICS INC

Far-field single-channel speech enhancement method

The invention relates to the technical field of speech enhancement, in particular to a far-field single-channel speech enhancement method based on an MFSE (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error)). The method comprises the following steps: step 1, processing a far-field voice signal to obtain a complex spectrogram of a noise voice signal; 2, inputting the compressed complex spectrogram into a feature encoder, processing the output of the complex spectrogram by N MamAttention blocks, and then sending the processed complex spectrogram into an amplitude mask decoder and a phase decoder to respectively predict a clean compressed amplitude mask and a phase spectrum; step 3, preheating and training the MamAttention model, and performing supervised confrontation training by taking the MamAttention model as a generator and the multi-resolution discriminator as a discriminator; and step 4, inputting test voice into the trained model to realize far-field single-channel voice enhancement. According to the method, the supervised adversarial training strategy and the MamAttention model are combined, so that the problems of signal attenuation, noise and reverberation interference in far-field voice are effectively solved, and the voice quality is remarkably improved in a scene that the distance of a loudspeaker exceeds 5 meters.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Voice conversation interaction method and system for industrial equipment

The invention provides a voice dialogue interaction method and system for industrial equipment, and relates to the technical field of man-machine interaction, and the method comprises the following steps: obtaining a voice signal of a user, and carrying out the voice enhancement processing of the voice signal of the user through an incremental adaptive filtering algorithm, and obtaining an enhanced voice; according to the enhanced voice, using an information gain transfer learning method to identify an interaction intention of the user, and generating a to-be-interacted voice based on the interaction intention of the user; based on the enhanced voice, utilizing a time delay estimation method to identify an interaction position of the user, and performing position prediction on the position of the user; according to the position prediction result, an industrial equipment horn output strategy is constructed; and outputting the to-be-interacted voice based on the industrial equipment loudspeaker output strategy so as to realize voice dialogue interaction between the user and the industrial equipment. According to the invention, convenience, high efficiency and accuracy of interaction between the user and the equipment in an industrial scene are greatly improved, and intelligent development of industrial production is facilitated.
Owner:CHINA APPLIED TECH CO LTD

Voice conversion authentic identification method and system based on knowledge distillation alignment

The invention discloses a voice conversion authentic identification method and system based on knowledge distillation alignment, and is applied to the technical field of voice authentic identification. The method comprises the following steps that a double-branch model used for voice authentic identification is constructed, pure voice serves as input of a teacher branch, noisy voice serves as input of a student branch, and the teacher branch and the student branch share a feature extraction network of the same structure; applying a speech enhancement technology at the front end of a student branch to generate an enhanced speech signal; deep features are extracted, and alignment of the deep features in a hidden space is restrained through a knowledge distillation loss function; carrying out dynamic weight fusion on the aligned deep features through a fusion weight; and training a classifier, performing joint optimization in combination with classification loss and knowledge distillation loss, and outputting a voice authentic identification result. According to the method, distribution alignment of pure and noise features is realized through knowledge distillation, forged traces in voice conversion are effectively recognized, and high detection precision is still kept in complex noise and unknown attack scenes.
Owner:ZHEJIANG UNIV

Coal conveying production voice system control platform

The invention discloses a coal conveying production voice system control platform, which comprises a voice acquisition module, a voice recognition module, an instruction analysis module, a control logic module, an equipment execution module, a feedback module, a communication module and a safety protection module, noise reduction and enhancement processing are carried out; the voice recognition module is connected with the voice acquisition module, and the voice recognition module is used for converting an acquired voice instruction into a text instruction. When the scheme is implemented, the equipment is quickly controlled through the voice instruction, so that manual operation steps are reduced, the response speed is increased, and the operation efficiency is improved; by combining voiceprint recognition and authority management, misoperation or illegal instructions are prevented, so that the security of the system during implementation is improved; and by adopting noise reduction and voice enhancement technologies, the recognition accuracy in a high-noise environment is ensured, so that the system can adapt to various working conditions.
Owner:HUANENG JILIN POWER GENERATION CO LTD CHANGCHUN THERMAL POWER PLANT

Multi-mode anti-interference communication method and system based on lip language recognition

The invention discloses a multi-mode anti-interference communication method and system based on lip language recognition, and belongs to the technical field of communication equipment. The method comprises the following steps: acquiring a face lip video stream and an audio signal; in response to the conventional mode trigger signal, feature extraction is performed on the lip video stream and the audio signal, and extraction results are fused to generate a fused feature vector; performing voice enhancement on the fusion feature vector in combination with lip motion information, and outputting an audio enhancement signal; and in response to the silent communication mode trigger signal, performing lip language recognition based on the face lip video stream to obtain a lip language recognition text, and converting the lip language recognition text into voice. Clear and stable communication in an ultra-strong noise environment can be realized by combining two types of modal information, and the problem that the communication quality is influenced by an existing high-noise environment and the requirement for mobile silent communication in a special scene are solved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Doll intelligent speech recognition method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence and voice recognition, and discloses a doll intelligent voice recognition method and system based on artificial intelligence, and the method comprises the steps: collecting audio through a multi-microphone array; locally executing speech enhancement, acoustic feature extraction and lightweight wake-up word detection; local speech recognition is started, semantic confidence is evaluated, and desensitized data are uploaded to a cloud only when the confidence is insufficient; and the cloud end performs secondary recognition of context perception by using the large model in combination with the language habits of the children, and returns an optimization result. The system comprises an audio acquisition module, a local processing module, a cloud collaboration module and a state management module. According to the invention, through an end-cloud collaborative architecture, the identification robustness and the interaction intelligence level in a complex environment are significantly improved while privacy and low delay are guaranteed.
Owner:ZHUHAI ZHIHUI HUACHUANG TECHNOLOGY CO LTD

Single-channel non-real-time speech enhancement method based on Taylor model, computer storage medium and product

The invention discloses a single-channel non-real-time speech enhancement method based on a Taylor model, a computer storage medium and a product. The method comprises the following steps: S1, collecting and preprocessing noisy speech data; s2, constructing a novel TaylorSENet neural network model, wherein the model comprises a zero-order block, a high-order block and a self-adaptive order selection module; s3, enhancing the noise-containing voice data preprocessed in the step S1 by using the model in the step S2, and outputting an enhanced complex spectrum; and S4, the novel TaylorSENet neural network model in the S2 is trained until the learning rate is kept unchanged. According to the invention, by constructing the novel TaylorSENet neural network model and through multi-scale feature coding and self-adaptive calculation order selection, the high speech enhancement performance is maintained, the calculation complexity and the reasoning time are significantly reduced, and calculation resources can be intelligently allocated according to the input signal-to-noise ratio.
Owner:TIANJIN AGRICULTURE COLLEGE

Loudspeaker speech enhancement method and device, electronic equipment and storage medium

The invention provides a loudspeaker speech enhancement method and device, electronic equipment and a storage medium, and belongs to the technical field of speech signal processing, and the method comprises the steps: obtaining an initial spectrum signal and an initial power spectrum signal corresponding to an initial audio signal; inputting the initial power spectrum signal into a feedback elimination network model to obtain a target frequency domain signal; and obtaining a target audio signal based on the target frequency domain signal. The method comprises the following steps: performing mask adjustment on the frequency domain amplitude of an initial frequency spectrum signal by using a real number signal mask to generate a frequency domain estimation signal, and performing mask adjustment on the frequency domain amplitude and phase of the initial frequency spectrum signal by using a complex number signal mask to obtain a frequency domain enhancement signal. And finally, a target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal, so that the output target frequency domain signal has higher fidelity in the aspects of amplitude and phase, environmental noise and loudspeaker feedback signals can be removed to the greatest extent, and the squeal phenomenon is effectively inhibited.
Owner:IFLYTEK CO LTD

Single-channel speech enhancement method and device based on time frequency-phase joint perception and CBAM attention mechanism

The invention discloses a single-channel speech enhancement method and equipment based on time-frequency-phase joint perception and a CBAM attention mechanism. The method comprises the following steps: acquiring a noise-added speech signal training data set; a single-channel speech enhancement network based on time frequency-phase joint perception and a CBAM attention mechanism is constructed, and the single-channel speech enhancement network specifically comprises a short-time Fourier transform module, an encoder, a double-path circulation network time sequence modeling unit, a first CBAM attention module, a second CBAM attention module, an amplitude decoder, a phase decoder and a signal reconstruction module; inputting a training data set into the single-channel speech enhancement network for network training; and inputting a noise-containing and reverberation-containing single-channel test voice signal to be enhanced into the trained single-channel voice enhancement network to obtain an enhanced single-channel voice signal. The enhancement effect is better, and the number of parameters is smaller.
Owner:SOUTHEAST UNIV

Microscale embedded system speech enhancement method and device based on neural network, electronic equipment and storage medium

The invention discloses a micro-magnitude embedded system speech enhancement method and device based on a neural network, electronic equipment and a storage medium. The method comprises the following steps: acquiring power spectrum data of voice data to be processed, and calculating a filter bank coefficient for a preset filter bank; the power spectrum data and the filter bank coefficient are input into a pre-training neural network, an initial Wiener filter coefficient is obtained, and the pre-training neural network is obtained through training in a ratio mask mode and a signal approximation mode; performing interpolation processing on the initial Wiener filter coefficient to obtain a target Wiener filter coefficient with the same number as the frequency points of the to-be-processed voice data; and processing the to-be-processed voice data through the target Wiener filter coefficient to obtain voice data after noise reduction processing. According to the technical scheme, the noise reduction effect is improved while the noise reduction calculation amount is reduced.
Owner:SHENZHEN JIAYZ PHOTO IND LTD

Environmental sound classification and noise reduction method and system for intelligent Bluetooth hearing-aid earphone

The invention discloses an environmental sound classification and noise reduction method and system for an intelligent Bluetooth hearing-aid earphone. The method comprises the following steps: S1, collecting background noise original audio signals and real-time binaural sound field signals in a typical scene; s2, Mel spectrum features and auditory perception features are extracted from each typical scene, fusion features are calculated, dimension compression is carried out, and a scene feature vector library is obtained; s3, extracting a live sound field feature, and extracting a scene feature vector matched with the live sound field in the scene feature vector library; s4, calculating a sound source direction correction coefficient, and calculating a corrected sound source direction; s5, calculating cross-modal interaction characteristics, then calculating noise suppression intensity and sound source gain intensity, and generating an audio after noise reduction; and S6, generating a synchronous stereo audio signal. The method can solve the problem that the traditional method is difficult to distinguish different types of background noise, so that the voice of a dialogue is inhibited by mistake, and the voice enhancement effect is poor.
Owner:HUNAN DINO INTELLIGENT TECHNOLOGY CO LTD

Brain-controlled speech enhancement method and system based on electroencephalogram signal neural decoding

The invention discloses a brain-controlled speech enhancement method and system based on electroencephalogram signal neural decoding. The method comprises the following steps: synchronously acquiring multi-channel speech signals and EEG signals; compact features related to auditory attention are extracted from the EEG signals through the EEG coding network and serve as auditory cognitive state representation, and an auditory consensus field continuously evolved on the slow time scale is constructed; performing time-frequency modeling on the multi-channel voice signals based on deep learning, and constructing a deep voice time-frequency analysis network; and performing voice time-frequency spectrum gain modulation and neural beam forming weight gain modulation on the signal by using an auditory consensus field, converting the signal back to a time domain, and outputting a finally enhanced target voice signal. According to the invention, by constructing an auditory consensus field driven by the EEG, neural modulation of a frequency spectrum characteristic flow of a deep speech time-frequency analysis network and a space beam forming weight of a neural network beam forming module is realized, so that speech enhancement in a complex acoustic environment and robustness and stability of a system are improved.
Owner:HUIZHOU UNIV

WebRTC speech enhancement system and method based on multi-modal large model

The invention discloses a WebRTC voice enhancement system and method based on a multi-modal large model, and belongs to the field of artificial intelligence and real-time communication cross technology, and the system comprises an audio and video collection module which is used for synchronously collecting original voice signals and corresponding video image data of a user side through a WebRTC protocol stack, high-precision alignment is realized through timestamp marking and a cache mechanism; the multi-modal feature extraction module is used for extracting audio features from the voice signals, extracting visual lip movement features from the video images and generating text semantic features through a voice recognition engine; a noise matching and updating module; a multi-modal semantic perception enhancement module; an audio reconstruction module; and a WebRTC integration module. According to the invention, high-precision noise suppression and millisecond-level delay are realized while the voice semantic integrity is guaranteed, and the requirements of industrial inspection, vehicle-mounted communication and other scenes on high fidelity, low delay and strong robustness are met.
Owner:浪潮智慧城市科技有限公司

Bandwidth extension and speech enhancement of audio

There is provided a system, apparatus and a method for audio processing. The operations include obtaining an input audio waveform, obtaining a mel-spectrogram by performing a short-time Fourier transform (STFT) operation on the input audio waveform, obtaining an updated mel-spectrogram by at least one or removing noise from the mel-spectrogram or restoring high frequency components by applying two-dimensional Unet convolutional blocks to the mel-spectrogram, converting the updated mel-spectrogram to a converted audio waveform in a waveform domain, correcting the converted audio waveform in a time domain, correcting the converted audio waveform in a frequency domain to remove artifacts or noise, processing the corrected audio waveform corrected in the time domain and corrected in the frequency domain with an one-dimensional convolutional layer, and outputting the processed audio waveform in the time domain and in the frequency domain.
Owner:SAMSUNG ELECTRONICS CO LTD

Bone conduction speech enhancement method based on double-flow self-attention fusion network

The invention provides a bone conduction speech enhancement method based on a double-flow self-attention fusion network. The bone conduction speech enhancement method comprises the following steps: acquiring and preprocessing a bone conduction speech signal to be processed; a double-flow self-attention fusion network is constructed, the double-flow self-attention fusion network comprises a double-flow encoder, a cross-modal feature fusion module, a multi-scale decoder and a waveform reconstruction module which are connected in sequence, the double-flow self-attention fusion network is trained, a trained model is obtained, and the performance of the model under different conditions is evaluated. According to the method, a double-flow structure is used for processing time domain information and frequency domain information respectively, a self-attention mechanism is introduced, so that the model can capture a long-distance dependency relationship in each mode, and the feature quality is improved. According to the invention, the time-frequency information can be better integrated through the cross-modal feature fusion module. According to the invention, the multi-scale loss function supervises the output of different levels of the decoder, so that the network can learn effective spectrogram representation at different resolutions, and a finer spectrum structure can be recovered.
Owner:TIANJIN UNIV +1