Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

308 results about "Speech enhancement" patented technology

Speech enhancement aims to improve speech quality by using various algorithms. The objective of enhancement is improvement in intelligibility and/or overall perceptual quality of degraded speech signal using audio signal processing techniques.

Hearing aid intelligent noise reduction and human voice enhancement technology based on electroencephalogram signals

The invention relates to a hearing aid intelligent noise reduction and human voice enhancement system based on electroencephalogram signals, and belongs to the field of biomedical engineering and acoustic signal processing. The system comprises an electroencephalogram signal acquisition module, a multi-channel acoustic sensor array, an embedded neural signal processor, an adaptive beam forming module, a dynamic speech enhancement engine and a dual-mode output device, and constructs electroencephalogram-acoustics joint features by extracting an alpha / theta wave power ratio, a P300 component and auditory cortical Gamma phase synchronism. A deep network is driven to separate target voice, a wave beam direction and a frequency response curve are dynamically adjusted based on neural feedback, a closed-loop calibration unit is innovatively adopted, gain is reversely adjusted according to N1-P2 wave amplitude, heart rate variability and eye movement data are fused to optimize decisions, and when the signal-to-noise ratio is-5dB, the voice recognition rate reaches 89%, the auditory fatigue is reduced by 37%, and the decision conflict rate is smaller than 6%. The defects of attention blind area, noise separation failure and physiological adaptation of a traditional hearing aid are overcome. The system is suitable for the fields of hearing impairment rehabilitation, special communication and intelligent cabins.
Owner:MAXSON GLOBAL GROUP INC

Method for training speech enhancement network, method for enhancing speech, and electronic device

PendingUS20250391419A1Speech analysisNoiseSpeech classification
A method for training a speech enhancement network, performed by an electronic device, includes: acquiring a first clean speech sample and a noise sample, and mixing them to generate a noisy speech sample; performing noise reduction on the noisy speech sample based on the speech enhancement network to obtain an enhanced speech sample; framing the enhanced speech sample into a plurality of enhanced speech frames, classifying speech effectiveness of the enhanced speech frames, and generating a first effectiveness distribution based on classification results of the enhanced speech frames; and determining a noise reduction accuracy based on the enhanced speech sample and the first clean speech sample, determining a speech classification accuracy based on the first effectiveness distribution, determining a speech enhancement accuracy based on the noise reduction accuracy and the speech classification accuracy, and training the speech enhancement network based on the speech enhancement accuracy.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Speech enhancement and high-precision recognition method and system in complex environment

PendingCN121641016ASpeech recognitionSpectral density estimationNerve network
The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

Time-frequency domain speech enhancement method based on KAN channel attention

The invention relates to the technical field of speech enhancement, in particular to a time-frequency domain speech enhancement method based on KAN channel attention, which comprises the following steps: processing a speech data set to obtain frequency domain representation; extracting local features in the frequency domain representation through an encoder to obtain output features; the output features are input into a TF-Transform block, noise components are recognized and suppressed, and output components are obtained; inputting the output component into a KAN-based channel attention module, sequentially obtaining a channel feature map and a spatial feature map, and successively multiplying the output component by the channel feature map and the spatial feature map and introducing jump connection to obtain refined features; and obtaining a time domain voice signal based on refined feature recovery, and processing the recovered time domain voice signal through an amplitude decoder and a phase decoder to correspondingly obtain an amplitude spectrum and a phase spectrum. And through a KAN-based channel attention module, the denoising and reconstruction performance of sparse and sensitive high-frequency components in the voice is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Single-microphone acoustic echo and noise suppression

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, the AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on the microphone signal and the reference audio signal and may suppress an echo component and a noise component of the microphone signal based on the determined set of masks.
Owner:SYNAPTICS INC

Far-field single-channel speech enhancement method

The invention relates to the technical field of speech enhancement, in particular to a far-field single-channel speech enhancement method based on an MFSE (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error)). The method comprises the following steps: step 1, processing a far-field voice signal to obtain a complex spectrogram of a noise voice signal; 2, inputting the compressed complex spectrogram into a feature encoder, processing the output of the complex spectrogram by N MamAttention blocks, and then sending the processed complex spectrogram into an amplitude mask decoder and a phase decoder to respectively predict a clean compressed amplitude mask and a phase spectrum; step 3, preheating and training the MamAttention model, and performing supervised confrontation training by taking the MamAttention model as a generator and the multi-resolution discriminator as a discriminator; and step 4, inputting test voice into the trained model to realize far-field single-channel voice enhancement. According to the method, the supervised adversarial training strategy and the MamAttention model are combined, so that the problems of signal attenuation, noise and reverberation interference in far-field voice are effectively solved, and the voice quality is remarkably improved in a scene that the distance of a loudspeaker exceeds 5 meters.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Doll intelligent speech recognition method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence and voice recognition, and discloses a doll intelligent voice recognition method and system based on artificial intelligence, and the method comprises the steps: collecting audio through a multi-microphone array; locally executing speech enhancement, acoustic feature extraction and lightweight wake-up word detection; local speech recognition is started, semantic confidence is evaluated, and desensitized data are uploaded to a cloud only when the confidence is insufficient; and the cloud end performs secondary recognition of context perception by using the large model in combination with the language habits of the children, and returns an optimization result. The system comprises an audio acquisition module, a local processing module, a cloud collaboration module and a state management module. According to the invention, through an end-cloud collaborative architecture, the identification robustness and the interaction intelligence level in a complex environment are significantly improved while privacy and low delay are guaranteed.
Owner:ZHUHAI ZHIHUI HUACHUANG TECHNOLOGY CO LTD

Single-channel non-real-time speech enhancement method based on Taylor model, computer storage medium and product

The invention discloses a single-channel non-real-time speech enhancement method based on a Taylor model, a computer storage medium and a product. The method comprises the following steps: S1, collecting and preprocessing noisy speech data; s2, constructing a novel TaylorSENet neural network model, wherein the model comprises a zero-order block, a high-order block and a self-adaptive order selection module; s3, enhancing the noise-containing voice data preprocessed in the step S1 by using the model in the step S2, and outputting an enhanced complex spectrum; and S4, the novel TaylorSENet neural network model in the S2 is trained until the learning rate is kept unchanged. According to the invention, by constructing the novel TaylorSENet neural network model and through multi-scale feature coding and self-adaptive calculation order selection, the high speech enhancement performance is maintained, the calculation complexity and the reasoning time are significantly reduced, and calculation resources can be intelligently allocated according to the input signal-to-noise ratio.
Owner:TIANJIN AGRICULTURE COLLEGE

Brain-controlled speech enhancement method and system based on electroencephalogram signal neural decoding

The invention discloses a brain-controlled speech enhancement method and system based on electroencephalogram signal neural decoding. The method comprises the following steps: synchronously acquiring multi-channel speech signals and EEG signals; compact features related to auditory attention are extracted from the EEG signals through the EEG coding network and serve as auditory cognitive state representation, and an auditory consensus field continuously evolved on the slow time scale is constructed; performing time-frequency modeling on the multi-channel voice signals based on deep learning, and constructing a deep voice time-frequency analysis network; and performing voice time-frequency spectrum gain modulation and neural beam forming weight gain modulation on the signal by using an auditory consensus field, converting the signal back to a time domain, and outputting a finally enhanced target voice signal. According to the invention, by constructing an auditory consensus field driven by the EEG, neural modulation of a frequency spectrum characteristic flow of a deep speech time-frequency analysis network and a space beam forming weight of a neural network beam forming module is realized, so that speech enhancement in a complex acoustic environment and robustness and stability of a system are improved.
Owner:HUIZHOU UNIV

Bone conduction speech enhancement method based on double-flow self-attention fusion network

The invention provides a bone conduction speech enhancement method based on a double-flow self-attention fusion network. The bone conduction speech enhancement method comprises the following steps: acquiring and preprocessing a bone conduction speech signal to be processed; a double-flow self-attention fusion network is constructed, the double-flow self-attention fusion network comprises a double-flow encoder, a cross-modal feature fusion module, a multi-scale decoder and a waveform reconstruction module which are connected in sequence, the double-flow self-attention fusion network is trained, a trained model is obtained, and the performance of the model under different conditions is evaluated. According to the method, a double-flow structure is used for processing time domain information and frequency domain information respectively, a self-attention mechanism is introduced, so that the model can capture a long-distance dependency relationship in each mode, and the feature quality is improved. According to the invention, the time-frequency information can be better integrated through the cross-modal feature fusion module. According to the invention, the multi-scale loss function supervises the output of different levels of the decoder, so that the network can learn effective spectrogram representation at different resolutions, and a finer spectrum structure can be recovered.
Owner:TIANJIN UNIV +1

Voice interaction method and device, electronic equipment and storage medium

The invention provides a voice interaction method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a real-time voice stream; under the condition that the wake-up word is detected from the real-time voice stream, voiceprint features of a target speaker are extracted from a voice segment, corresponding to the wake-up word, in the real-time voice stream; based on the correlation between the voiceprint feature and the voice feature of the real-time voice stream, performing voice enhancement on the real-time voice stream to obtain an enhanced voice stream of the target speaker; and performing voice interaction based on the enhanced voice stream of the target speaker. According to the method and device, the electronic equipment and the storage medium provided by the invention, under the condition that the wake-up word is detected, the voiceprint feature of the target speaker is extracted, and targeted voice enhancement is performed based on the voiceprint feature, so that the enhanced voice stream for the target speaker is accurately extracted from the complex real-time voice stream, and the voice enhancement efficiency is improved. The influence of interference voice on voice interaction is greatly reduced, and the accuracy and reliability of voice interaction are ensured.
Owner:XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD

Intelligent tea scale data management system based on voice recognition

The invention provides an intelligent tea scale data management system based on voice recognition, which comprises a front-end processing module, a voice enhancement module, a semantic recognition module and a weighing management module, and is characterized in that the front-end processing module is used for collecting an original voice signal, generating an initial noise model of the original voice signal, and outputting the initial noise model; carrying out adaptive denoising on the original voice signal based on the initial noise model to obtain a primary voice signal; the voice enhancement module is used for extracting administrator voiceprint features from the administrator voice and performing cross attention fusion and mask enhancement on the primary voice signal to obtain an enhanced voice signal; the semantic recognition module is used for performing semantic recognition and keyword matching on the enhanced voice signal to obtain recognized tea data; and the weighing management module is used for collecting real-time weighing data and carrying out price settlement and data storage on the identified tea data according to the real-time weighing data to obtain settled tea data.
Owner:HANGZHOU DASHANG INTELLIGENT TECH CO LTD

Noise isolation and target sound enhancement system based on specific voice pre-storage

The invention relates to the technical field of voice processing and enhancement, in particular to a noise isolation and target voice enhancement system based on specific voice pre-storage, and the system extracts pre-stored features from a pre-stored target voice database module through collecting environment audio signals in real time and carrying out frame segmentation and frequency domain transformation; adaptive modulation of amplitude and phase of each sub-band signal is realized by combining energy sensing sub-band modulation, then environmental noise energy is identified through dynamic noise isolation and multi-stage suppression is carried out, and finally continuous and clear target voice output is generated through enhanced fusion. The method can achieve the effective enhancement and environmental noise suppression of the target voice in a complex and changeable environment, improves the voice recognition precision and definition, supports the dynamic switching of multiple teachers and multiple classes, adapts to the target voiceprint change, and achieves the high-robustness voice enhancement and noise reduction in teaching, conference and multi-target scenes.
Owner:UNIV OF JINAN

Differential microphone array speech enhancement device and method based on reference point optimization

The invention provides a differential microphone array speech enhancement device and method based on reference point optimization, and relates to the technical field of speech signal processing. The device comprises a microphone array, a pre-processing module, a reference point selection module and a signal processing module. According to the invention, the reference point of the microphone array is improved from the first microphone in the prior art to the geometric center of the invention, the maximum value of the distance rm from the microphones in the microphone array to the reference point is reduced by half, the high frequency band is significantly reduced, and the approximation precision of Jacobi-Anger expansion is ensured. Through reference point optimization, the invention aims to realize a beam pattern with invariable frequency in a broadband range and combined optimization of DF and WNG. While the beam pattern approximation requirement and no distortion constraint are met, the white noise gain can be maximized, and the robustness is improved.
Owner:WUHAN UNIV

Mining intelligent voice and video interactive communication method and system

The invention relates to the technical field of communication, discloses a mining intelligent voice and video interactive communication method and system, and aims to solve the problems of poor audio and video signal quality, low bandwidth utilization rate, lack of semantic understanding and interaction lag in a complex mine environment. The method comprises the following steps: acquiring an original signal through an intrinsic safety type audio and video terminal; voice enhancement is carried out by adopting sound source orientation estimation and a deep complex network; reconstructing a high dynamic range video based on the physical imaging model; realizing audio and video semantic alignment and key event extraction by using a lightweight spatial-temporal feature fusion network; the bandwidth is dynamically allocated by combining the event confidence coefficient and the channel state, and high / common priority data is transmitted in a graded manner; and ground and underground real-time communication is realized through a bidirectional interaction channel. According to the invention, the signal availability, the bandwidth efficiency and the emergency response capability are remarkably improved, and high-stability, low-delay and high-fidelity communication is guaranteed.
Owner:JINAN HUAKE ELECTRICAL DEVICE

Speech enhancement

In accordance with implementations of the subject matter described herein, a solution for speech enhancement is proposed. In this solution, a target time-frequency representation at least indicating intensities of an input audio signal at different frequencies over time is obtained. The input audio signal comprises a speech component and a noise component. Frequency correlation information and time correlation information of the input audio signal is determined based on the target time-frequency representation. A target feature representation is generated based on the frequency correlation information, the time correlation information, and the target time-frequency representation. The target feature representation is for distinguishing the speech component and the noise component. An output audio signal is generated based on the target feature representation and the target time-frequency representation. The speech component is enhanced relative to the noise component in the output audio signal. In this way, the performance of speech enhancement can be improved.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Multi-modal speech enhancement method and device based on deep learning model

PendingCN121963735AImplement adaptive bindingAchieve natural bindingSpeech recognitionSound source locationSound sources
The invention discloses a multi-modal speech enhancement method and device based on a deep learning model, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining the head posture data, binaural audio signals and visual context information of a user in a virtual reality environment, coding the binaural audio signals into three-dimensional space acoustic features, and carrying out the coding of the three-dimensional space acoustic features; and extracting virtual sound source position features and lip motion features from the visual context information, inputting the features into an immersive fusion enhancement network, selectively enhancing or inhibiting acoustic features from different spatial directions, generating an enhanced audio stream, and outputting the enhanced audio stream through a binaural rendering engine. According to the method, the technical problems that in the prior art, due to the fact that multi-modal prior information cannot be effectively fused, voice enhancement lacks spatial selectivity, and an interference sound source irrelevant to vision is difficult to restrain are solved, and head posture dynamic attention and lip motion cross-modal constraint are fused; the technical effects of natural binding of auditory attention and a visual focus and effective suppression of an interference sound source are achieved.
Owner:SHUTIAN (HANGZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Dialect speech enhancement method and device based on cross-modal and confrontation verification

The invention discloses a dialect speech enhancement method and device based on cross-modal and confrontation verification, and belongs to the technical field of speech recognition and enhancement. According to the method, the dialect voice and the lip movement video are jointly modeled, so that the accuracy and the naturalness of a dialect generation task are remarkably improved; a set of generation-adversarial-feedback closed-loop enhancement framework is constructed, and a multi-dimensional adversarial verification mechanism is introduced, so that the model is self-evolved, and enhancement data is ensured to be diversified and accord with real use habits of dialects. The dialect knowledge graph and the voice generation model are creatively fused, double enhancement of semantic and cultural levels is achieved, the generated dialect voice content better conforms to the real language environment and cultural background of the dialect voice content, and the method is particularly suitable for protection and inheritance of endangered or small dialects. The method can be directly applied to subsequent dialect recognition, synthesis or protection and the like.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Intelligent glasses double-target speech enhancement method and system based on multi-microphone array, terminal equipment and medium

The invention discloses an intelligent glasses double-target speech enhancement method and system based on a multi-microphone array, terminal equipment and a medium, and relates to the technical field of speech enhancement. The method is applied to the intelligent glasses, and comprises the following steps: collecting noise-containing multichannel original time domain voice signals through a multi-microphone array of the intelligent glasses, and performing time-frequency domain conversion to obtain a frequency spectrum set; frequency domain complex spectrum amplitude-phase characteristics of each channel are extracted, a local time sequence context is introduced, and multi-channel space acoustic characteristics of amplitude-phase decoupling are constructed; and regulating and controlling the double-tone-area complex spectrum through pre-trained deep neural network modeling to obtain a target complex spectrum, and performing inverse time-frequency domain conversion and reconstruction to obtain an enhanced independent wearer voice signal and a preset direction target voice signal. According to the method, double-target voice parallel enhancement and non-target interference suppression are realized, the power consumption is low, the real-time performance is high, the method adapts to the harsh performance constraint of the end side of the intelligent glasses, and the robustness of voice enhancement in a complex noise scene is excellent.
Owner:ELEVOC TECH CO LTD

Speech enhancement noise reduction system based on generative adversarial network

The invention discloses a speech enhancement and noise reduction system based on a generative adversarial network, and the system comprises a time-frequency alignment module which is used for collecting and preprocessing original noise speech; the three-component coding module is used for extracting voice content features, voiceprint features and noise characterization; the dual-domain collaborative generation module is used for generating candidate enhanced waveform frequency spectrums and performing collaborative correction; the multi-head discrimination module is used for generating a noise correction factor through a time domain discriminator, a frequency domain discriminator and an index discrimination head; the noise re-projection module is used for correcting noise representation, generating heavy noise voice and obtaining a noise self-adaptive enhancement result set; and the consistency reconstruction module is used for spectrum consistency correction and phase reconstruction. According to the method, the double-domain three-decoupling VoiceGAN is used for voice enhancement and noise reduction, and the method has the advantages of being high in naturalness, high in intelligibility and good in generalization.
Owner:HANGZHOU QINGOU TECHNOLOGY CO LTD

Bluetooth earphone call intelligent noise reduction method based on cloud collaboration

The invention discloses a Bluetooth earphone call intelligent noise reduction method based on cloud cooperation, and the method comprises the following steps: S1, obtaining a call voice signal and an environment state parameter, carrying out the preprocessing, and constructing a voice spectrum sequence and an environment feature sequence; s2, performing noise estimation and suppression processing on the speech spectrum sequence by using a DCCRN model to obtain a spectrum feature sequence; s3, constructing a cooperative processing sequence based on the environment feature sequence and the spectrum feature sequence; s4, using a Conv-TasNet model to execute speech enhancement on the co-processing sequence, and generating a target spectrum sequence; s5, performing phase reconstruction and inverse frequency spectrum transformation based on the target frequency spectrum sequence to generate a target voice signal; and S6, constructing a loss function according to the target voice signal, and updating the Conv-TasNet model. According to the method, the DCCRN model and the like are fused, and the method has the advantages of high noise reduction precision, high environment adaptability and sustainable optimization.
Owner:深圳市美迪声科技有限公司

Speech enhancement method, apparatus, device, medium, and program product

The application provides a speech enhancement method, device, equipment, medium and program product. The speech enhancement method comprises the following steps: performing speech enhancement on a microphone receiving signal to obtain a first enhanced signal; performing speech enhancement on a signal in a low frequency band of the first enhanced signal to obtain a second enhanced signal; determining a target enhanced signal based on the first enhanced signal and the second enhanced signal; and performing sound amplification on the target enhanced signal to obtain an amplified signal. The application can reduce the probability of howling in a sound amplification environment.
Owner:IFLYTEK CO LTD

Binaural speech enhancement method based on lightweight spatial perception complex network

The invention discloses a binaural speech enhancement method based on a lightweight spatial perception complex network. The binaural speech enhancement method comprises the following steps: acquiring speech data, noise data and binaural transfer function data of different speakers; preprocessing the acquired data, applying short-time Fourier transform to obtain a complex spectrum, and retaining a unilateral spectrum; layered feature extraction is carried out through an encoder, and abstract feature representation is formed through plural batch normalization and an activation function; the coding features are input into a lightweight attention module, and the time-frequency features of the voice dominant region are enhanced through an attention mechanism; inputting the time-frequency features into a space-time-frequency feature dynamic fusion module to obtain fusion features; recovering the resolution of the fusion features through a decoder, generating a complex ratio mask, and calibrating whether the frequency dimensions are matched or not; and applying the calibrated complex ratio mask to the noisy voice complex spectrum, and obtaining enhanced left and right ear time domain signals through inverse short-time Fourier transform.
Owner:GUANGZHOU MARITIME INST

Speech enhancement methods, devices, electronic devices, storage media, and software products

This application discloses a speech enhancement method, apparatus, electronic device, storage medium, and program product, belonging to the field of artificial intelligence technology. The method includes: during a call conducted through a head-mounted device, determining a first real-world object in a first scene based on the gaze information of a first user, wherein the first user is a user wearing the head-mounted device, and the first real-world object is a real-world object exhibiting speaking behavior and for which the first user intends to perform sound suppression processing; acquiring first voice feature information of the first real-world object; and suppressing the voice signal corresponding to the first real-world object in the detected voice signal of the first scene based on the first voice feature information and the first user's second voice feature information.
Owner:VIVO MOBILE COMM CO LTD

Speech recognition method and device based on environmental noise enhancement and medium

The invention discloses a speech recognition method and device based on environmental noise enhancement and a medium, and the method comprises the steps: carrying out the framing processing of a to-be-recognized target speech signal, and obtaining at least one frame of speech signal; determining respective first power spectrums of the at least one frame of voice signal, and further determining respective Mel spectrums of the at least one frame of voice signal; determining respective noise feature vectors of the at least one frame of voice signal based on a noise perception model; and target model parameters of the speech enhancement model are determined based on the parameter generation model, so that the model parameters of the speech enhancement model are adaptively generated according to the to-be-recognized target speech signal. Then, inputting the respective first power spectrum of the at least one frame of voice signal into a voice enhancement model, and determining the respective second power spectrum of the at least one frame of voice signal based on the voice enhancement model; and finally, determining a target text sequence corresponding to the target voice signal based on the voice recognition model. Therefore, the accuracy of the determined target text sequence is improved.
Owner:WEBANK (CHINA)

Speech enhancement method based on noise fusion

The invention discloses a voice enhancement method based on noise fusion, and relates to the technical field of voice enhancement. According to the speech enhancement model, a frequency band division recurrent neural network is taken as a core architecture, noise-containing mixed speech and additional noise signals are input, and after processing of a noise fusion module, generated enhanced speech is output. Noise input is composed of a plurality of actually recorded noise segments, and the model can efficiently learn diversified noise modes in real time. The noise fusion comprises signal level splicing fusion, feature level splicing fusion, conditional normalization fusion and cross attention fusion. According to the method, the performance upper limit and robustness of speech enhancement are greatly improved on the premise of keeping the advantages of light weight and calculation efficiency of an original model.
Owner:SHANGHAI JIAOTONG UNIV

Hearing aid speech enhancement method and device based on artificial intelligence, and medium

The invention discloses a hearing aid speech enhancement method and device based on artificial intelligence, and a medium, and relates to the technical field of speech signal processing, and the method comprises the steps: inputting a standardized speech enhancement input feature into a direction perception high-frequency speech enhancement model, extracting a coding feature through a shared coding part in the direction perception high-frequency speech enhancement model, and carrying out the recognition of the coding feature; determining a target sound source direction through a direction information generation part in a direction perception high-frequency speech enhancement model, converting the target sound source direction into direction-related information, and combining the direction-related information with the encoding features to generate direction enhancement features; and inputting the direction enhancement feature into a high-frequency recovery part of a direction perception high-frequency speech enhancement model, performing enhancement processing on a high-frequency band corresponding to the target sound source to obtain high-frequency enhancement data, and compensating the high-frequency enhancement data according to the audiogram of the wearer to obtain individualized high-frequency enhancement frequency domain amplitude data. According to the invention, the recognition capability of the wearer on the target voice is improved.
Owner:HEARING EARTH (GUANGZHOU) TECHNOLOGY CO LTD

Information processing apparatus, information processing method, and non-transitory recording medium

PendingUS20260253597A1Information processingNoise
An information processing apparatus including: a constraint loss calculation unit that calculates a constraint loss using an estimated speech enhancement mask output by a speech enhancement mask estimation model in case a noise-mixed speech is input, and an estimated noise enhancement mask output by a noise enhancement mask estimation model in case the noise-mixed speech is input; and a parameter update unit that updates parameters included in the speech enhancement mask estimation model and parameters included in the noise enhancement mask estimation model, using a speech enhancement mask loss indicating a difference between the estimated speech enhancement mask and a target speech enhancement mask calculated in case a speech and a noise in the noise-mixed speech are known, and a noise enhancement mask loss indicating a difference between the estimated noise enhancement mask and a target noise enhancement mask calculated in case a speech and a noise in the noise-mixed speech are known, and the constraint loss.
Owner:NEC CORP

Artificial intelligence-based speech processing method and apparatus, computer device, and medium

PendingCN122658332AEngineeringVoice data
The application belongs to the technical field of artificial intelligence, and relates to a voice processing method based on artificial intelligence, which comprises the following steps: receiving an input voice signal; pre-processing the voice signal to obtain voice data; calling a preset voice processing model; wherein the voice processing model comprises a modeling module, an attention module and a convolutional neural network module; performing feature extraction on the voice data based on the modeling module to obtain feature data; performing global modeling processing on the feature data based on the attention module to obtain global features; processing the global features based on a skip connection mechanism in the convolutional neural network module to generate a target complex spectrum; and performing inverse transformation processing on the target complex spectrum to obtain a target voice signal. The application also provides a voice processing device based on artificial intelligence, a computer device and a storage medium. The application can be applied to the voice enhancement processing scene in the fields of financial technology and digital medical treatment, and effectively improves the generation quality of the target voice signal.
Owner:PING AN TECH (SHENZHEN) CO LTD

Ultra-low computing resource speech enhancement method based on Half-UNet architecture

The invention discloses an ultra-low computing resource speech enhancement method based on a Half-UNet framework. According to the method, a decoder of a UNet is simplified, a recurrent neural network module is arranged between an encoder and the decoder to construct a Half-UNet architecture, and in combination with feature fusion, adaptive frequency band division and power law compression penalty technologies, the calculation complexity and parameter quantity are greatly reduced while enhancing performance is ensured. The method comprises the following steps: carrying out frequency band combination on frequency spectrums of input noise voice by using a filter obtained by training of a self-adaptive frequency band division module, and reducing high-frequency characteristic redundancy; carrying out feature extraction by using Half-UNet, and reconstructing a frequency spectrum; a power law is used to compress penalty terms to enhance weak detail features, and the weak detail features are prevented from being submerged by strong noise features, so that the voice quality of the model in a low signal-to-noise ratio environment is improved. The method has the advantages that under the condition that only about 23.7 K parameters and 25.42 MMACs operand are needed, the voice enhancement effect equivalent to that of a large-scale deep model is achieved, and the method is particularly suitable for resource-limited equipment such as earphones and hearing aids.
Owner:EAST CHINA NORMAL UNIV +1