Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "SPEECH DISTORTION" patented technology

Speech Sound Disorders. Substitutions- one sound is used in place of another. Distortions- a sound is changed slightly and made incorrectly. Omissions- certain sounds are completely left out.

Ultra-low delay signal sequence conversion method and system

The invention provides an ultra-low time delay signal sequence conversion method and system, which can be widely applied to various time sequence conversion scenes, including hearing enhancement systems and equipment, such as hearing aid, denoising, simultaneous translation, sound conversion signal sequence, intelligent glasses, brain-computer systems and the like. According to the invention, the rapid high-fidelity noise elimination method is adopted, so that the limitation and difficulty in the traditional technology can be overcome. And for the problems of noise uncertainty and voice distortion in the sound related field, the high-quality voice can be recovered by recognizing the voice content as the intermediate language representation with high generalization, and meanwhile, the voice of the same speaker is synthesized by using a pre-trained artificial intelligence module. And the ultra-low time delay is realized by using a variable multi-partition block and a multi-head prediction mechanism in the process. Therefore, in the specific hearing enhancement field, the voice of the target speaker can be focused by using the voice separation technology, and the uncertainty of background noise is avoided.
Owner:SHANGHAI PEDAWISE INTELLIGENT TECH CO LTD

Speech recognition method based on self-supervised pre-training and interactive fusion network

The application discloses a speech recognition method based on self-supervised pre-training and an interactive fusion network, constructs a speech recognition model, uses a self-supervised pre-training model as a feature extraction part after a speech enhancement module, effectively combines the speech enhancement module and the self-supervised pre-training method, and relieves speech distortion caused by speech enhancement; an interactive feature fusion method is used to fuse enhanced features and original audio features, so that information loss in the speech enhancement process is made up. By using the method, low-resource speech recognition results are more accurate, and the recognition accuracy of low resources in a complex environment is improved.
Owner:BEIJING TECH & BUSINESS UNIV

An automatic control system for audio output amplitude and method thereof

The present application relates to the technical field of short wave frequency band communication equipment, in particular to an automatic control system of audio output amplitude and a method thereof. The present application is convenient to build and easy to realize; by audio sampling, the present audio output amplitude and the change trend are actively identified, and the amplifier inside the DSP chip is controlled to dynamically compensate the audio digital signal sent to the speech compression circuit, so that the audio output amplitude meets the index requirement, the small short wave audio output amplitude change is ensured in the extreme environment, the speech distortion degree is reduced, the speech quality of the short wave communication is improved, and the communication requirement in the extreme environment is met.
Owner:SHAANXI FENGHUO ELECTRONICS

Voice noise reduction method, device and storage medium

The application relates to a voice noise reduction method, device and storage medium, wherein the voice noise reduction method comprises the following steps: acquiring a voice signal, extracting a frequency domain component of the voice signal, and dividing the frequency domain component into a high-frequency component and a low-frequency component according to the frequency size, wherein the frequency of the high-frequency component is greater than that of the low-frequency component; acquiring a sound source distance of the voice signal; and performing noise reduction on the voice signal, and in the case that the sound source distance is not lower than a preset threshold, the noise reduction intensity of the high-frequency component is not higher than that of the low-frequency component, thereby solving the problem of voice distortion when the sound source distance is far, and improving the quality of the voice signal.
Owner:ZHEJIANG HUACHUANG VISION TECH CO LTD

Language enhancement method and device, computer equipment and storage medium

The invention relates to the technical field of voice processing, can be applied to the fields of finance and medical treatment, and discloses a language enhancement method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining an input voice signal with noise; the input voice signal with noise is converted into noise embedded data through a pre-trained generative audio encoder; performing denoising processing on the noise embedded data through a denoising encoder to obtain clean embedded data; and converting the clean embedded data into an enhanced target voice signal through a pre-trained vocoder. According to the method, the naturalness of the enhanced voice and the consistency of the speaker are effectively improved, the modeling difficulty of complex noise distribution is reduced, the voice distortion is reduced, meanwhile, the model parameter quantity and the training complexity are greatly reduced, the reasoning speed is improved, and the real-time application can be realized in a low-resource environment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation

PendingCN121354580ASpeech analysisHigh level techniquesComputation complexitySpeech reconstruction
The invention discloses a real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation, and relates to a real-time single-microphone voice noise reduction algorithm. The invention aims to solve the problem that conversation voice is buried by noise due to noisy and diverse background noise of an interphone in a special communication scene. Noise power spectrum estimation is optimized by fusing a continuous minimum value tracking algorithm, a speech enhancement residual error is introduced as a real noise approximate value to participate in a recursive average process, and the response speed and accuracy of noise estimation are improved; a gain function is calculated in combination with an optimal correction logarithm MMSE estimator, and a closed expression approximates exponential integration to reduce the calculation complexity. The algorithm specifically comprises the steps of preprocessing and framing, noise power spectrum estimation, gain function calculation, voice reconstruction and the like, the segmentation signal-to-noise ratio is remarkably improved on the premise that low voice distortion is guaranteed, and the real-time processing requirement is met. The invention belongs to the technical field of voice signal processing.
Owner:HARBIN INST OF TECH

A doctor-patient communication information online synchronization system for a ward round vehicle

InactiveCN122117297AHealthcare resources and facilitiesSpeech recognitionPhysician patient communicationData acquisition
The application relates to the technical field of medical instruments, and discloses a ward round vehicle medical patient communication information online synchronization system, which comprises a data acquisition module, a data preprocessing module, a sound quality analysis module, a tone analysis module, a patient comprehensive analysis module, a report review module and a report generation module. Compared with traditional ward round technology, the system realizes comprehensive intelligentization and dataization of the ward round process through the integrated multi-module collaborative work. In terms of sound quality processing, the system solves the speech distortion problem in a high-noise environment through dynamic noise reduction and audio enhancement technology; in terms of dialect recognition, the system significantly improves the accuracy of dialect communication through tone analysis technology; in terms of report generation, the system greatly improves the efficiency and accuracy through intelligent decision-making and automatic generation mechanism. Overall, the system optimizes the ward round process, improves the medical patient communication quality, reduces the medical error rate, and provides more reliable support for clinical decision-making.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV

Compression method, system and equipment of speech synthesis model

The invention relates to a method, a system and equipment for compressing a speech synthesis model, and relates to the field of artificial intelligence. According to the technical scheme, the problem of traditional model deployment is solved, storage in an ONNX format occupies 112 MB originally, the hardware requirement is high, embedded equipment is difficult to bear, the model is only 30 MB after compression, reasoning is accelerated, and smooth operation can be achieved on edge equipment such as an ARM chip and a low-power-consumption MCU; moreover, the scheme is superior to a traditional compression means, solves the problem of proneness to voice distortion caused by static quantization and violent modification of a model structure, adopts quantization perception training to control errors, and ensures that voice details and tone quality are basically flush with those of an original model; in addition, according to the scheme, the training cost is reduced, 1000 epochs can reach the target on the basis of the original pre-training model, the GPU investment is reduced, the GPU cycle is shortened, and the threshold of enterprises and developers is reduced.
Owner:LOOTOM TELCOVIDEO NETWORK WUXI

A directional sound pickup and adaptive noise reduction system and method for a multi-microphone array

PendingCN122290623AReduce processing burdenavoid damageNoiseFeedback control
This invention relates to the field of audio signal processing technology, specifically to a directional sound pickup and adaptive noise reduction system and method using a multi-microphone array. The invention acquires signals through a multi-microphone array and performs time-frequency conversion. A beamforming unit forms a directional beam according to the target direction and constructs spatial nulls in the interference direction. A noise reduction unit adaptively suppresses noise in the beam output, simultaneously extracting residual noise energy and speech distortion feature values. A closed-loop feedback control unit executes a zero-point adaptive optimization algorithm based on the joint constraints of residual noise energy and speech distortion. Using the aforementioned feature values, a cost factor is calculated and a beam zero-point direction update is generated, which is fed back to the beamforming unit to achieve closed-loop adaptive adjustment of the null direction. This breaks the limitation that traditional beamforming and post-noise reduction are independent, and can simultaneously improve noise suppression capability and speech fidelity in complex acoustic environments.
Owner:SHENZHEN FUDEYUAN DIGITAL TECH CO LTD

Voiceprint-driven voice noise reduction method and terminal equipment

The invention relates to the technical field of voice processing, and discloses a voiceprint-driven voice noise reduction method and terminal equipment, and the method comprises the steps: collecting a terminal user voice sample, and generating a user exclusive voiceprint packet; collecting an audio input by a user, analyzing the input audio through an AI algorithm, separating voiceprint feature information in the audio, and obtaining input voiceprint data; performing similarity comparison on the input voiceprint data and a user exclusive voiceprint packet pre-stored in the terminal, setting a similarity threshold value, retaining audio signals of which the similarity reaches the threshold value, and filtering noise signals of which the similarity does not reach the threshold value; and extracting syllable parameters in the reserved audio signal, combining rhythm information, optimizing voice integrity through a feature completion algorithm, and outputting the voice integrity. According to the method, the exclusive voiceprint of the user is taken as a core screening basis, the method is not limited by the noise type, frequency and intensity, all interference signals without target voiceprint can be filtered, the method is adaptive to a complex use scene, and voice distortion caused by traditional noise reduction is avoided.
Owner:WESTVALLEY DIGITAL TECH

Voice interaction method, apparatus, device, medium and program product

The application provides a voice interaction method, device, equipment, medium and program product. The voice interaction method comprises the following steps: determining a first confidence degree of a first original signal and a second confidence degree of a first voice enhanced signal; the first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal; determining a target signal based on the first confidence degree, the second confidence degree, a second original signal and a second voice enhanced signal; the second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal; and performing voice interaction with a target device based on the target signal. The application can reduce voice distortion caused by voice enhancement.
Owner:IFLYTEK CO LTD

An underwater sound acquisition device capable of avoiding voice distortion

ActiveCN224418931Uaccurate identificationAvoid voice distortionSound waveBass (sound)
The utility model discloses an underwater sound acquisition device capable of avoiding voice distortion, which comprises a shell, a sound wave filtering assembly, a sound wave receiving assembly and a signal transmission part, the shell has a waterproof sealed cavity inside, the sound wave filtering assembly is installed in the waterproof sealed cavity, is used for filtering bass parts formed due to resonance of the sealed space, the sound wave receiving assembly is installed in the waterproof sealed cavity, is used for receiving sound waves filtered through the sound wave filtering assembly, and the signal transmission part is used for realizing signal transmission between the sound wave receiving assembly and external equipment. The underwater sound acquisition device can be installed on a waterproof breathing mask for use to collect human voices of underwater personnel. The sound wave filtering assembly arranged can filter out bass parts formed due to resonance of the sealed space, can avoid voice distortion, the sound wave receiving assembly receives filtered sound waves, and then transmits the sound waves to external equipment through the signal transmission part, so that the sound wave receiving party can receive clear voice information and correctly identify voice content.
Owner:DIVEVOLK ZHUHAI INTELLIGENCE TECH CO LTD

Intelligent manufacturing workshop safety interaction control method and system based on large model

The invention relates to the field of industrial control, in particular to an intelligent manufacturing workshop safety interaction control method and system based on a large model, and the method comprises the steps: collecting a sound signal of a workshop environment, calculating the instantaneous energy of a sliding window, constructing an energy concentration ratio, and discriminating local impact noise; introducing adaptive Kalman filtering by taking an energy concentration ratio as a prior factor, tracking a smooth energy change rate, calculating a window length parameter, and adaptively determining a dynamic window length; and after short-time Fourier transform noise reduction, multiplexing or reconstructing window length according to frame similarity, and inputting pure voice into the large voice recognition model to realize safe interaction control. According to the invention, the problems of noise trailing, insufficient frequency resolution and voice distortion easily caused by a traditional fixed window length are solved, the equipment impact noise and the artificial voice instruction can be effectively distinguished, the signal processing precision and the anti-interference capability are improved, and the method is suitable for high-reliability voice safety interaction in a complex industrial environment.
Owner:SHANDONG BLUEBIRD IND INTERNET CO LTD

Dynamic feature enhancement method for noise robustness speech recognition

The invention relates to the technical field of voice signal processing, in particular to a dynamic feature enhancement method for noise robustness voice recognition. The method comprises the following steps: acquiring a multi-dimensional acoustic context information set, and based on the multi-dimensional acoustic context information set, analyzing a masking and interference relationship between a disaster site noise dynamic characteristic and a rescue worker voice key characteristic to obtain a voice characteristic dynamic damage assessment parameter set; based on the multi-dimensional acoustic context information set, analyzing a competition and masking relationship between the target human voice of the rescue personnel and interference human voice and noise in a disaster site in an acoustic feature space, and obtaining a target sound source dynamic separation degree evaluation parameter set; and based on the voice feature dynamic damage evaluation parameter set and in combination with the target sound source dynamic separation degree evaluation parameter set, analyzing the composite stress intensity of residual noise and voice distortion on final recognition performance, executing a dynamically optimized feature reconstruction and enhancement instruction, and generating and outputting a target human voice enhancement process log. And the reliability of rescue communication is improved.
Owner:SHENZHEN BOSHITE TECH CO LTD

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium, relates to the technical field of artificial intelligence, and can be applied to a medical speech generation scene or a financial speech generation scene. Through a semantic token path, the modeling capability of voice on a deep semantic structure is enhanced, and the continuity and definition of semantic expression are ensured; and acoustic details related to rhythm, timbre and emotion are reserved through the text condition and the voice prompt path, so that the naturalness and expressive force of the voice can be improved. And an adjustable weight or a self-adaptive mechanism is introduced in the fusion process, so that the system can realize smooth adjustment between semantic controllability and voice naturalness according to different application requirements, thereby improving the flexibility and the application range of the whole system. According to the method, the problems of voice distortion, emotion weakening or single expression caused by information loss of a single path are effectively reduced, and the method has more stable and superior synthesis performance in complex scenes such as cross-speaker, cross-emotion and dialogue generation.
Owner:PING AN TECH (SHENZHEN) CO LTD

Non-invasive methods for enhancing speech distortion suppression for robust speech recognition

This invention discloses a non-intrusive method for enhancing speech distortion suppression for robust speech recognition. The method includes the following steps: S1: Input the original complex spectrum and the enhanced complex spectrum; S2: Obtain the distortion suppression coefficient based on the input in step S1; S3: Apply the distortion suppression coefficient to a distortion suppression interpolation algorithm to obtain the output corrected spectrum. This invention achieves low computational complexity and compatibility with existing streaming and non-streaming speech enhancement models by using a non-intrusive front-end and back-end bridging module; the enhancement model requires a small amount of training data and can quickly adapt to a small amount of labeled data; it does not change the output signal of the enhancement model, effectively maintaining the auditory gain of different enhancement algorithms for different aspects of enhanced speech.
Owner:SHANGHAI JIAOTONG UNIV

A deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling

PendingCN122314000ASaliency mapPESQ
This invention provides a deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling, comprising: preprocessing the acquired near-end microphone signal and far-end reference signal to obtain a complex spectrum with uniform time-frequency resolution; constructing an echo cancellation model based on the complex spectrum using a cross-modal gating attention mechanism, a conditional Transformer, and a multi-branch decoder; and performing echo cancellation on the newly acquired mixed speech signal based on the echo cancellation model to obtain enhanced near-end speech. This method can still learn echo saliency maps through pseudo-labels / weak supervision in unlabeled scenarios, and inference only requires a single forward computation, thereby improving PESQ / STOI and ERLE and reducing speech distortion.
Owner:ANHUI UNIV

Children inquiry method and device based on large language model, medium and program product

The invention discloses a children inquiry method and device based on a large language model, a medium and a program product, relates to the field of intelligent auxiliary medical treatment, and aims to solve the problems of voice distortion and inaccurate positioning caused by the fact that a single-channel noise reduction method cannot give consideration to voice quality and spatial information in an intelligent inquiry process. The method comprises the steps of performing short-time Fourier transform on inquiry voice data to obtain spatial information of the inquiry voice data to obtain a multi-channel time-frequency mask, obtaining a voice feature matrix based on the multi-channel time-frequency mask and a short-time Fourier transform coefficient, generating a generalized mutual information entropy value through kernel density estimation, quantifying direction consistency between channels, and obtaining a multi-channel mutual information entropy value. Noise reduction can be ensured, spatial information is effectively reserved, and the efficiency of intelligent inquiry is effectively improved.
Owner:XIAO ER FANG HEALTH TECH (BEIJING) CO LTD