Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9 results about "SPEECH DISTORTION" patented technology

Speech Sound Disorders. Substitutions- one sound is used in place of another. Distortions- a sound is changed slightly and made incorrectly. Omissions- certain sounds are completely left out.

Ultra-low delay signal sequence conversion method and system

The invention provides an ultra-low time delay signal sequence conversion method and system, which can be widely applied to various time sequence conversion scenes, including hearing enhancement systems and equipment, such as hearing aid, denoising, simultaneous translation, sound conversion signal sequence, intelligent glasses, brain-computer systems and the like. According to the invention, the rapid high-fidelity noise elimination method is adopted, so that the limitation and difficulty in the traditional technology can be overcome. And for the problems of noise uncertainty and voice distortion in the sound related field, the high-quality voice can be recovered by recognizing the voice content as the intermediate language representation with high generalization, and meanwhile, the voice of the same speaker is synthesized by using a pre-trained artificial intelligence module. And the ultra-low time delay is realized by using a variable multi-partition block and a multi-head prediction mechanism in the process. Therefore, in the specific hearing enhancement field, the voice of the target speaker can be focused by using the voice separation technology, and the uncertainty of background noise is avoided.
Owner:SHANGHAI PEDAWISE INTELLIGENT TECH CO LTD

A doctor-patient communication information online synchronization system for a ward round vehicle

InactiveCN122117297AHealthcare resources and facilitiesSpeech recognitionPhysician patient communicationData acquisition
The application relates to the technical field of medical instruments, and discloses a ward round vehicle medical patient communication information online synchronization system, which comprises a data acquisition module, a data preprocessing module, a sound quality analysis module, a tone analysis module, a patient comprehensive analysis module, a report review module and a report generation module. Compared with traditional ward round technology, the system realizes comprehensive intelligentization and dataization of the ward round process through the integrated multi-module collaborative work. In terms of sound quality processing, the system solves the speech distortion problem in a high-noise environment through dynamic noise reduction and audio enhancement technology; in terms of dialect recognition, the system significantly improves the accuracy of dialect communication through tone analysis technology; in terms of report generation, the system greatly improves the efficiency and accuracy through intelligent decision-making and automatic generation mechanism. Overall, the system optimizes the ward round process, improves the medical patient communication quality, reduces the medical error rate, and provides more reliable support for clinical decision-making.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV

A directional sound pickup and adaptive noise reduction system and method for a multi-microphone array

PendingCN122290623AReduce processing burdenavoid damageNoiseFeedback control
This invention relates to the field of audio signal processing technology, specifically to a directional sound pickup and adaptive noise reduction system and method using a multi-microphone array. The invention acquires signals through a multi-microphone array and performs time-frequency conversion. A beamforming unit forms a directional beam according to the target direction and constructs spatial nulls in the interference direction. A noise reduction unit adaptively suppresses noise in the beam output, simultaneously extracting residual noise energy and speech distortion feature values. A closed-loop feedback control unit executes a zero-point adaptive optimization algorithm based on the joint constraints of residual noise energy and speech distortion. Using the aforementioned feature values, a cost factor is calculated and a beam zero-point direction update is generated, which is fed back to the beamforming unit to achieve closed-loop adaptive adjustment of the null direction. This breaks the limitation that traditional beamforming and post-noise reduction are independent, and can simultaneously improve noise suppression capability and speech fidelity in complex acoustic environments.
Owner:SHENZHEN FUDEYUAN DIGITAL TECH CO LTD

Voiceprint-driven voice noise reduction method and terminal equipment

The invention relates to the technical field of voice processing, and discloses a voiceprint-driven voice noise reduction method and terminal equipment, and the method comprises the steps: collecting a terminal user voice sample, and generating a user exclusive voiceprint packet; collecting an audio input by a user, analyzing the input audio through an AI algorithm, separating voiceprint feature information in the audio, and obtaining input voiceprint data; performing similarity comparison on the input voiceprint data and a user exclusive voiceprint packet pre-stored in the terminal, setting a similarity threshold value, retaining audio signals of which the similarity reaches the threshold value, and filtering noise signals of which the similarity does not reach the threshold value; and extracting syllable parameters in the reserved audio signal, combining rhythm information, optimizing voice integrity through a feature completion algorithm, and outputting the voice integrity. According to the method, the exclusive voiceprint of the user is taken as a core screening basis, the method is not limited by the noise type, frequency and intensity, all interference signals without target voiceprint can be filtered, the method is adaptive to a complex use scene, and voice distortion caused by traditional noise reduction is avoided.
Owner:WESTVALLEY DIGITAL TECH

An underwater sound acquisition device capable of avoiding voice distortion

ActiveCN224418931Uaccurate identificationAvoid voice distortionSound waveBass (sound)
The utility model discloses an underwater sound acquisition device capable of avoiding voice distortion, which comprises a shell, a sound wave filtering assembly, a sound wave receiving assembly and a signal transmission part, the shell has a waterproof sealed cavity inside, the sound wave filtering assembly is installed in the waterproof sealed cavity, is used for filtering bass parts formed due to resonance of the sealed space, the sound wave receiving assembly is installed in the waterproof sealed cavity, is used for receiving sound waves filtered through the sound wave filtering assembly, and the signal transmission part is used for realizing signal transmission between the sound wave receiving assembly and external equipment. The underwater sound acquisition device can be installed on a waterproof breathing mask for use to collect human voices of underwater personnel. The sound wave filtering assembly arranged can filter out bass parts formed due to resonance of the sealed space, can avoid voice distortion, the sound wave receiving assembly receives filtered sound waves, and then transmits the sound waves to external equipment through the signal transmission part, so that the sound wave receiving party can receive clear voice information and correctly identify voice content.
Owner:DIVEVOLK ZHUHAI INTELLIGENCE TECH CO LTD

Intelligent manufacturing workshop safety interaction control method and system based on large model

The invention relates to the field of industrial control, in particular to an intelligent manufacturing workshop safety interaction control method and system based on a large model, and the method comprises the steps: collecting a sound signal of a workshop environment, calculating the instantaneous energy of a sliding window, constructing an energy concentration ratio, and discriminating local impact noise; introducing adaptive Kalman filtering by taking an energy concentration ratio as a prior factor, tracking a smooth energy change rate, calculating a window length parameter, and adaptively determining a dynamic window length; and after short-time Fourier transform noise reduction, multiplexing or reconstructing window length according to frame similarity, and inputting pure voice into the large voice recognition model to realize safe interaction control. According to the invention, the problems of noise trailing, insufficient frequency resolution and voice distortion easily caused by a traditional fixed window length are solved, the equipment impact noise and the artificial voice instruction can be effectively distinguished, the signal processing precision and the anti-interference capability are improved, and the method is suitable for high-reliability voice safety interaction in a complex industrial environment.
Owner:SHANDONG BLUEBIRD IND INTERNET CO LTD

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium, relates to the technical field of artificial intelligence, and can be applied to a medical speech generation scene or a financial speech generation scene. Through a semantic token path, the modeling capability of voice on a deep semantic structure is enhanced, and the continuity and definition of semantic expression are ensured; and acoustic details related to rhythm, timbre and emotion are reserved through the text condition and the voice prompt path, so that the naturalness and expressive force of the voice can be improved. And an adjustable weight or a self-adaptive mechanism is introduced in the fusion process, so that the system can realize smooth adjustment between semantic controllability and voice naturalness according to different application requirements, thereby improving the flexibility and the application range of the whole system. According to the method, the problems of voice distortion, emotion weakening or single expression caused by information loss of a single path are effectively reduced, and the method has more stable and superior synthesis performance in complex scenes such as cross-speaker, cross-emotion and dialogue generation.
Owner:PING AN TECH (SHENZHEN) CO LTD

A deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling

PendingCN122314000ASaliency mapPESQ
This invention provides a deep echo cancellation method based on cross-domain prior interaction gating and feature decoupling, comprising: preprocessing the acquired near-end microphone signal and far-end reference signal to obtain a complex spectrum with uniform time-frequency resolution; constructing an echo cancellation model based on the complex spectrum using a cross-modal gating attention mechanism, a conditional Transformer, and a multi-branch decoder; and performing echo cancellation on the newly acquired mixed speech signal based on the echo cancellation model to obtain enhanced near-end speech. This method can still learn echo saliency maps through pseudo-labels / weak supervision in unlabeled scenarios, and inference only requires a single forward computation, thereby improving PESQ / STOI and ERLE and reducing speech distortion.
Owner:ANHUI UNIV

Children inquiry method and device based on large language model, medium and program product

The invention discloses a children inquiry method and device based on a large language model, a medium and a program product, relates to the field of intelligent auxiliary medical treatment, and aims to solve the problems of voice distortion and inaccurate positioning caused by the fact that a single-channel noise reduction method cannot give consideration to voice quality and spatial information in an intelligent inquiry process. The method comprises the steps of performing short-time Fourier transform on inquiry voice data to obtain spatial information of the inquiry voice data to obtain a multi-channel time-frequency mask, obtaining a voice feature matrix based on the multi-channel time-frequency mask and a short-time Fourier transform coefficient, generating a generalized mutual information entropy value through kernel density estimation, quantifying direction consistency between channels, and obtaining a multi-channel mutual information entropy value. Noise reduction can be ensured, spatial information is effectively reserved, and the efficiency of intelligent inquiry is effectively improved.
Owner:XIAO ER FANG HEALTH TECH (BEIJING) CO LTD