Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Speech spectrum" patented technology

Speech spectrum - the average sound spectrum for the human voice. acoustic spectrum, sound spectrum - the distribution of energy as a function of frequency for a particular sound source.

AI digital human intelligent creation management method and system

PendingCN122289480AFeature extractionAnimation
This application discloses an AI digital human intelligent creation management method and system, relating to the field of artificial intelligence technology. The method includes: determining the digital human image and speech content based on user-input digital human creation instructions; extracting features from the speech content to obtain speech spectrum features and emotional feature parameters; generating an initial lip-sync sequence based on the speech spectrum features; correcting the initial lip-sync sequence based on the emotional feature parameters to obtain a target lip-sync sequence; driving the digital human image based on the target lip-sync sequence and emotional feature parameters to generate a digital human animation; and creating an AI digital human based on the digital human animation and speech content. This application, by jointly driving the digital human image with the corrected target lip-sync sequence and emotional feature parameters, enables the generated AI digital human to adaptively follow the emotional fluctuations of the speech in both lip-sync dynamics and overall performance, enhancing the realism and vividness of the digital human video.
Owner:WUHAN BAOJI NEW MEDIA TECHNOLOGY CO LTD

Voice modification

ActiveUS12670918B2TimbreSpeech sound
A computing system that receives an audio waveform representing speech from an individual and produces as output a modified version of the audio waveform that maintains the speaker's speech characteristics as well as prosody for specific utterances (e.g., voice timbre, intonation, timing, intensity). The system uses a bottleneck-based autoencoder with speech spectrograms as input and output. To produce the output audio waveform, the system includes a reconstruction error-based loss function with two additional loss functions. The second loss function is speaker “real vs fake” discriminator that penalizes for the output not sounding like the speaker. The third loss function is a speech intelligibility scorer that penalizes the output for speech that is difficult for the target population to understand. The produced modified audio waveform is an enhanced speech output that delivers speech m a target accent without sacrificing the personality of the speaker.
Owner:SRI INTERNATIONAL

Cross-lingual corpus synthesis method, speech synthesis model training method, and related devices

The application provides a cross-language corpus synthesis method, a speech synthesis model training method and related equipment. The cross-language corpus synthesis method comprises the following steps: obtaining a cross-language text, a target speaker embedding vector, and a language embedding vector corresponding to each language included in the cross-language text; determining a character embedding vector corresponding to each character included in the cross-language text; determining a speech spectrum corresponding to the cross-language text according to the language embedding vector corresponding to each language included in the cross-language text, the character embedding vector corresponding to each character included in the cross-language text, and the target speaker embedding vector; and composing a cross-language corpus from the speech spectrum and the cross-language text. The application can synthesize speech spectrums of the same speaker switching between languages, thereby obtaining a cross-language corpus composed of the speech spectrums and the cross-language text, so that a speech synthesis model with higher naturalness of synthesized speech can be constructed based on the obtained cross-language corpus subsequently.
Owner:UNIV OF SCI & TECH OF CHINA

Microphone voice noise reduction method and device, earphone and computer readable storage medium

ActiveCN117041787BAdaptive filterHeadphones
The application discloses a microphone voice noise reduction method and device, earphones and a computer readable storage medium, and is applied to the technical field of earphones. The method determines microphone full-band voice spectrum energy according to an adaptive filter and collected bone conduction voice spectrum and microphone noise voice spectrum, carries out noise estimation according to the microphone full-band voice spectrum energy and the microphone noise voice spectrum, obtains first microphone noise spectrum energy, acquires noise spectrum energy to be corrected, corrects the noise spectrum energy to be corrected through the first microphone noise spectrum energy, obtains corrected microphone noise spectrum energy, and finally carries out noise reduction on the microphone noise voice spectrum through the corrected microphone noise spectrum energy.
Owner:NANJING GOERTEK ACOUSTICS TECH CO LTD

A high-precision voice emergency response method and system

This invention relates to the field of speech recognition technology and proposes a high-precision speech emergency response method and system. The method includes: performing waveform phase comparison between initial mixed audio and background noise audio, and superimposing the active cancellation sound wave signal with the initial mixed audio to reconstruct the target speech signal; performing time-frequency domain transformation on the target speech signal to obtain speech spectrum data, and extracting the narrowband alarm command characteristic frequency band within the emergency command transmission frequency band; generating an alarm confirmation signal when the consistency comparison passes and the matching degree analysis exceeds a dynamic threshold; responding to the alarm confirmation signal and executing an alarm output action according to the alarm level, while monitoring the generation quality of the target speech signal, and switching to a backup trigger mode when the signal quality is lower than a failure threshold, and obtaining an emergency trigger signal through an independent trigger path to forcibly execute the alarm output action; this invention can improve the efficiency of a high-precision speech emergency response.
Owner:SHENZHEN MEIAN TECH