Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Speech code" patented technology

Using machine learning speech synthesizer using synthetic analytic speech coding

PendingCN122374817ASpeech codeSpeech synthesis
An apparatus includes a memory configured to store data associated with a machine learning (ML) based speech synthesis model. The apparatus also includes a speech encoder including the ML based speech synthesis model. The speech encoder is configured to perform a synthetic analysis operation of an input speech signal including generating a synthetic version of the input speech signal by the ML based speech synthesis model.
Owner:QUALCOMM INC

Method and apparatus for adjusting speech coding, electronic device, and storage medium

The application provides a speech coding adjustment method and device in dynamic spectrum sharing, electronic equipment and storage medium, and relates to the technical field of wireless communication. The method comprises the following steps: receiving a notification of whether a next first preset time length occupies a shared frequency band sent by a long term evolution (LTE) network at a current time; in response to the notification including that the LTE network needs to occupy the shared frequency band in the next first preset time length, correcting channel quality information, and adjusting a speech coding rate based on the corrected channel quality information. At least the problem that the LTE system has high priority and occupies the shared spectrum for a long time when the traffic volume is large, resulting in poor voice service quality of the UMTS system, is solved. The application is suitable for spectrum sharing optimization, voice service optimization and the like.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Deep learning-based adaptive speech recognition system

PendingCN122337201AFeature extractionSpeech code
This invention discloses a deep learning-based adaptive speech recognition system, relating to the field of speech recognition technology. The system processes raw speech signal data acquired by a speech acquisition module and dynamically adapts it to user identification information to obtain speech coding feature data. Furthermore, it extracts features from environmental metadata to obtain environmental embedding feature data. Based on the environmental embedding feature data and the speech coding feature data, feature recognition processing is performed to calculate recognition feature coefficients. The speech recognition module compares these coefficients with preset recognition feature thresholds and determines the speech quality based on the comparison results. This enables recognition under dynamically changing user and environmental conditions, improving the accuracy of speech recognition.
Owner:IANGSU COLLEGE OF ENG & TECH

A voice-driven three-dimensional virtual figure complex emotional facial motion generation method

The application discloses a speech-driven three-dimensional virtual image complex emotion facial action generation method, relates to the technical field of virtual image animation, and comprises the following steps: receiving a speech segment and an emotion guide image as input, extracting speech features and emotion coding features through a speech coding module and an emotion coding module respectively; randomly sampling a fixed-length noise sequence, splicing the fixed-length noise sequence with time step embedded and processed default shape parameters to form an initial noise input; based on a conditional diffusion model, sequentially guiding a diffusion process with the speech features and the emotion coding features as conditions, generating a three-dimensional facial expression animation with synchronized lip shapes and speech and consistent emotions and input pictures; the method can effectively handle complex emotion scenes, effectively express mixed emotions and hidden emotion characteristics, and quantitative evaluation results show that the method has excellent performance.
Owner:BEIJING INST OF TECH

Speech recognition methods, speech recognition systems, computer equipment and storage media

ActiveCN116959424Bavoid collectingaccurate identificationSpeech recognitionSpeech codeSpeech classification
This application provides a speech recognition method, a speech recognition system, a computer device, and a storage medium, belonging to the field of financial technology. The method includes: inputting target speech with a preset emotion category into a pre-trained multi-task speech recognition model; encoding the target speech using a first speech coding sub-model to obtain initial speech features; performing speech attention processing on the initial speech features using a first attention sub-model to obtain first target attention features; encoding the initial speech features using a second speech coding sub-model to obtain hidden speech features; performing hidden attention processing on the first target attention features and the hidden speech features using a second attention sub-model to obtain second target attention features; and performing speech classification on the second target attention features using a multi-task classification sub-model to obtain a target speech label. This application embodiment can improve the recognition accuracy of multi-task speech recognition.
Owner:PING AN TECH (SHENZHEN) CO LTD

A system and method for processing audio data

PendingCN122313998AData processing systemNoise
This invention relates to the field of speech coding technology, specifically to an audio data processing system and method. The system includes an audio signal framing module, a spectrum threshold routing module, a threshold balance calibration module, a peak link recognition module, and a curve audio correction module. In this invention, after the audio signal is framed and transformed, the amplitude is compared with adjacent differences to trigger dynamic adjustments, enabling fine correction of spectral abrupt change regions. Local threshold adaptive reconstruction maintains frequency band energy coordination, half-value iterative compensation suppresses single-point deviations and stabilizes the energy structure, and inverse-range weighted calculation of peak positions makes frequency band transitions smoother, reducing energy abrupt changes. Offset correction adjusts the amplitude distribution based on the predicted curve difference, optimizing spectral continuity and residual smoothness, and overall improving transient fidelity and low-noise balance, allowing compressed audio to remain clear and natural in complex scenes.
Owner:NANJING CODE NOTE NETWORK TECH CO LTD