Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7 results about "Vocal organ" patented technology

N any of the organs involved in speech production. a movable speech organ. the vocal apparatus of the larynx; the true vocal folds and the space between them where the voice tone is generated.

Multi-language automatic identification method and system

The invention relates to a multi-language automatic recognition method and system, and belongs to the technical field of language recognition, and the recognition method comprises the steps: receiving an original voice signal, and carrying out the preprocessing of the original voice signal, and obtaining a preprocessed voice frame sequence; synchronously extracting an acoustic feature vector, a vocal organ motion feature matrix and a rhythm feature vector from the voice frame sequence to form a feature triple; performing language family classification according to the acoustic feature vector and the vocal organ motion feature matrix, and outputting a candidate language family set; inputting the candidate language family set and the rhythm feature vector into a dialect clustering model, and outputting a refined dialect cluster tag; calculating an acoustic feature weight value, a vocal organ motion feature weight value and a rhythm feature weight value, and carrying out weighted operation on the feature triple to generate a weighted feature vector; and inputting the weighted feature vector and the refined dialect cluster label into a language decision model, and outputting a language recognition result containing a language label and a confidence value. According to the invention, the accuracy and robustness of multilingual recognition in a complex environment are improved.
Owner:BEIJING HIZHI TECH CO LTD

English pronunciation correction auxiliary system

PendingCN121617386ASpeech recognitionVocal organPersonalized learning
The English pronunciation correction auxiliary system comprises a microphone for collecting English pronunciation input of a user, a storage unit for storing a standard English pronunciation model and historical pronunciation data of the user, a processor for comparing and analyzing the pronunciation of the user and the standard pronunciation in real time, and a feedback unit for outputting correction feedback information, the output end of the microphone is connected with the input end of the processor, the storage unit is in bidirectional connection with the processor, and the output end of the processor is connected with the input end of the feedback unit. According to the invention, accurate correction and visual guidance of pronunciation errors are realized through a multi-mode feedback mechanism. The pronunciation of the user can be compared with the standard model in real time, the motion difference of vocal organs is visually displayed through the dynamic mouth shape animation and the highlight mark, and the user is helped to quickly understand the error source. By continuously updating the user pronunciation data and continuously optimizing the analysis model, a personalized learning path is formed, and the pronunciation accuracy and learning efficiency are effectively improved.
Owner:张鹏鹏

Cross-modal silent voice reconstruction method and system based on ear canal air pressure micro-motion perception

The invention discloses a cross-modal silent voice reconstruction method and system based on ear canal air pressure micro-motion sensing, and belongs to the technical field of man-machine interaction and wearable computing. The method comprises the following steps: acquiring a non-acoustic air pressure sequence caused by vocal organ movement by using a micro air pressure sensing unit arranged in the in-ear earphone; constructing a robust feature space through adaptive baseline drift suppression and rhythm perception data enhancement processing; and further mapping the TPVS into a high-fidelity acoustic Mel spectrum by using an end-to-end deep neural network including domain adversarial adaptation, cross-modal semantic alignment, coarse-grained Mel spectrum generation and residual detail correction. The technical bottleneck that low-frequency mechanical signals lack high-frequency acoustic features is effectively broken through, and high-precision silent voice instruction analysis in a mobile and noise scene is achieved; and coupling quality evaluation gating and trigger type start / stop control are introduced at a reasoning end, so that invalid reasoning is inhibited and power consumption is reduced under the condition of poor wearing or no trigger.
Owner:DONGHUA UNIV

Phonetic symbol display system, phonetic symbol display program, and speech output system

To provide a pronunciation symbol display system, a pronunciation symbol display program and a voice output system for displaying pronunciation symbols capable of visually recognizing the image of a voice and suitably expressing various pronunciations.SOLUTION: In the symbol display section 21A, a plurality of graphic pronunciation symbols are displayed side by side, the graphic pronunciation symbols using consonant-oriented morphology symbols which are symbols imitating the morphology of the pronunciation organs and correspond to the articulation parts, vowel-oriented morphology symbols in which a component part representing the front-back position of the tongues by the position in the left-right direction and a component part representing the size of the gap from the upper jaw to the tongues by the position in the up-down direction are combined, and auxiliary symbols, and pronunciations output as voices can be represented by the graphic pronunciation symbols corresponding to the respective pronunciations.SELECTED DRAWING: Figure 2
Owner:今井 聖一郎

A speech-based emotion recognition method

PendingCN122347964APhonic TicLinear predictive coding
The application relates to the technical field of speech recognition, in particular to a speech-based emotion recognition method, which comprises the following steps: according to an input speech signal, frame division and windowing are carried out. The application can obtain more profound insights into emotional expression by decomposing the input speech signal into two physical sources of glottal excitation and vocal tract response for independent modeling and analysis, obtaining a vocal tract transfer function set via linear predictive coding operation, and applying inverse filtering to the original speech signal to reconstruct an approximate glottal pulse sequence, effectively stripping the influence of vocal tract resonance on the signal, so that perturbation parameters, open quotient and closed quotient and the like representing vocal cord vibration patterns can be directly calculated, at the same time, the vocal tract transfer function set is used to identify and track the dynamic trajectory of the formant, and the phase difference cosine mean between adjacent frames is combined to quantify the sound production stability, and subtle dynamic adjustment and control stability of sound production organs such as the oral cavity and tongue position caused by emotional changes are captured.
Owner:SHANGHAI LIXIN UNIV OF ACCOUNTING & FINANCE

Multi-mode Chinese tone mouth shape auxiliary teaching system based on multiple attention mechanisms

The invention relates to the technical field of language teaching, and provides a multi-modal Chinese tone mouth shape auxiliary teaching system based on multiple attention mechanisms, and the system comprises a plurality of modules: a self-adaptive vocal cord vibration signal extraction module which is used for collecting and purifying a voice signal, extracting a fundamental frequency and scoring; the multi-scale tone feature fusion module receives the fundamental frequency signals and scores, extracts features and constructs vectors; the hierarchical tone recognition module processes the feature vectors through a neural network and an attention mechanism, recognizes tones and extracts voiceprint features; the personalized vocal organ motion modeling module selects and generates a personalized vocal organ motion model based on vocal print features; the mouth shape and tongue position dynamic generation module creates a three-dimensional visual mouth shape animation; the pronunciation feedback guidance module compares the pronunciation of the user with the standard pronunciation and provides multi-modal feedback; the tone recognition accuracy is improved through the adaptive signal extraction technology, and an effective auxiliary tool is provided for Chinese teaching.
Owner:KUNMING UNIV OF SCI & TECH

A multilingual automatic recognition method and system

The application relates to a multilingual automatic recognition method and system, and belongs to the technical field of language recognition. The recognition method comprises the following steps: receiving an original voice signal and performing pretreatment to obtain a pretreated voice frame sequence; simultaneously extracting an acoustic feature vector, a vocal organ movement feature matrix and a prosody feature vector from the voice frame sequence to form a feature triple; performing language family classification according to the acoustic feature vector and the vocal organ movement feature matrix to output a candidate language family set; inputting the candidate language family set and the prosody feature vector into a dialect clustering model to output a refined dialect cluster label; calculating acoustic feature weight values, vocal organ movement feature weight values and prosody feature weight values, performing weighted operation on the feature triple to generate a weighted feature vector; and inputting the weighted feature vector and the refined dialect cluster label into a language decision model to output a language recognition result containing a language label and a confidence value. The application improves the accuracy and robustness of multilingual recognition in a complex environment.
Owner:BEIJING HIZHI TECH CO LTD