Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Formant" patented technology

In speech science and phonetics, a formant is the spectral shaping that results from an acoustic resonance of the human vocal tract. However, in acoustics, the definition of a formant differs slightly as it is defined as a peak, or local maximum, in the spectrum. For harmonic sounds, with this definition, it is therefore the harmonic partial that is augmented by a resonance. The difference between these two definitions resides in whether "formants" characterise the production mechanisms of a sound or the produced sound itself. In practice, the frequency of a spectral peak can differ from the associated resonance frequency when, for instance, harmonics are not aligned with the resonance frequency. In most cases, this subtle difference is irrelevant and, in phonetics, formant can mean either a resonance or the spectral maximum that the resonance produces. Formants are often measured as amplitude peaks in the frequency spectrum of the sound, using a spectrogram (in the figure) or a spectrum analyzer and, in the case of the voice, this gives an estimate of the vocal tract resonances. In vowels spoken with a high fundamental frequency, as in a female or child voice, however, the frequency of the resonance may lie between the widely spaced harmonics and hence no corresponding peak is visible.

Depression state prediction method and system based on voice multi-scale time domain perception

The invention discloses a depression state prediction method and system based on voice multi-scale time domain perception, and the method comprises the steps: collecting voice signals of a participant reading a unified standardized text, and carrying out the preprocessing of the voice signals, and generating a corresponding Mel spectrogram; extracting time context features by using a spectrum-time domain feature extraction algorithm to obtain joint feature representation; carrying out multi-scale division on the combined feature along the time dimension, and fusing the features of each scale into a global feature; and obtaining depression state discrimination results of the participants based on the global features. By combining the spectrum-time domain feature extraction algorithm and the frame-level time attention, the time domain locality features of depression speech such as inter-sentence pause, hesitant pause, inter-word transition and formant blurring can be explicitly analyzed and accurately positioned on the Mel spectrum, so that the recognition accuracy and the result stability are remarkably improved.
Owner:NANJING MEDICAL UNIV

A formant extraction method for continuous speech based on peak selection

ActiveCN115064180BSpeech analysisFrequency spectrumFormant
The present invention discloses a continuous speech formant extraction method based on peak selection, comprising: performing a preprocessing operation on a single frame of input speech; using a linear prediction method to preliminarily estimate the peak value in the spectral envelope of the speech frame; establishing a reference point and a formant trough, and then using a peak selection method to establish a mapping relationship between the peak value and the reference point; using the mapping relationship between the peak value and the reference point and the formant trough to determine the formant of the speech frame; and performing formant estimation on the continuous speech: dividing the continuous speech into frames according to different frame numbers, using the above algorithm to loop 100 times to obtain the formant parameters under different frame number tests, averaging the results after 100 loops, and obtaining the final result after smoothing. The method of the present invention can eliminate the influence of merged peaks and false peaks, and has a fast convergence speed and strong robustness.
Owner:NANJING UNIV OF POSTS & TELECOMM

Artificial intelligence-based recruitment interview evaluation method and system, and storage medium

The application discloses a recruitment interview evaluation method and system based on artificial intelligence and a storage medium, relates to the technical field of recruitment interview evaluation, and solves the technical problem that the prior art only completes morphological analysis and pragmatic reasoning through voice recognition and corpus comparison, does not deeply mine dynamic characteristics of voice signals, and does not perform correlation analysis on personality descriptions in resumes and real-time emotional performances in interviews, so that the evaluation dimension is single, and it is difficult to comprehensively and objectively reflect the language expression ability, psychological state and personality matching degree of a job seeker; the application obtains voice data and emotional data of a job seeker; the voice data is preprocessed to obtain a formant sequence corresponding to the voice data; a voice evaluation coefficient is calculated based on the formant sequence; an emotional similarity is calculated based on the emotional data; an interview evaluation coefficient is calculated based on the voice evaluation coefficient and the emotional similarity; and whether the job seeker is qualified is judged based on the interview evaluation coefficient, so that the above technical problem is solved.
Owner:YI ZHANYI (GUANGDONG) TECH INFORMATION CO LTD

Foreign language teaching effect evaluation method and system based on artificial intelligence

The invention discloses a foreign language teaching effect evaluation method and system based on artificial intelligence. The method comprises the steps of data acquisition, voice feature extraction, multi-modal fusion evaluation modeling and teaching effect evaluation. The invention relates to the technical field of teaching effect intelligent evaluation, in particular to a foreign language teaching effect evaluation method and system based on artificial intelligence. According to the scheme, a pronunciation definition composite index is constructed through formant difference analysis, the acoustic semantic coupling degree is calculated, and the phoneme transfer stability is evaluated; the definition, the accuracy and the continuity of the voice can be accurately quantified in a multi-noise environment; according to the scheme, modal attention is introduced, full-interconnection cross-modal cross attention is utilized to realize multi-modal information bidirectional interaction, a complementary gating mechanism is innovatively introduced to adaptively adjust a modal fusion proportion, a light quantum network is combined to generate a local score, a meta-decision network dynamically distributes weight, and finally end-to-end optimization is carried out through a differential weighting strategy. And teaching effect evaluation of fineness and accuracy is realized.
Owner:湖南工商大学

Multi-mode voice conversion method and system based on user behaviors

The invention relates to the technical field of speech recognition and synthesis, and discloses a multi-mode speech conversion method and system based on user behaviors, and the method comprises the steps: carrying out the vibration separation of a bone conduction signal, obtaining a separation vibration signal, mapping the face muscle micro-current into a muscle deformation gradient, and analyzing the intention intensity of a user through implicit behavior data; mapping the muscle deformation gradient into a voice fundamental frequency of the user, linearly converting the separation vibration signal into a voice formant of the user, and performing noise injection on a non-voice active section based on intention intensity to obtain noise injection voice; constructing user behavior voice corresponding to the voice fundamental frequency, the voice formant and the noise injection voice through an orthogonal projection layer in a preset orthogonal regularization vocoder; and synthesizing the speech feature vector and the user behavior speech into bone conduction propagation speech through a bone conduction synthesis layer in an orthogonal regularization vocoder. According to the invention, the accuracy of the multi-mode voice conversion technology can be improved.
Owner:SHENGZHEN BEIHAI RALL TRANSIT CENTURY TECHNOLOGY CO LTD

System and Method Configured for Analysing Acoustic Parameters of Speech to Detect, Diagnose, Predict and / or Monitor Progression of a Condition, Disorder or Disease

The present invention relates to a system and method configured for analysing acoustic parameters of speech to detect, diagnose, predict and / or monitor progression of a condition, disorder, or disease, and more particularly, any of paediatric and adult neurological and central nervous system conditions including but not limited to low back pain, multiple sclerosis, stroke, seizures, Alzheimer's disease, Parkinson's disease, dementia, motor neuron disease, muscular atrophy, acquired brain injury, cancers involving neurological deficits, paediatric developmental conditions and rare genetic disorders such as spinal muscular atrophy. The system and method extracts a first formant data set from words spoken by an individual and uses these to classify the vowels in the words on a first computing device, such as a mobile smart phone equipped with a microphone into which an individual speaks. The system stores at least some of these frequencies for the vowel formants in a second formant data set as a recorded file and provides the second formant data set as input to acoustic metrics to generate score data from which an assessment is made to determine the articulation level of the vowels in the words spoken by the individual, allowing allow for detection, diagnosis, prediction and / or monitoring progression of the condition, disorder, or disease.
Owner:BEATS MEDICAL

Artificial intelligence-based foreign language teaching effect evaluation method and system

The application discloses a foreign language teaching effect evaluation method and system based on artificial intelligence, and the content comprises data collection, speech feature extraction, multi-modal fusion evaluation modeling and teaching effect evaluation. The application relates to the technical field of intelligent evaluation of teaching effect, and particularly discloses a foreign language teaching effect evaluation method and system based on artificial intelligence. The scheme constructs a pronunciation articulation composite index through formant difference analysis, calculates acoustic semantic coupling degree, and evaluates phoneme transfer stability, so that the articulation, accuracy and coherence of speech can be accurately quantified in a multi-noise environment. The scheme introduces modal attention, realizes bidirectional interaction of multi-modal information by using full-interconnection cross-modal cross-attention, and innovatively introduces a complementary gating mechanism to adaptively adjust the modal fusion ratio. Local scores are generated by combining a lightweight network, meta-decision network dynamically allocates weights, and finally, end-to-end optimization is realized by using a differentiable weighting strategy, so that the teaching effect evaluation of fineness and accuracy is realized.
Owner:湖南工商大学

A networked audio product collaborative testing method based on a distributed network

The present application relates to the technical field of acoustic device resource scheduling and conflict determination, first, a standardized multi-type excitation signal is applied to the measured device, audio input and output and near-field sound pressure signals are collected and fused, mel-frequency cepstral coefficients, harmonic-to-noise ratios and formant frequencies are extracted and combined to generate acoustic fingerprints, a multi-dimensional semantic association graph of acoustic fingerprints and physical resource nodes is constructed, and resource real-time mapping and occupation relationship tracking are realized. When scheduling a new task, multi-dimensional features and time sequence overlap queries are performed based on the fingerprint graph, and multi-level conflict state determination is combined with parameters such as phase response and formant shift. Through adaptive adjustment of the graph weight, dynamic optimization of resource conflict prediction is realized. The method improves the accuracy of task scheduling and the resource utilization efficiency in a complex acoustic test environment, and effectively reduces the test risk caused by resource competition.
Owner:SHENZHEN FENDA TECH CO LTD

Method, device and equipment for anti-voice deepfake based on quantum resonance peak disturbance

The application relates to a method, device and equipment for anti-deepfake speech based on quantum resonance peak disturbance, wherein the method comprises the following steps: reading an original speech signal, extracting a fundamental frequency and a formant frequency of the original speech signal; creating a parameterized quantum circuit to generate quantum noise; adding the quantum noise to the extracted fundamental frequency and formant frequency to obtain a disturbed signal; and preprocessing the disturbed signal to obtain a reconstructed speech signal. The optimized quantum noise generated by the quantum neural network can effectively interfere with the learning and generation of speech features by a deepfake model. Through the design of a loss function and the adjustment of an optimization algorithm, the speech after adding the noise is ensured to have the smallest difference in hearing from the original speech.
Owner:RELATED (BEIJING) TECHNOLOGY CO LTD

Digital human generation system and method based on language model

PendingCN121811908ASpeech analysisTongue tipSpeech sounds
The invention discloses a digital human generation system and method based on a language model, and relates to the technical field of artificial intelligence. A complete digital human pronunciation modeling path from voice signal acquisition, formant frequency extraction, deviation analysis and pronunciation stability evaluation to motion compensation control and three-dimensional animation fusion is realized. Compared with the existing mode of driving the digital human pattern only based on energy envelope or mouth shape classification, the method has the advantages that the multi-frame formant frequency of the user voice is extracted and processed, so that the digital human can generate dynamic lingual surface and tongue tip actions according to the sound channel change in the real pronunciation process of the user; therefore, the physical correspondence and linguistic consistency of the pronunciation actions of the digital human are improved. According to the method, the direct mapping between the pronunciation characteristics and the three-dimensional tongue actions is realized, and the authenticity of the pronunciation animation of the digital human and the guidance in a teaching scene are remarkably improved.
Owner:SUZHOU LVHUA TECH CO LTD

Personalized inner voice synthesis using adaptive acoustic parameter modification using demographic data

The systems and methods disclosed herein generate a personalized inner voice audio output that replicates or otherwise shares the acoustic characteristics of a speaker's self-perceived voice, which can differ from their externally perceived voice due to bone conduction effects. The systems and methods disclosed herein can generate a provisional voice clone (e.g., parameters of a voice model) of the speaker's voice (e.g., using a recording of the speaker's voice), and can apply a frequency shift to compensate for the absence of bone-conducted low-frequency adjustment that occurs during natural speech production. A trained artificial intelligence model predicts and applies one or more additional acoustic parameter adjustment values (e.g., formant structure, spectral envelope, prosodic patterns) to the provisional voice clone based on one or more factors (e.g., demographic, anatomical, content, environment). The provisional voice clone can be iteratively refined based on received user feedback (e.g., until the output aligns with the speaker's perception of their inner voice).
Owner:RANDOLPH VENTURES LLC

AI pet implementation method based on multi-dimensional emotion model and personalized sound effect

The invention relates to the technical field of AI pets, and discloses an AI pet implementation method based on a multi-dimensional emotion model and a personalized sound effect, and the method comprises the steps: representing the emotion state of an AI pet through employing a four-dimensional dynamic parameter; constructing character features of the AI pet through the four-dimensional dynamic parameters on the basis of a Puluxik emotion model; according to character characteristics, sequentially performing filter equation, frequency translation amount and formant offset parameter adjustment on the original sound subjected to short-time Fourier transform; and performing inverse short-time Fourier transform processing on the sound after parameter adjustment to generate the sound effect of the AI pet. Four-dimensional dynamic parameters are combined into various composite emotions through the Puluxik emotion model, the traditional binary emotion classification limitation is broken through, the character characteristics of the AI pet are constructed, and the original sound is subjected to parameter adjustment in combination with the character characteristics, so that the sound effect of the AI pet is generated. The fixed sound effect limitation is broken through, the personalized sound effect is realized, and the user experience is improved.
Owner:呜呦(长沙)智能科技有限公司

Method for speech synthesis and apparatus therefor

PCT designated stageWO2026014591A1Speech synthesisSynthesis methodsFormant
A method for speech synthesis and an apparatus therefor are disclosed. The method for speech synthesis, according to one embodiment, may comprise the steps of: generating, on the basis of pitch information associated with the input text, an excitation representation of target speech corresponding to an input text; refining the excitation representation through a diffusion model; generating a formant representation of the target speech on the basis of the input text regardless of the diffusion model; obtaining a spectrogram by aggregating the refined excitation representation and the formant representation; and synthesizing the target speech on the basis of the spectrogram. According to the method, both prosody expression and pronunciation accuracy of a target speech can be improved.
Owner:DEEPBRAIN AI INC

Artificial intelligence-based recruitment interview evaluation method and system, and storage medium

The invention discloses a recruitment interview evaluation method and system based on artificial intelligence, and a storage medium, relates to the technical field of recruitment interview evaluation, and solves the problems that in the prior art, lexical analysis and pragmatic reasoning are completed only through voice recognition and corpus comparison, dynamic features of voice signals are not deeply mined, and the efficiency of recruitment interview evaluation is improved. The technical problems of single evaluation dimension and difficulty in comprehensively and objectively reflecting the language expression ability, the psychological state and the character matching degree of the applicant due to the fact that the character description in the resume and the real-time emotional expression in the interview are not subjected to association analysis in the prior art are solved. The method comprises the following steps: acquiring voice data and emotion data of an applicant; the method comprises the following steps: preprocessing voice data to obtain a formant sequence corresponding to the voice data; calculating a voice evaluation coefficient based on the formant sequence; performing calculation based on the emotion data to obtain emotion similarity; calculating an interview evaluation coefficient based on the voice evaluation coefficient and the emotion similarity; and judging whether the applicant is qualified based on the interview evaluation coefficient. The technical problem is solved.
Owner:YI ZHANYI (GUANGDONG) TECH INFORMATION CO LTD

Children voiceprint feature extraction method and system

The invention discloses a child voiceprint feature extraction method and system. The method comprises the following steps: acquiring an original voice signal; identifying a voiced segment and a multi-peak energy structure of the original voice signal so as to judge whether a plurality of sound sources are overlapped or not; when a plurality of sound sources are overlapped, analyzing a current frame frequency spectrum of the original voice signal, and comparing the current frame frequency spectrum with a pre-established child voice feature mask, marking and suppressing non-target interference frequency components to identify a child sound source so as to obtain a time-frequency mask track; performing weighted suppression on a frequency spectrum in the original voice signal by using the time-frequency mask track, eliminating noise by using an adaptive Wiener filter, and reconstructing a target child voice signal; multi-dimensional voiceprint features are extracted from the target child voice signal, normalization processing is carried out, a high-dimensional voiceprint feature vector is constructed, and the multi-dimensional voiceprint features comprise MFCC, fundamental frequency and formant. By implementing the method provided by the invention, the accuracy and stability of child voiceprint recognition are improved.
Owner:HANGZHOU ZHONGDA CHENG TECHNOLOGY DEVELOPMENT CO LTD

Sound standardization evaluation system based on multi-dimensional acoustic parameters and dynamic analysis of emotional scenes

The present invention provides a sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenes, aiming to achieve structured and objective evaluation of speech performance; the system comprises an audio acquisition module, an audio processing module, a standard evaluation module and an integrated processing module; the audio processing module performs noise reduction, pre-emphasis, sampling rate adjustment, framing and slicing operations on original speech data; the standard evaluation module performs pronunciation accuracy analysis, acoustic basic skill feature evaluation or emotion and scene adaptability analysis based on multimodal fusion according to the evaluation type set by the user; parameters including fundamental frequency, resonance peak, sound intensity, fundamental frequency perturbation, amplitude perturbation, etc. are extracted during the evaluation, and classification prediction is completed in combination with a machine learning model; finally, the integrated processing module generates an evaluation report including scoring results, question prompts and personalized training suggestions, providing closed-loop feedback to support voice training and expression optimization.
Owner:GUANGZHOU SENJI SOFTWARE TECH CO LTD

Enhanced wireless communication handover management system

To effectively resolve the issue of inappropriate handovers during communication sessions, methods, systems, and machine-readable mediums which utilize speech features captured by microphones in the original wireless peripheral device and / or the wireless peripheral device to which the communication session is to be handed over to determine if a handover should proceed or should be reversed. Speech features refer to the composite attributes of spoken language that encompass both acoustic and linguistic features. Acoustic features characterize the sound properties of speech and include, but are not limited to, timbre, pitch, intonation, speaking rate, articulation, prosody, melody, spectral features, formant frequencies, and the like. Linguistic features pertain to the actual content conveyed, comprising words, phrases, syntax, and semantics. This includes the analysis of words, phrases, syntax, and semantics to understand the context and continuity of the conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Personality perception pet multi-mode emotion recognition method

The invention relates to the technical field of emotion recognition, in particular to a personalized perception pet multi-mode emotion recognition method, which comprises the following steps of: collecting sound signals and all sound production events in a long time period in a pet calm state, extracting Mel-frequency cepstrum coefficients, pitch envelopes, formants and short-time energy of the sound signals, and calculating a pet emotion recognition result; and establishing a personalized emotional steady state baseline. According to the method, the sound signals of the pets are collected in the calm state, the multi-dimensional acoustic features are extracted, acoustic behavior modes of different pets under stable emotions can be effectively captured, feature deviation measurement is carried out on new sound production samples based on the base line, acoustic emotion migration vectors with directivity and personality differences are formed, and the accuracy of emotion migration is improved. The accuracy and pertinence of emotion perception are improved, while acoustic processing is carried out, visual signals are synchronously collected, the tail shaking frequency, the ear posture angle and other behavior parameters highly related to the emotion are extracted, and the time sequence characteristics of visual behaviors and acoustic characteristics form a logic pairing relation.
Owner:SHENZHEN SONGSHI INTELLIGENT TECH CO LTD

Dynamically optimized pronunciation learning support system and method integrating acoustic statistics, visualization, and history control

We provide a pronunciation learning support system that improves the efficiency, acceptability, and continuity of pronunciation learning by acoustically quantitatively evaluating learners' pronunciation errors, visually presenting their directionality and magnitude, and individually optimizing teaching materials based on error history. [Solution] This invention processes speech on a phoneme-by-phoneme basis, extracts formants F1 and F2 from each phoneme, and performs error evaluation based on Z-scores and normalized distance Z_norm. Error information is visually displayed as points, vectors, and ellipses in F1F2 space, intuitively presenting the statistical error structure to the learner. Error information is recorded as a history and analyzed over time, and error trends are classified using clustering processing. The teaching material presentation means calculates presentation priority based on scoring based on the amount of error and frequency of recurrence, and is equipped with a mechanism for gradually controlling the difficulty of the teaching material according to the recurrence rate and improvement trend.
Owner:池上 さくら

Chinese multi-dialect-oriented dysarthria detection and grading method

PendingCN121884868ASpeech recognitionDysarthriaSemantic feature
The invention discloses a Chinese multi-dialect-oriented dysarthria detection and grading method, and belongs to the field of speech processing. According to the method, detection and severity grading of dysarthria are synchronously realized through a multi-task deep learning framework. The collected count-off voice is preprocessed, the sampling rate is unified, and the voice is segmented into short segments; a Chinese optimized wav2vec2.0 model is adopted to extract deep semantic features, and key acoustic features such as pronunciation fatigue, tone stability and formant dynamics are extracted in combination with a specially designed predefined feature module; thirdly, a confusion perception attention mechanism is introduced, feature interference caused by dialect variation and dysarthria pronunciation similarity is inhibited, and the generalization ability of the model in different regional crowds is improved; and synchronously outputting health / illness judgment and 0-3-level severity rating through a multi-task classification network in combination with a dynamic weighted loss function. The objectivity, accuracy and practicability of dysarthria detection are improved.
Owner:BEIJING UNIV OF TECH

Chinese speech signal segmentation method and device, equipment and storage medium

ActiveCN121306099ASpeech recognitionSpeech synthesisFormantVowel
The embodiment of the invention discloses a Chinese speech signal segmentation method and device, equipment and a storage medium, and the method comprises the steps: carrying out the sampling of a target audio signal containing the speech corresponding to a target Chinese text, and obtaining signal amplitudes corresponding to a plurality of sampling points; performing voice endpoint detection on the target audio data based on the signal amplitudes corresponding to the plurality of sampling points to obtain a plurality of voice segments in the target audio data; determining a vowel position sequence corresponding to the target voice segment based on the formant energy of the voice signal in the target voice segment; determining syllable segmentation points and initial and final segmentation points of the target speech segment based on signal amplitudes of sampling points between two adjacent vowel positions in the vowel position sequence and initial and final consonants of the Chinese text segment corresponding to the target speech segment; and based on the syllable segmentation points and the initial and final segmentation points of the target voice segment, segmenting the target voice segment to obtain segmented voice elements. According to the embodiment of the invention, the accuracy of voice signal segmentation can be improved.
Owner:BEIJING INFORMATION TECH COLLEGE

A personalized pet multi-modal emotion recognition method

The present application relates to the technical field of emotion recognition, in particular to a pet multi-modal emotion recognition method with individuality perception, comprising the following steps: collecting sound signals and all sound events in a long time period in a pet's calm state, extracting the mel-frequency cepstral coefficient, pitch envelope, formant and short-time energy of the sound signals, and establishing an individualized emotion steady-state baseline. The present application can effectively capture the acoustic behavior patterns of different pets in a stable emotional state by collecting pet sound signals in a calm state and extracting multi-dimensional acoustic features. Based on the baseline, the feature deviation of new sound samples is measured to form an acoustic emotion offset vector with directionality and individuality difference, which improves the accuracy and pertinence of emotion perception. At the same time of acoustic processing, visual signals are synchronously collected, and behavior parameters highly related to emotion such as tail shaking frequency and ear posture angle are extracted, so that the timing characteristics of visual behavior and acoustic characteristics form a logical pairing relationship.
Owner:SHENZHEN SONGSHI INTELLIGENT TECH CO LTD

A speech therapy treatment instrument abnormal voice detection method and system

PendingCN122290639ASyllableAbnormal voice
This invention relates to the field of pediatric therapeutic instrument technology, specifically to a method and system for detecting abnormal speech in a speech therapy instrument. The method involves acquiring voltage signals and filtering discrete speech frame sequences, extracting formant trajectories to analyze frequency change trends, comparing the articulation direction distribution of children with standard pronunciation, identifying directional deviation segments, and using formant neighborhood envelope peak comparison to pinpoint envelope offset segments. It also compares spectral peak and valley distribution characteristics, performs multidimensional temporal overlap comparison, and finally outputs the abnormal speech detection results. This invention utilizes amplitude cohesion to filter stable speech frame sequences, extracts structured formant trajectories, correlates syllable segment change directions to enhance dynamic trend discrimination, analyzes neighborhood envelope peak offsets to refine spectral characterization, cross-validates peak and valley distribution and temporal overlap segments, strengthens multidimensional consistency, achieves hierarchical identification and precise localization of pronunciation abnormalities, and improves the stability and distinguishability of detection results.
Owner:WOMEN & CHILDRENS MEDICAL CENTER AFFILIATED WITH GUANGZHOU MEDICAL UNIVERSITY