Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

55 results about "Vowel" patented technology

A vowel is a syllabic speech sound pronounced without any stricture in the vocal tract. Vowels are one of the two principal classes of speech sounds, the other being the consonant. Vowels vary in quality, in loudness and also in quantity (length). They are usually voiced, and are closely involved in prosodic variation such as tone, intonation and stress.

Vocal music training vowel pronunciation quality evaluation method based on auditory and visual spatio-temporal feature fusion

The invention provides a vocal music training vowel pronunciation quality evaluation method based on auditory and visual spatial-temporal feature fusion, and the method comprises the steps: collecting vowel pronunciation audio signals and corresponding videos of a singer, and constructing a multi-modal data set; generating a fractional order Mel spectrogram for the audio signal through short-time fractional order Fourier transform of an adaptive order; extracting time sequence features and spatial features of the fractional order Mel spectrogram, and fusing the time sequence features and the spatial features through a gating mechanism to generate audio spatio-temporal features; face visual features in the video are extracted and fused with the audio spatio-temporal features through a cross attention mechanism, and the cross attention mechanism is integrated with a periodic modeling network; the fused features are input into a classifier, a dynamic weight multi-mode cosine loss function training model is adopted, the dynamic weight multi-mode cosine loss function dynamically adjusts the sample weight through a confusion matrix, and the weight is increased for the samples with classification errors based on the historical frequency mistaken division times of the samples; and outputting a pronunciation quality evaluation result.
Owner:FUZHOU UNIV

Acoustic analysis-based dysarthria speech evaluation system

The invention discloses a dysarthria speech evaluation system based on acoustic analysis. The dysarthria speech evaluation system comprises a speech evaluation module, a system management module and a cloud management platform, the speech evaluation module is used for performing vowel composition evaluation, voice evaluation, definition evaluation and composition motion evaluation on the user; the system management module is used for carrying out login authentication management, authority management, data management and evaluation record management on a user, and maintaining the safety of system operation and the integrity of data; and the cloud management platform is used for performing real-time data transmission and data analysis on the speech evaluation module and the system management module to complete dysarthria speech evaluation. According to the invention, objectivity and standard consistency of dysarthria assessment are improved.
Owner:CHINA REHABILITATION RES CENT

Method, system and device for predicting pulmonary nodules through sound and storage medium

The invention relates to the technical field of audio signal processing. The invention provides a method, system and device for predicting pulmonary nodules through sound and a storage medium. The method comprises the following steps: in a quiet test environment, directionally collecting a sound signal of a patient through an audio sensor; wherein the sound signals comprise breath sounds, cough sounds, vowels and sentences; performing noise reduction processing on the collected sound signals to obtain a noise reduction data set; extracting multi-dimensional voiceprint features from the noise reduction data set, and constructing a sound feature matrix; and performing multi-modal fusion analysis on the sound feature matrix and clinical data of the patient, inputting the sound feature matrix and the clinical data into a sound prediction model optimized by a focus loss function, and outputting a pulmonary nodule risk prediction result and a visual diagnosis report. The problems that the misdiagnosis and missed diagnosis rate is high, the radiation risk is large and the equipment dependence is strong when the pulmonary nodules are examined by the existing iconography, and the characteristic analysis is single, the noise immunity is poor and the screening accuracy is insufficient in the sound detection technology are solved.
Owner:SHANXI QINLING QIYAO COLLABORATIVE INNOVATION CENT CO LTD +1

Speech Analysis Method and System for Key Feature Parameters of Freezing Gait Symptoms in Parkinson's Disease Based on AdaBoost Algorithm

The present invention discloses a voice analysis method for key feature parameters of freezing gait symptoms in Parkinson's disease based on the AdaBoost algorithm. Step 1: Collect continuous and stable vowels of Parkinson's disease patients and record whether the Parkinson's disease patients have freezing gait symptoms; Step 2: Perform denoising preprocessing on the voice signals and remove the silent segments; Step 3: Extract various voice features; Step 4: Use the CART algorithm to perform feature selection on the original features and screen out the key features that can effectively represent the information of freezing gait symptoms; Step 5: Train the AdaBoost model; Step 6: Input the feature vector of the voice to be measured into the model to obtain the key feature parameters of the freezing gait symptoms in Parkinson's disease. The present invention uses the AdaBoost algorithm to analyze the freezing gait symptoms in Parkinson's disease, improves the model accuracy by using ensemble learning, and reduces the cost of early analysis of the freezing gait symptoms in Parkinson's disease.
Owner:NANJING UNIV OF POSTS & TELECOMM

System and Method Configured for Analysing Acoustic Parameters of Speech to Detect, Diagnose, Predict and / or Monitor Progression of a Condition, Disorder or Disease

The present invention relates to a system and method configured for analysing acoustic parameters of speech to detect, diagnose, predict and / or monitor progression of a condition, disorder, or disease, and more particularly, any of paediatric and adult neurological and central nervous system conditions including but not limited to low back pain, multiple sclerosis, stroke, seizures, Alzheimer's disease, Parkinson's disease, dementia, motor neuron disease, muscular atrophy, acquired brain injury, cancers involving neurological deficits, paediatric developmental conditions and rare genetic disorders such as spinal muscular atrophy. The system and method extracts a first formant data set from words spoken by an individual and uses these to classify the vowels in the words on a first computing device, such as a mobile smart phone equipped with a microphone into which an individual speaks. The system stores at least some of these frequencies for the vowel formants in a second formant data set as a recorded file and provides the second formant data set as input to acoustic metrics to generate score data from which an assessment is made to determine the articulation level of the vowels in the words spoken by the individual, allowing allow for detection, diagnosis, prediction and / or monitoring progression of the condition, disorder, or disease.
Owner:BEATS MEDICAL

Chinese pinyin sliding input system and method based on multidirectional sliding gestures

The invention discloses a multidirectional sliding Chinese character pinyin input method and a matrix type keyboard layout system, and belongs to the technical field of human-computer interaction of touch screen equipment (IPC classification number: G06F3 / 0488, G06F3 / 0323G06F 40 / 274). Aiming at the problems of complex vowel operation redundancy, initial and final combination conflict and the like in the existing sliding input method, the invention provides the following schemes: 1, the multidirectional sliding input method comprises the following steps of: starting sliding from any initial consonant or vowel (containing a zero initial consonant 0 / er key) on the basis of a matrix keyboard layout; different vowels or vowel combinations are mapped every time sliding is performed by one key in different directions or continuous sliding is performed. 2, multi-direction sliding input logic: defining an 8-direction sliding mapping er to be rightward-or-generated from e in a layout E; the method has the technical effects that the input efficiency is improved by more than 60% (high-frequency syllable steps are reduced by 70%), the false touch rate is lower than 5%, the national standard pinyin library is covered by 100%, and the method is obviously superior to a traditional touch screen input scheme.
Owner:江远

Letter keyboard for Chinese teaching

The utility model discloses an alphabetical keyboard for Chinese teaching, which comprises a keyboard main body, a pressing plate and a chip, the side wall of the keyboard main body is provided with alphabetical key positions, the alphabetical key positions are provided with single vowel key boards, the top of each single vowel key board is provided with an X-shaped groove, the X-shaped groove is internally provided with a movable groove, and the movable groove is provided with a through hole. A movable groove is formed in the top of the pressing plate, a tone electric contact is arranged in the movable groove, a movable column is arranged in the X-shaped groove, a first reset spring is arranged on the inner wall of the X-shaped groove, a column body extending into the X-shaped groove is arranged at the bottom of the pressing plate, a universal head is arranged at the bottom of the column body, and a light-sound electric contact is arranged at the top of the chip. The chip is electrically connected with the tone electric contacts through signal lines respectively, a second reset spring is arranged at the top of the chip, and the top of the second reset spring is fixedly connected with the base of the universal joint. According to the utility model, the tone of pinyin of a Chinese character can be directly observed, and the corresponding tone of the Chinese character can be learned.
Owner:CHENYANG NEW HUB DIGITAL TECHNOLOGY CO LTD

Method for learning to read using specialized text

A method of providing an instructional scaffold to person learning to read English language comprises displaying printed matter to the learner in the form of lists. Preferably text stories. The normal spacing of the letters in words is kept intact; in a first part of the method the rime portions of monosyllable words are bolded and made larger than the onset portions; and in a second part of the method, which also may be used independently, the sequential syllables of multisyllable words are emphasized by alternating plain font with bolded font and the long vowels are identified, for example by underscore, again with the words of any text being kept intact.
Owner:NOAH TEXT LLC

The word of god (WOG): the 1,197,000 letter string of encoded hebrew letters underlying the original bible

A data structure and associated methods for analysis of a continuous 1,197,000-letter unvocalized Hebrew string referred to as the Word of God (WOG). The data structure contains only the twenty-two classical Hebrew letters and their five final forms, with no spacing, punctuation, vowelization, or editorial symbols. Intrinsic placement of the final letters enables deterministic segmentation of the string into 305,490 lexical units and 23,206 verses without external conventions. Fixed letter-number assignments provide a numeric architecture for evaluating substrings, detecting alterations, identifying encoded mathematical correspondences, and performing pattern analysis. The system preserves full semantic range by supporting multiple morphologically valid interpretations of unvocalized Hebrew strings. Methods for segmentation, numeric evaluation, reconstruction, integrity verification, semantic analysis, and mathematical pattern detection are provided thereby providing a reproducible foundation for computational and linguistic research.
Owner:JURAVIN DON KARL

A method for speech recognition using tactile stimulation

: A method for enabling comprehension of speech by tactile stimulation is provided. The method includes steps as follows: receiving an audio input; performing an analog-to-digital conversion on the audio input; performing a DFT to convert the audio input to a frequency domain; extracting key features of the audio input based on a band of the audio input, wherein the key features contain frequencies that are representative of at least one vowel; and converting frequencies of the extracted key features to a range of tactile frequencies, which comprises performing a mapping process on the frequencies to shift or map them into a subset, such that the mapping process transfers the frequencies from a high-frequency interval to a low-frequency interval, wherein the low-frequency interval is a perceivable range of tactile frequencies.
Owner:IREROBOT LTD

Type II diabetes intelligent detection method based on improved logistic regression integrated model

The invention provides a type II diabetes intelligent detection method based on an improved logistic regression integrated model, and belongs to diabetes intelligent detection, and the method comprises the steps: designing a voice vowel data collection scheme of a diabetic patient and a normal person, and constructing a data set: extracting voice audio features in the data set through the Praat voice analysis, and carrying out the standardization processing; feature selection is carried out based on Lasso regression and recursive feature elimination, and a training set, a verification set and a test set are divided; constructing a tree-shaped logistic regression model based on the attention mechanism, performing training by using the training set, performing verification by using the verification set, and dynamically determining an optimal prediction threshold value to obtain a trained tree-shaped logistic regression model based on the attention mechanism; and taking the test set as input, testing the trained tree logic regression model based on the attention mechanism, and obtaining a detection result. According to the method, the detection accuracy is improved under the condition of not depending on demographic characteristics.
Owner:HENAN UNIVERSITY OF TECHNOLOGY

An input method, device and system based on number sequence and number sequence rime

PendingCN122633055ASymbol mappingWord list
The application discloses an input method, device and system based on number sequence and number sequence vowel definition and symbol mapping, which comprises the following steps: defining a number sequence and a number sequence vowel set, wherein the number sequence is a fixed sequence of numbers 1-9, 0 and an extended character, the number sequence vowel set comprises 11 groups of vowels arranged according to the number sequence, and each group of vowels corresponds to a unique number sequence coding bit; establishing a mapping relationship between the 11 groups of vowels and a number sequence input unit, wherein the number sequence input unit comprises at least 10 digital key positions and an extended character key position arranged according to the number sequence; and generating a candidate word list according to the defined number sequence and number sequence vowel set and the mapping relationship in response to a user input operation. The application solves the problem of inconsistent coding logic and multiple keystrokes of vowels in the prior art, significantly improves the input efficiency of Chinese pinyin, supports multiple input modes to share a unified coding system and reduces the learning cost of users.
Owner:梁晨

Pinyin keyboard optimized arrangement scheme

The invention discloses an optimized arrangement scheme of a pinyin keyboard. The method is mainly based on a standard QWERTY keyboard layout, single vowel character keys, single vowels A, O, E, I and U and corresponding 24 tone character keys including grand, Yang, upper sound and lower sound which are unique to Chinese phonetic alphabets are expanded, optimization design is carried out on the layout of newly-added keys, a user can input the Chinese phonetic alphabets conveniently, rapidly and more completely, and the user experience is improved. And in cooperation with related computing equipment software, quick Chinese character input conforming to the natural language expression habit of Chinese can be carried out, and learning and popularization of Chinese and Chinese pinyin are facilitated.
Owner:陈大威

English teaching material

ActiveJP2025159743ATeaching apparatusTongue tipMouth shape
To provide English teaching materials helping students intuitively acquire English pronunciation.SOLUTION: There is provided an English teaching material comprising a position identification member and a plurality of vowel cards. The position identification member includes a bimaxillary display for roughly representing a lateral cross-section of both jaws of a person, and a plurality of identifiers for respectively identifying a plurality of positions of a tip of the tongue during pronunciation. The plurality of vowel cards are each assigned to a different English vowel. The plurality of vowel cards have, on each one side, a face symbol imitating a human face and one or two of a plurality of identification displays. The face display includes a mouth shape diagram for graphically representing the mouth shape when pronouncing the corresponding vowel. The plurality of identification displays respectively have the same form as any of the plurality of identifiers.SELECTED DRAWING: Figure 1
Owner:浜家 優子

Method, apparatus, and medium for recognizing a voice tone

Provided are a method, apparatus, and medium for recognizing tones of speech. The method comprises: detecting the duration position of a vowel of a vowel in a speech to be recognized; detecting a tone core portion of the speech to be recognized based on the duration position of the vowel of the vowel; and identifying the tone category of the speech to be recognized based on the tone core portion. Thus, according to at least one embodiment of the present disclosure, the tone core portion can be more accurately detected within the duration position of the vowel of the vowel, thereby more accurately identifying the tone category.
Owner:NEW ORIENTAL EDUCATION & TECH GRP CO LTD

Information input device and pinyin input method thereof

The invention provides an information input device and a pinyin input method thereof. In a pinyin input mode, selective input of information is realized through an operation signal of a physical key, and the information input device comprises an initial input unit which is configured to switch initial pages in response to an operation signal of a first key and determine initial in the initial pages in response to an operation signal of a first rocker; the vowel input unit is configured to switch vowel pages in response to the operation signal of the second key and determine vowels in the vowel pages in response to the operation signal of the second rocker; and the Chinese character input unit is configured to respond to a pinyin confirmation signal of the third key, generate candidate Chinese characters according to the determined initial consonants and finals, and input selected Chinese characters according to operation signals of the cross key and the fourth key. The invention provides a technical scheme which accords with the operation characteristics of a general handle and can remarkably improve the pinyin input efficiency.
Owner:四川长虹新网科技有限责任公司

Modifying facial feature based on speech signal

A computer-implemented method can include determining a speech transition within a speech signal, the speech transition including a change of sound; determining a mouth state based on the speech transition; determining a vowel transition during a vowel sound within the speech signal; and modifying a facial feature of an avatar based on the mouth state and the vowel transition.
Owner:GOOGLE LLC

Vowel recovery method and device, electronic equipment and storage medium

The invention provides a vowel recovery method and device, electronic equipment and a storage medium, and belongs to the technical field of natural language processing, the vowel recovery method comprises the steps that a first to-be-processed text is acquired, and the first to-be-processed text comprises a first text needing vowel recovery; the first to-be-processed text is input into the vowel recovery model for variable note label synchronous prediction, prediction labels output by the vowel recovery model are obtained, and the prediction labels are three types of variable note labels corresponding to each character in the first text; and determining a target text after vowel recovery based on the first to-be-processed text and a prediction label output by the vowel recovery model. In the process, the three types of variable note tags corresponding to each character in the first text can be synchronously predicted, and conflicts among different variable note tags are avoided, so that the variable note tag is accurately predicted for each character in the first to-be-processed text, and the accuracy of vowel recovery is improved.
Owner:IFLYTEK CO LTD

Vocal singing practice quality grade evaluation method based on fractional order spectrogram deep learning

The present application proposes a vocal singing practice quality grade evaluation method of fractional order spectrogram deep learning, including the following steps; Step S1, collect the audio signal of the singer's vocal vowel singing practice, and according to the singing index, label the corresponding quality grade, construct the sample data set, and use it for model training, testing and verification; Step S2, convert the audio signal into a series of fractional order spectrograms; Step S3, construct the fractional order spectrogram deep feature extraction network based on DenseNet and channel attention mechanism in the model, input the extracted fractional order spectrogram deep features into the BiLSTM network, and extract the time sequence features of the singing practice signal; Step S4, the quantum fireworks algorithm is used to optimize the hyperparameters of the kernel extreme learning machine of the model, and the extracted time sequence features are mapped to a high-dimensional space for quality grade decision, forming the evaluation result; Step S5, train the model; The present application can better adapt to the characteristics of non-stationary signals and provide more accurate spectrum analysis.
Owner:FUJIAN NORMAL UNIV

English syllable u feature wavelet coefficient extraction method based on spline interpolation wavelet neural network

The invention discloses an English syllable u feature wavelet coefficient extraction method based on a spline interpolation wavelet neural network, and the method comprises the steps: firstly carrying out the normalization and threshold preprocessing of an audio signal, and effectively removing the interference of environment noise; secondly, constructing a three-layer neural network feedback matrix based on a six-order spline wavelet function, and generating a feedback matrix through transposition operation and inverse matrix calculation of a global matrix psi; then, inverse discrete Fourier transform is utilized to construct a criterion function containing frequency domain errors; and finally, dynamically adjusting the weight of an output layer through iterative training, and terminating training when the modulus value of the criterion function is smaller than a training error iteration ending condition, so as to obtain a wavelet coefficient set representing the characteristics of the longhairy sound u. According to the method, the frequency domain localization characteristic of the six-order spline wavelet is innovatively combined with the adaptive learning of the neural network, the anti-noise performance is improved while the feature extraction precision is ensured, and the problem of individual pronunciation difference is solved.
Owner:ARMY ENG UNIV OF PLA

Lightweight voice band extension method and device for edge device, terminal and medium

The edge device-oriented lightweight voice band expansion method, device, terminal and medium provided by the application belong to the technical field of voice signal processing, and the method comprises the following steps: obtaining a logarithmic domain narrowband audio signal amplitude spectrum and a logarithmic domain mixed amplitude spectrum; constructing a white noise amplitude spectrum; inputting the logarithmic domain narrowband audio signal amplitude spectrum, the logarithmic domain mixed amplitude spectrum and the white noise amplitude spectrum into a trained voice bandwidth expansion model to generate a first high-frequency component corresponding to a consonant in the narrowband audio signal and a second high-frequency component corresponding to a vowel in the narrowband audio signal, and then obtaining a predicted amplitude spectrum; expanding phase information of the narrowband audio signal, generating a wideband audio signal according to the predicted amplitude spectrum and the expanded phase information of the narrowband audio signal, and outputting the wideband audio signal. The voice bandwidth expansion model is used to reconstruct the consonant, so that the reconstructed consonant component has higher fidelity, and the intelligibility of the reconstructed voice is ensured.
Owner:ELEVOC TECH CO LTD

Vowel restoration method and apparatus, electronic device, and storage medium

The application provides a vowel restoration method and device, electronic equipment and storage medium, and belongs to the technical field of natural language processing, and comprises the following steps: obtaining a first to-be-processed text, the first to-be-processed text comprising a first text requiring vowel restoration; inputting the first to-be-processed text into a vowel restoration model for multi-vowel symbol label synchronous prediction to obtain a predicted label output by the vowel restoration model, the predicted label being three types of vowel symbol labels corresponding to each character in the first text; and determining a target text after vowel restoration based on the first to-be-processed text and the predicted label output by the vowel restoration model. In this process, three types of vowel symbol labels corresponding to each character in the first text can be synchronously predicted, avoiding conflicts between different vowel symbol labels, so that the vowel symbol label of each character in the first to-be-processed text can be accurately predicted, thereby improving the accuracy of vowel restoration.
Owner:IFLYTEK CO LTD

Chinese phonetic alphabet tone corner code labeling method

The Chinese phonetic alphabet tone corner code labeling method is used for supplementing and expanding a currently executed Chinese phonetic alphabet scheme in actual use and is used for reading correct Chinese phonetic alphabets. An existing Chinese pinyin scheme is composed of letters and tones, and a standard mode that pronunciation of Chinese characters is displayed in an integral up-and-down structure mode is adopted. However, on the basis of an information digital technology, information is expressed in a one-way structure mode from left to right, and pronunciation of Chinese characters is often directly expressed by phonetic alphabets without upper tones. According to the scheme, a single form of an up-and-down structure of which tones are marked on finals in a Chinese pinyin scheme is expanded in a left-and-right structure form, and the scheme is more suitable for a left-to-right structure expression mode of a modern information digital technology. The method can complement the missing tones of the existing Chinese pinyin scheme in use, provides correct Chinese character pronunciation, and does not affect listening, speaking, reading and writing of Chinese characters. Strengthening Chinese phonetic alphabets is an important role as a tool for assisting Chinese character pronunciation.
Owner:俞羿君 +1

Chinese speech signal segmentation method, device and equipment and storage medium

This disclosure provides a method, apparatus, device, and storage medium for segmenting Chinese speech signals. The method includes: sampling a target audio signal containing speech corresponding to a target Chinese text to obtain signal amplitudes corresponding to multiple sampling points; performing speech endpoint detection on the target audio data based on the signal amplitudes corresponding to the multiple sampling points to obtain multiple speech segments in the target audio data; determining the vowel position sequence corresponding to the target speech segment based on the formant energy of the speech signal in the target speech segment; determining syllable segmentation points and initial / final segmentation points of the target speech segment based on the signal amplitudes of sampling points between two adjacent vowel positions in the vowel position sequence and the initials and finals of the corresponding Chinese text segment; and segmenting the target speech segment based on the syllable segmentation points and initial / final segmentation points to obtain segmented speech primitives. This disclosure improves the accuracy of speech signal segmentation.
Owner:BEIJING INFORMATION TECH COLLEGE

Hearing system comprising a hearing instrument and method for operating a hearing instrument

A hearing system (2) having a hearing instrument (4) worn in or at a user's ear and a method for operating the hearing instrument (4) are provided. During operation of the hearing instrument (4), an ambient sound signal from the hearing instrument (4) is captured and processed to support the user's hearing. The processed sound signal (O) is output to the user. In a start point identification step at the beginning of the signal processing, the captured sound signal (I) is analyzed to identify a start point of a speech vowel. In a start point enhancement step, the processed sound signal is instantaneously modified when the start point of the speech vowel is identified to enhance the start point. In this context, the processed sound signal (O) is instantaneously modified by selectively increasing the signal level of the processed sound signal in a first frequency sub-range.
Owner:SIVANTOS PTE LTD

Gnn-lstm method for chinese lip speech classification based on node multi-association graph information fusion

The application belongs to the technical field of lip language mouth shape analysis, and discloses a GNN-LSTM Chinese lip language classification method based on node multi-association graph information fusion. Lip key points are represented by three structures of adjacency graph, symmetric graph and upper and lower lip relationship graph. High-dimensional space-time features are extracted under the synergistic effect of graph convolutional neural network and long short-term memory network, the space-time global correlation between lip key points is effectively captured, and the mouth shape class is divided based on initial and final vowels. The multi-level collaborative relationship between initial and final vowels is considered, the influence of mouth shape similarity and visual ambiguity on model performance is reduced, a mouth shape library is established by using the extracted high-dimensional space-time features, the lip shape and corresponding pinyin are more discriminatively mapped and induced, the influence of mouth shape similarity and visual ambiguity on model performance is reduced, the method can adapt to complex changes and many-to-one mapping phenomena in actual pronunciation processes, and effectively enhances the accuracy and robustness of subsequent mouth shape classification.
Owner:XIANGJIANG LAB

A method for extracting characteristic wavelet coefficients of English long vowel u based on spline interpolation wavelet neural network

The present invention discloses a method for extracting characteristic wavelet coefficients of the English long vowel u based on a spline interpolation wavelet neural network. The method first normalizes and thresholds the audio signal to effectively remove environmental noise interference. Secondly, a three-layer neural network feedback matrix based on a sixth-order spline wavelet function is constructed, and the feedback matrix is generated by transposing the global matrix Ψ and calculating the inverse matrix. Then, a criterion function containing frequency domain errors is constructed using an inverse discrete Fourier transform. Finally, the output layer weights are dynamically adjusted through iterative training. When the criterion function modulus is less than the training error iteration end condition, the training is terminated to obtain a set of wavelet coefficients representing the characteristics of the long vowel u. The present invention innovatively combines the frequency domain localization characteristics of the sixth-order spline wavelet with the adaptive learning of the neural network, improving the noise resistance performance while ensuring feature extraction accuracy, and solving the problem of individual pronunciation differences.
Owner:ARMY ENG UNIV OF PLA

Human voice style recognition method based on time-frequency refinement analysis

This invention discloses a method for recognizing vocal styles based on time-frequency refined analysis. The process is as follows: First, a fundamental frequency estimation and harmonic labeling step is performed. The short-duration vowel portion of the vocal signal is first used for fundamental frequency estimation. This estimation employs a frequency estimation algorithm combining time-domain autocorrelation and narrowband spectral energy. Then, based on the estimated fundamental frequency, adaptive harmonic labeling is performed on the spectrum to accurately identify all harmonics. Next, a time-frequency refined analysis step is performed. The vocal signal is refined and features are extracted from both the time and frequency domains, with a focus on periodic variations and harmonic structures. Finally, a support vector machine (SVM) model training and recognition step is performed. The extracted features and corresponding vocal style labels are used to train the SVM model. After training, the model can be used for vocal style recognition. The style is obtained by using the feature vectors extracted from the vocal signal as input.
Owner:SOUTH CHINA UNIV OF TECH