Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

64 results about "Vocal pitch" patented technology

Natural statement decoding method and device based on high-density electrocorticogram

The invention discloses a natural statement decoding method and device based on high-density electrocorticogram. The method comprises the following steps: acquiring an electroencephalogram signal acquired based on the high-density electrocorticogram; different frequency band signals are extracted from the electroencephalogram signals; the signals of different frequency bands comprise high gamma frequency band signals and at least one frequency band signal with the frequency lower than that of the high gamma frequency band signals; acquiring a voice starting point corresponding to each frequency band signal, and determining a target voice starting point based on the voice starting point corresponding to each frequency band signal; after the target voice starting point, acquiring a syllable classification result and a tone decoding result corresponding to each frequency band signal, determining a target syllable classification result based on the syllable classification result corresponding to each frequency band signal, and determining a target tone decoding result based on the tone decoding result corresponding to each frequency band signal; and determining a target tone language corresponding to the electroencephalogram signal based on the target syllable classification result and the target tone decoding result. According to the scheme, the natural statement decoding accuracy can be improved.
Owner:SHANGHAI TECH UNIV +1

Suzhou dialect medical voice electronic medical record conversion system and method

The invention provides a Suzhou dialect medical voice electronic medical record conversion system and method, and the system comprises an acoustic feature extraction module which is used for extracting the acoustic features of an input voice signal; the dialect tone recognition module is used for recognizing a multi-tone system of Suzhou dialects; the voiced sound processing module is used for detecting voiced sound initial consonants in Suzhou dialects and performing acoustic feature mapping; the medical term mapping module comprises a corresponding relation library of dialect medical vocabularies and standard medical terms and a context-based ambiguity resolution unit; the speech recognition engine comprises an acoustic model, a pronunciation dictionary and a language model; and the medical record generation module is used for converting the identification result into a structured electronic medical record. According to the method, a sliding window processing strategy is adopted, continuous voice input of a doctor can be effectively processed, a long-time voice input scene is supported, and various requirements in actual clinical application are met.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Chinese learner-oriented tone evaluation and improvement method

PendingCN121415801ASpeech recognitionTime domainVocal tract
The invention relates to the technical field of speech recognition, in particular to a Chinese learner-oriented tone evaluation and improvement method, which comprises the following steps of: extracting a fundamental frequency F0 curve and sound channel parameters of a speech signal, constructing a turbid and clear adaptive fusion model, and fusing to generate fundamental frequency related characteristics and fundamental frequency unrelated characteristics; constructing a tone error corpus, carrying out tone classification labeling, generating a multi-level tone feature set through a turbid and clear adaptive fusion model, aligning a voice signal with a reference voice time domain, and extracting and decoupling a tone shape feature vector and a tone domain feature vector; and based on the decoupled feature vectors, establishing a dual-channel evaluation path, hierarchically calculating the tone type distance and the tone domain distance of the tones, carrying out weighted fusion, generating a final evaluation score, and finally generating feedback information for the Chinese learner. According to the method, a complete method from feature decoupling to two-channel evaluation is constructed, the pronunciation problem of the learner is quantitatively diagnosed, and the pertinence and efficiency of Chinese tone learning are improved.
Owner:GUANGXI UNIV

Suzhou dialect speech recognition system and method based on tone track neural field

The invention provides a Suzhou dialect speech recognition system and method based on a tone track neural field, and the system comprises a tone track neural field module which is used for modeling the tone change of Suzhou dialects into a continuous space-time neural field; the bidirectional semantic memory network module comprises a forward prediction memory bank and a backward correction memory bank; the phoneme-font coupling error corrector is used for realizing polyphone disambiguation and homonym error correction by establishing association mapping between a phoneme sequence and a font sequence; the semantic entropy calculation module is used for evaluating the uncertainty of the recognition result; and the self-adaptive fusion decision module is used for generating a final recognition text. According to the invention, through an online learning mechanism, the system can continuously accumulate experience from actual use, automatically discover a new language mode and update an identification strategy. The self-improvement capability enables the system to adapt to the dynamic change of languages, and the performance is continuously improved along with the increase of the use time. Each use of the user helps the system to become more intelligent and accurate.
Owner:程思民

Melody note sequence generation method and system based on tone features and machine learning, electronic device, and storage medium

The application belongs to the technical field of audio processing, and specifically provides a melody note sequence generation method and system based on tone features and machine learning, an electronic device, and a storage medium. The method comprises obtaining a tone sequence based on an original Chinese text; constructing a rich context feature vector; pre-training a gradient boosting model; inputting the context feature vector into the pre-trained gradient boosting model to obtain the probability distribution of each category under the melody mode; concatenating the probability distribution of each category under the melody mode in order to generate a melody note sequence separated by spaces; and obtaining slow, clear, and rhythm-exaggerated singing audio suitable for aphasia MIT rehabilitation based on the melody note sequence and the Chinese text. The application automatically extracts Chinese tone features and predicts the melody mode with context awareness, thereby realizing high-quality generation from Chinese text to natural melody note sequence.
Owner:CHANGSHA LIANYU TECHNOLOGY CO LTD

A smart accompaniment singing method and system

ActiveCN121122218BEngineeringAudio frequency
This invention belongs to the field of audio processing technology and provides an intelligent accompaniment singing method and system. The method can collect the user's vocal pitch and voiceprint information, then divide the song selected based on the song selection command into vocal data and accompaniment data. The pitch of the accompaniment data is adjusted by comparing the vocal data with the user's vocal pitch, and the audio is played based on the adjusted accompaniment data. The user's singing data is continuously acquired and monitored, and when the user's singing data meets preset conditions, accompaniment vocals generated using the voiceprint information and vocal data are added to the audio. This proposed method can tailor the song's pitch to the user's individual needs, adjusting the accompaniment pitch to a suitable range for the user, ensuring easier singing. It can also selectively add accompaniment vocals based on the user's singing performance, thereby guiding the user to adjust their singing style, improve their singing level, and enhance their singing experience.
Owner:SHENZHEN WANSHENG CULTURE TECH CO LTD

An AI-generated text detection method, system, device and medium

The application discloses an AI generated text detection method, system, device and medium, and relates to the technical field of text processing. The application converts a Chinese text sentence to be detected into a tone category sequence composed of tone categories. In this process, the categories of Chinese tones are used as a quantitative index of phonological structure, and text analysis is converted from a high-dimensional and high-cost semantic space to a low-dimensional and high-efficiency phonological feature space. The application is not dependent on training data of a specific model and is adaptable to text detection of multiple models. Then, N-gram analysis is performed on the generated tone category sequence, and the occurrence frequency of each N-gram combination is calculated to form a feature vector. Finally, an AI generated text is judged through a pre-trained classification model. In this process, the calculation type in processing is mainly string processing and frequency statistics, and the inconsistency problem of the model in multiple model invocations is avoided, so that the AI generated text can be accurately detected.
Owner:FOSHAN UNIVERSITY

Chinese pronunciation defect recognition method and system based on speech recognition

The invention provides a Chinese pronunciation defect recognition method and system based on speech recognition, and the method comprises the steps: obtaining a Chinese speech signal, generating an initial time-frequency spectrum through short-time Fourier transform, extracting a frequency feature sequence from the initial time-frequency spectrum, carrying out the segmentation processing of the speech signal through a time-domain signal segmentation technology, and carrying out the recognition of a Chinese pronunciation defect. Preliminary segmentation fragments are obtained; for the abnormal tone distribution diagram and the abnormal airflow distribution diagram, fusing tone duration and breathing rhythm mode features to generate a comprehensive feature matrix, performing boundary optimization on an abnormal interval through tone smoothing processing, and determining feature distribution after smoothing; and positioning a specific pronunciation defect syllable through the defect candidate interval in combination with conjoint analysis of a tone change mode and an airflow amplitude threshold value, tracing correlation characteristics of a frequency abrupt change point location and signal energy distribution, and judging a defect cause combination to obtain a final defect positioning result.
Owner:PUYANG VOCATIONAL & TECHN COLLEGE

An unsupervised speech recognition modeling method based on phoneme segment level representation discretization

ActiveCN119446136BSpeech recognitionCluster algorithmLaotian language
The present application relates to an unsupervised speech recognition modeling method based on phoneme segment level representation discretization, belonging to the field of speech recognition. The present application uses an IFMF model to extract phoneme features from the original audio through a speech feature discretization module, then trains a K-Means clustering algorithm to cluster the audio representation, and obtains the clustering index as a discrete label; through an adversarial learning module, a generator network and a discriminator are trained; through a speech discrete representation decoder module, a language model is used to decode the phoneme discretization-based model output obtained by unsupervised training. The present application fuses multiple speech features, adopts a phoneme segment level representation discretization method, considers the influence of tone information on the Lao language, and as much as possible reduces the influence of information redundancy caused by fine-grained features on cross-modal modeling. The method of the present application achieves competitive results compared with traditional speech recognition methods.
Owner:KUNMING UNIV OF SCI & TECH

Speech recognition method and system based on big data

The invention relates to the technical field of speech recognition, in particular to a speech recognition method and system based on big data, and the method comprises the steps: collecting each speech signal frame and the fundamental frequency thereof in a target region, obtaining a first matching index of each voice signal frame with a positive tone, a second matching index of each voice signal frame with a negative tone, a third matching index of each voice signal frame with a negative tone, and a fourth matching index of each voice signal frame with a positive tone; taking the maximum value of the first matching index, the second matching index, the third matching index and the fourth matching index corresponding to each voice signal frame as a characteristic value of each voice signal frame, obtaining a label of each characteristic value, and inputting each voice signal frame and the label of the characteristic value into a trained machine learning model, and obtaining the voice content of each voice signal frame. According to the invention, the problem of low voice content recognition accuracy is solved.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

A national vocal dialect prosody intelligent correction method and system

This invention proposes an intelligent error correction method and system for ethnic vocal music dialect prosody. The method includes: collecting multimodal data, extracting features to form a training dataset, constructing and updating a dynamic dialect prosody map, and performing cross-modal comparative analysis to obtain a joint representation of cross-modal error features. A neural vocoder direct-connect correction model is constructed, inputting the joint representation to generate a corrected Mel spectrum. A cross-modal generative adversarial network is constructed, including a generator and three discriminators. The generator generates a joint latent representation based on high-fidelity corrected audio and three features. The three discriminators respectively judge the naturalness of the audio, the compliance of the text tone, and the rhythmic coordination of the musical score. This invention solves the problem of manual error correction by constructing a prosody map, overcomes the bottleneck of manual detection by utilizing cross-modal comparative analysis, improves the generalization ability of the model by combining meta-learning algorithms, and optimizes the correction results with the help of adversarial networks, thereby reducing manual costs and improving the error correction efficiency of ethnic vocal music works.
Owner:GUIZHOU RADIO & TV UNIV

Tibetan speech recognition method fusing tone perception hybrid expert and search correction

The application discloses a Tibetan speech recognition method fusing tone perception hybrid experts and retrieval error correction, and belongs to the technical field of signal processing in the electronic industry. The specific steps of the recognition method are as follows: I: Tibetan speech signals are acquired, and corresponding log-mel spectrogram features and fundamental frequency contour features are extracted; II: the tone gating weight is calculated based on the fundamental frequency contour features, and is dynamically routed to the corresponding hybrid expert network to acquire dialect-independent acoustic hidden layer features. The application effectively solves the model interference problem caused by the presence or absence of tones among multiple dialects, avoids the parameter conflict between tone dialects and non-tone dialects, significantly improves the recognition performance of multi-dialect hybrid training, corrects the homonym heterograph error commonly existing in Tibetan, improves the performance of multi-dialect Tibetan speech recognition, can be used for converting Tibetan speech into characters, and is helpful for protecting and mining Tibetan culture.
Owner:CHINA UNIVERSITY OF POLITICAL SCIENCE AND LAW

A method for evaluating and correcting spoken pronunciation

ActiveCN121862161BImprove the pertinence of correctionsImprove efficiencySpeech analysisSpoken languageSpeech sound
This invention relates to the field of speech recognition technology, specifically a spoken pronunciation correction and evaluation method. The method includes: constructing an audio time-series sequence containing multiple phonemes based on the standard text spoken by the user; comparing the audio features corresponding to different phonemes in the audio time-series sequence based on the standard pronunciation of each phoneme to locate the pronunciation deviation position corresponding to each phoneme; if the current phoneme has a pronunciation deviation position, performing contextual scene analysis on the pronunciation deviation position by combining the duration, interval, and tone of the preceding and following phonemes to determine the deviation type in the current phoneme's context; identifying the user's abnormal pronunciation patterns; and configuring correction weights for each abnormal pronunciation pattern based on the difference feature values ​​between any two abnormal pronunciation patterns, and determining the correction direction corresponding to each phoneme according to the correction weights of the abnormal pronunciation patterns. This achieves targeted and efficient spoken pronunciation correction.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Multi-Host Redundant Audio Interface System with Real-Time Vocal Pitch Correction and Enhanced Signal Processing

An audio interface comprises a CPU (central processing unit) that includes a DSP (digital signal processor), a MIDI (musical instrument digital interface) processor, and two Voc FX (vocal effects) modules; MIDI IN, MIDI OUT, and MIDI HOST ports connected to the MIDI processor; first and second microphone ports connected to the first and second Voc FX modules, respectively; a plurality of OUTPUT ports connected to the DSP; and first and second HOST ports connected to the DSP; wherein the audio interface provides multi-host redundancy, real-time signal processing and MIDI control,
Owner:APOGEE ELECTRONICS CORP

Speech synthesis method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, and provides a speech synthesis method, device and equipment and a medium, which are applied to the fields of medical treatment, finance and the like, and the method comprises the following steps: integrating initial speech data to obtain target speech data; extracting the target voice data to obtain target phoneme features; fusing the target voice data and the target phoneme features to obtain target fusion features; and synthesizing the target fusion feature based on a synthesis strategy to obtain a target synthesized speech. According to the embodiment of the invention, the method achieves the integration processing of the initial voice data to obtain the target voice data, carries out the extraction and fusion processing of the target voice data to obtain the target fusion features, and carries out the synthesis processing of the target fusion features based on the synthesis strategy to obtain the target synthesis voice. The tone accuracy and the phoneme alignment accuracy of the synthesized voice are improved, the method has good field adaptability and rhythm continuity, high fidelity of small sample voice cloning is ensured, and therefore the synthesis efficiency is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Spoken language pronunciation correction evaluation method based on deep learning

The invention relates to the technical field of speech recognition, in particular to a spoken language pronunciation correction evaluation method based on deep learning, which comprises the following steps: constructing an audio time sequence containing a plurality of phonemes on the basis of a standard text when a user pronunciates; on the basis of the standard pronunciation of each phoneme, comparing audio features corresponding to different phonemes in the audio time sequence, and positioning a pronunciation deviation position corresponding to each phoneme; if the current phoneme has the pronunciation deviation position, performing context scene analysis on the pronunciation deviation position by combining the duration, interval and tone of the previous and next phonemes, and determining the deviation type of the current phoneme in the scene; identifying an abnormal pronunciation mode of the user; and based on the difference characteristic value between any two abnormal pronunciation modes, the correction weight of each abnormal pronunciation mode is configured, and the correction direction corresponding to each phoneme is determined according to the correction weight of the abnormal pronunciation mode. The pertinence and the processing efficiency of spoken language pronunciation correction are realized.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Song identification method, electronic equipment, storage medium and product

The invention discloses a song recognition method, electronic equipment, a storage medium and a product. The method comprises the following steps: acquiring first audio data; determining first syllable information and first tone information based on the first audio data; determining a first syllable duration based on the first syllable information; determining second audio data corresponding to the first audio data based on the first tone information, the first syllable duration and a database; the database comprises at least one piece of second audio data and second tone information and second syllable duration corresponding to the second audio data; the first syllable duration represents the duration of each syllable in the first audio data, and the second syllable duration represents the duration of each syllable in the second audio data. The song is searched by using the tone in the audio and the duration of each syllable, so that the song identification flexibility is improved, and the identification accuracy is improved.
Owner:MIGU CO LTD +1

Dialect speech recognition method, system and model

The present application belongs to the technical field of speech recognition, and particularly relates to a dialect speech recognition method, system and model. The dialect speech recognition method comprises: S1, generating pseudo labels and screening by a plurality of teacher models; S2, sub-model training; S3, data fusion; S4, student model training, repeating steps S3 to S4 until round fusion training is performed, and student models of each region are obtained after training. The present application combines multi-teacher model knowledge distillation and K nearest neighbor dynamic data fusion training to construct an end-to-end speech transcription method suitable for a multi-dialect scene, and is particularly suitable for a less resourceful dialect environment with a complex tone system, rich speech variants and scarce data.
Owner:GUIZHOU UNIV +1

Han-Zhuang translation and speech synthesis system based on Tacotron2 and HiFi-GAN

The invention discloses a Chinese-Zhuang translation and speech synthesis system based on Tacotron2 and HiFi-GAN, the system comprises text preprocessing, Chinese-Zhuang mapping and speech synthesis module realization, the text preprocessing comprises text standardization, text legality verification and word segmentation and part-of-speech tagging, the Chinese-Zhuang mapping is used for utilizing international phonetic symbol correspondence, and the Chinese-Zhuang translation and speech synthesis module is used for realizing Chinese-Zhuang translation and speech synthesis. The tone characteristics of Chinese can be reasonably converted into the tone mode of Zhuang language, and the implementation of the speech synthesis module comprises Tacotron2 model processing and a vocoder. Compared with the prior art, the method has the advantages that a seamless conversion system platform from Chinese text to Zhuang language speech synthesis is provided, the minority language is protected and inherited, convenient speech interaction experience is provided for Zhuang language users, the accessibility of information is enhanced, and the user experience is improved. The method is especially suitable for visually impaired people or scenes requiring hand-eye liberation.
Owner:GUANGXI UNIV FOR NATITIES +1

Courseware making method and device based on artificial intelligence and storage medium

PendingCN122002103AImprove participationEnhance sense of accomplishmentSelective content distributionElectrical appliancesMachine learningData science
The invention relates to the technical field of artificial intelligence, in particular to a courseware making method and device based on artificial intelligence and a storage medium, and the method comprises the steps: constructing a tone problem distribution table based on tone accuracy scores and machine bias labels; determining a target training group based on the tone question distribution table, further determining a courseware generation parameter, extracting a recent continuous practice sequence based on the target training group, and further generating a training intensity adjustment parameter to adjust the courseware generation parameter; constructing an individual cognitive bias error model based on listen-and-read entries in a user historical period, and generating a bias error pair list; generating training content based on the courseware generation parameters, the target training group, the bias error pair list and the tone problem distribution table; and arranging a teaching process based on the training content, and generating a courseware script. According to the invention, the pronunciation data, the personal bias mode and the linguistic knowledge base of the learner can be dynamically combined, and an accurate closed loop from data diagnosis to targeted teaching is realized.
Owner:BEIJING QIXING INTERACTIVE EDUCATION TECHNOLOGY CO LTD

Multi-language voice test method and device, computer equipment and storage medium

The embodiment of the invention discloses a multi-language voice testing method and device, computer equipment and a storage medium, and relates to the technical field of voice testing. The method comprises the following steps: acquiring standard pronunciation data and accent feature data; extracting frequency spectrum features and tone features of the accent feature data; determining an accent model according to the spectrum features and the tone features; generating an accent voice test instruction based on the accent model and the standard pronunciation data through a preset neural voice synthesizer; and testing a voice assistant based on the accent voice test instruction to obtain a test result. According to the automatic voice test process constructed by the invention, the systematic bottleneck that the traditional voice test depends on manual recording is fundamentally solved through multi-stage technology collaboration. The core value of the method is that the voice test is converted into an extensible digital generation process from labor-intensive operation.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Multi-mode Chinese tone mouth shape auxiliary teaching system based on multiple attention mechanisms

The invention relates to the technical field of language teaching, and provides a multi-modal Chinese tone mouth shape auxiliary teaching system based on multiple attention mechanisms, and the system comprises a plurality of modules: a self-adaptive vocal cord vibration signal extraction module which is used for collecting and purifying a voice signal, extracting a fundamental frequency and scoring; the multi-scale tone feature fusion module receives the fundamental frequency signals and scores, extracts features and constructs vectors; the hierarchical tone recognition module processes the feature vectors through a neural network and an attention mechanism, recognizes tones and extracts voiceprint features; the personalized vocal organ motion modeling module selects and generates a personalized vocal organ motion model based on vocal print features; the mouth shape and tongue position dynamic generation module creates a three-dimensional visual mouth shape animation; the pronunciation feedback guidance module compares the pronunciation of the user with the standard pronunciation and provides multi-modal feedback; the tone recognition accuracy is improved through the adaptive signal extraction technology, and an effective auxiliary tool is provided for Chinese teaching.
Owner:KUNMING UNIV OF SCI & TECH

Dialect speech recognition method, system and model

The invention belongs to the technical field of speech recognition, and particularly relates to a dialect speech recognition method, system and model. The dialect speech recognition method comprises the following steps: S1, a plurality of teacher models generate pseudo tags and screen the pseudo tags; s2, training a sub-model; s3, data fusion; and S4, performing student model training, repeating the steps S3 to S4 until round fusion training is performed, and obtaining the student model which is trained in each region. According to the invention, through combination of multi-teacher model knowledge distillation and K-nearest neighbor dynamic data fusion training, an end-to-end voice transcription method suitable for a multi-dialect scene is constructed, and the method is especially suitable for a low-resource dialect environment with a complex tone system, rich voice variants and scarce data.
Owner:GUIZHOU UNIV +1

Sample set generation method, device and computer equipment for training speech recognition model

The present application relates to the technical field of speech recognition, and aims to solve the problem of lack of training samples of heavy accent speech recognition model. A method, device and computer equipment for generating a sample set for training a speech recognition model are provided, wherein the method comprises: decoding a target command word into a toneless original pinyin sequence; based on a heavy accent rule library (heavy accent refers to non-standard pronunciation of initial / final change) constructed according to common non-standard pronunciation rules, a heavy accent pinyin sequence is generated through coding conversion; a heavy accent audio is generated through a text-to-speech audio generation tool; an audio input is input into a preset recognition model to obtain a recognition result and convert it into a recognized pinyin sequence; a screening rule is constructed based on the original and / or heavy accent pinyin, and the recognized pinyin sequence is compared to screen the audio; and the audio meeting the requirements is collected to form a training sample set. Through rule-based generation and accurate screening, high-quality heavy accent samples can be efficiently obtained, and the heavy accent recognition performance of the model can be improved.
Owner:深圳市友杰智新科技有限公司

Multi-dimensional voice training plan generation method and system

The invention discloses a method and a system for generating a multi-dimensional voice training plan. The method comprises the following steps: firstly, collecting age, gender and voice samples of a patient, and obtaining seven-dimensional evaluation parameters including sound pressure, amplitude perturbation, maximum sounding duration, fundamental frequency perturbation, vowel space area, intonation damage and speech understanding degree score; and then, according to a preset logic, sequentially performing judgment based on the parameters, and dynamically combining different training modules of loudness, breath, pitch, vowel, gliding, tone, consonant and the like into a personalized voice training plan. Through multi-dimensional evaluation and dynamic module matching, the problem of training scheme solidification in the prior art is solved, and the individuation degree and rehabilitation effect of voice training are remarkably improved.
Owner:BEIJING REHABILITATION HOSPITAL CAPITAL MEDICAL UNIVERSITY(BEIJING WORKERS SANATORIUM)

Device for changing tones of trumpet

The utility model provides a trumpet tone changing device which comprises a trumpet sound tube and first to third pistons which are arranged side by side and are communicated through different movable channels, the third piston is communicated with a blowing nozzle guide tube, and first to third adjusting tubes are respectively and correspondingly arranged on the first to third pistons in a telescopic and penetrating mode. The horn sound tube is bent from front to back to be connected with a bent connection tube, a tone tube connection portion is arranged on the first piston in a backward extending mode, tone tubes with different tones are wound according to the lengths needed by the different tones, and the two ends of each tone tube extend forwards and can be connected between the bent connection tube and the tone tube connection portion respectively. After the tone tubes with different tones are replaced by the trumpet, the trumpet can make different tones with the same fingering.
Owner:HOXON GAKKI CORP

A deep learning-based pronunciation evaluation scoring method

The present application relates to the technical field of speech evaluation, in particular to a pronunciation evaluation and scoring method based on deep learning.The present application uses a speech recognition model to recognize the true text result of the audio.Then an HMM-DNN model is used to obtain the posterior probability of the audio.Finally, a scoring model is used to score the phonemes.Before forced alignment, the speech recognition model is used to identify the correct text of the audio, avoiding the situation that the audio and the text are inconsistent and cannot be aligned to the correct position during the forced alignment process.Meanwhile, the scoring model is constructed using a deep neural network, which can fit multiple information such as posterior probability, initial and final consonants, part of speech, tone, pronunciation duration, etc., making the phoneme scoring more reasonable and more accurate.
Owner:SUZHOU ZHIYAN INFORMATION TECH CO LTD

Intelligent speech recognition method and device, electronic equipment and storage medium

The present application relates to speech processing technology, disclose a kind of intelligent speech recognition method, comprising: obtaining speech signal, the frame processing of speech signal is carried out, and the frame speech signal of obtaining is obtained;The characteristic parameter of the frame speech signal is analyzed, and the phoneme information parameter and the tone information parameter are obtained;According to the phoneme information parameter and the tone information parameter frame synchronization word recognition is carried out to the frame speech signal, and the frame word sequence of obtaining is obtained;Inquiry the starting end and the end of each sentence in the frame word sequence, the frame word sequence is split into the word sequence segment under each sentence;Word reorganization is carried out to the word sequence segment in turn using pre-constructed finite state machine, and the recognition sentence is generated.The present application also proposes a kind of intelligent speech recognition device, electronic equipment and storage medium.The present application can solve the problem that there is sequence disorder in speech recognition result.
Owner:PING AN TECH (SHENZHEN) CO LTD