Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9 results about "Phoneme recognition" patented technology

Phoneme recognition is carried out using the acoustic model. The acoustic model is created using machine learning algorithms. The machine learning is divided into two phases: training and testing.

Phoneme-based speech recognition method and device of pronunciation-free dictionary, computer equipment and storage medium

PendingCN121768394Aimprove accuracyImprove adaptabilitySpeech recognitionPhoneme recognitionAgrammatism
The invention relates to a phoneme-based speech recognition method and device of a pronunciation-free dictionary, computer equipment and a storage medium. The method is applied to a server, a speech recognition framework is deployed in the server, the speech recognition framework comprises a phoneme recognition model, a phoneme-to-character conversion model and an auxiliary character-to-phoneme conversion model, and the method comprises the following steps: decoding speech information to be recognized based on the phoneme recognition model to generate a plurality of candidate phoneme sequences; decoding the plurality of candidate phoneme sequences based on a phoneme-word conversion model to generate a plurality of candidate text sequences; scoring the plurality of candidate text sequences according to an auxiliary character-to-pronunciation conversion model to obtain a speech recognition result; wherein the speech recognition framework is obtained by training an initial speech recognition framework based on a joint random approximation (JSA) algorithm in advance. And the accuracy of speech recognition is improved.
Owner:TSINGHUA UNIVERSITY +1

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342BResolve driver conflictsImprove viewing experienceTelevision system detailsColor television detailsPhoneme recognitionEngineering
The application discloses a virtual anchor real-time driving system based on facial motion capture, aims to solve the problems of insufficient precision, lack of naturalness and weak scene adaptation of virtual anchor facial motion in the prior art, and through multi-thread parallel collection of audio and video streams, optimizes data quality through adaptive preprocessing; adopts multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face function area semantic segmentation and confidence dynamic weight adjustment technology, realizes accurate representation and collaborative fusion of expression and lip feature; finally, through the lip-expression collaborative driving mechanism, the real-time driving characteristics of the adaptive virtual image are output; thereby effectively improving the precision and naturalness of virtual anchor facial motion, enhancing the complex scene adaptability, giving consideration to real-time response and low-cost deployment, and being applicable to virtual live broadcast, online education and other scenes, and having good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

Oral English test system and method based on multi-dimensional factors

ActiveCN120808759Bovercome limitationsImprove authoritySpeech recognitionPhoneme recognitionText matching
The application discloses a spoken language evaluation system and method based on multiple dimensions, and relates to the technical field of data processing.The system comprises the following steps: obtaining audio text, a phoneme list and a phoneme time boundary list corresponding to speech data according to a pre-trained text phoneme recognition model; calculating a fluency score according to the phoneme list and the phoneme time boundary list; obtaining a semantic score of the audio text through a pre-trained semantic model; obtaining a syntax score of the audio text through a pre-trained syntax model; matching the audio text and an answer text according to a text matching method based on an edit distance to obtain a text matching score; obtaining a pronunciation score of the audio data finally through phoneme confidence of the phoneme list; dynamically adjusting the weight of each dimension score according to the length of the audio text, and calculating a final spoken language evaluation score; and the application comprehensively evaluates the spoken language ability through multi-dimensional evaluation and dynamic weight adjustment according to the scene, and realizes the comprehensiveness, fairness, accuracy and high efficiency of the spoken language evaluation.
Owner:读书郎教育科技有限公司

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342Aexact matchPrecise linkageTelevision system detailsColor television detailsPhoneme recognitionEngineering
The invention discloses a virtual anchor real-time driving system based on facial motion capture, and aims at solving the problems that in the prior art, virtual anchor facial motions are insufficient in accuracy, lack of natural sense and weak in scene adaptation. Audio and video streams are collected in parallel through multiple threads, and the data quality is optimized through self-adaptive preprocessing; the technologies of multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face functional region semantic segmentation and confidence coefficient dynamic weight adjustment are adopted, and accurate representation and collaborative fusion of expression and mouth shape features are achieved; and finally, outputting real-time driving characteristics matched with the virtual image through a mouth shape-expression cooperative driving mechanism. Therefore, the accuracy and the natural sense of the face action of the virtual anchor are effectively improved, the adaptability of complex scenes is enhanced, real-time response and low-cost deployment are both considered, and the method is suitable for scenes such as virtual live broadcast and online education and has good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

Audio data rhythm analysis and playing method and device and storage medium

The present application relates to the field of audio analysis, and discloses a kind of audio data's tone analysis and play method, equipment and storage medium.The method comprises: receiving audio data;According to preset phoneme recognition algorithm, the phoneme recognition processing is carried out to the audio data, and phoneme sequence is obtained;According to preset tone rhythm data, the phoneme sequence is divided and marked, and tone sequence is obtained;Using preset description library, the tone sequence is described and matched, and the description text corresponding to the audio data is obtained;When the audio data plays, based on the playing position of the audio data, the text field corresponding to the playing position in the description text is shown.In the embodiment of the present application, the technical problem that the current analysis technology lacks tone analysis of audio and cannot meet the demand of people for audio music appreciation is solved.
Owner:SHENZHEN TONGXINGZHE TECH

Speech phoneme recognition method, medium, apparatus, and computing device

ActiveCN116453504BSpeech recognitionPhoneme recognitionAcoustics
Embodiments of the present disclosure provide a speech phoneme recognition method, medium, device and computing device. The method comprises: inputting speech data to be recognized into a pre-trained phoneme recognition model, outputting a phoneme sequence corresponding to the speech data, and the phoneme sequence comprising the occurrence order of each phoneme in the speech data. The present disclosure solves the problem that speech recognition in the related art cannot effectively match the lip movement of a virtual character image and speech. The speech data is disassembled into a set of phonemes that can be represented by the lip movement of AI, and is represented in a sequence form, so that the AI can perform representation through lip movement in turn according to the duration of each phoneme by reading the phoneme sequence, thereby achieving accurate matching of lip movement and speech, and significantly improving the experience of the audience.
Owner:HANGZHOU NETZHIYI INNOVATION TECH CO LTD

Digital population profile driven method, system and media based on phoneme recognition

PendingCN122290625APhoneme recognitionFeature extraction
This application provides a digital human mouth shape driving method, system, and medium based on phoneme recognition. The method includes: acquiring input raw speech signals and processing the raw speech signals; performing feature extraction and phoneme recognition on preprocessed data information to obtain accurate phoneme category and time sequence information; constructing a preset phoneme and mouth shape parameter mapping library, the mapping library including mouth shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles; dynamically matching the phoneme time sequence information with the mapping library to obtain corresponding mouth shape driving parameters and muscle deformation parameters; inputting a digital human mouth and surrounding muscle model, driving the model to deform, obtaining the driving result, and performing synchronous processing of mouth shape and speech pronunciation on the model to obtain synchronization result information; through phoneme-level fine-grained recognition, dynamic mapping matching, and precise driving of muscle models, real-time synchronization of mouth shape and speech pronunciation is achieved, improving the naturalness and realism of digital human voice interaction.
Owner:SHENZHEN LUDIE SOFTWARE TECHNOLOGY CO LTD

Learning evaluation method and system for English word pronunciation through audio recognition

PendingCN121963789APinpoint weak areasavoid ambiguityData processing applicationsBiological modelsPhoneme recognitionSpeech sound
The invention relates to a learning evaluation method for English word pronunciation through audio recognition, which comprises the following steps: constructing a speech recognition system, preprocessing received user pronunciation audio, and then realizing phonetic symbol recognition of recognized words through vowel phonetic symbol recognition and screening and consonant phonetic symbol recognition and screening or full phoneme recognition. The word recognized by the final recognition phonetic symbol sequence obtained by the speech recognition system is compared with the stored standard pronunciation of the word, the pronunciation of the practicer is scored, and specific practice suggestions are given.
Owner:LAIWU VOCATIONAL & TECHNICAL COLLEGE

Data calibration method, device, equipment and medium

PendingCN122024734ASpeech recognitionPhoneme recognitionPrediction probability
The invention relates to the technical field of data processing, and discloses a data calibration method and device, equipment and a medium, and the method comprises the steps: carrying out the phoneme recognition of a target audio through a target phoneme recognition model, obtaining the prediction phoneme and prediction probability of each audio frame, and obtaining the prediction probability of each audio frame according to a preset mapping table; determining a mapping phoneme of each sub-word in the transcription text corresponding to the target audio, for any audio frame, according to the predicted phoneme of the audio frame, the prediction probability corresponding to the predicted phoneme and the mapping phoneme of each sub-word, respectively matching the audio frame with each sub-word, and according to a matching result, determining a target sub-word corresponding to the audio frame, and according to the mapping phonemes of the target sub-words, the predicted phonemes of the audio frames are calibrated, and calibration results of the predicted phonemes of all the audio frames are fused to form a final calibration result. The method can be applied to the field of financial science and technology, the accuracy of phoneme recognition is improved, and a more reliable acoustic representation basis is provided for a downstream speech processing task of phoneme recognition.
Owner:PING AN TECH (SHENZHEN) CO LTD