Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Grapheme" patented technology

In linguistics, a grapheme is the smallest unit of a writing system of any given language. An individual grapheme may or may not carry meaning by itself, and may or may not correspond to a single phoneme of the spoken language.

HMM decoding compensation for speech recognition and multi-structured decoding for low-resource command detection

Techniques are described for recognizing a spoken wake word (WW) or a spoken command for a human-machine interface using a speech recognition system that does not require WW / command matching speech data for training. The system uses the text or grapheme representation of the WW or commands for pre-deployment training. The technique involves receiving a target phrase for recognition by a speech recognition model. It also involves analyzing a sequence of acoustic units representative of the target phrase as it is spoken to generate offline analysis data. Furthermore, the technique includes building the speech recognition model based on this offline analysis data to decode speech signals of the target phrase according to the acoustic units.The technique also includes processing speech based on the speech recognition model to detect the presence of the target phrase.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

Systems and methods for grapheme-phoneme correspondence learning

Systems and methods are described for grapheme-phoneme correspondence learning. In an example, a display of a device is caused to output a grapheme graphical user interface (GUI) that includes a grapheme. Audio data representative of a sound made by the human user is received based on the grapheme shown on the display. A grapheme-phoneme model can determine whether the sound made by the human corresponds to a phoneme for the displayed grapheme based on the audio data. The grapheme-phoneme model is trained based on augmented spectrogram data. A speaker is caused to output a sound representative of the phoneme for the grapheme to provide the human with a correct pronunciation of the grapheme in response to the grapheme-phoneme model determining that the sound made by the human does not correspond to the phoneme for the grapheme.
Owner:617 EDUCATION INC

Legend recognition method, device and electronic equipment

The application relates to the technical field of computer-aided design, in particular to a legend recognition method and device and electronic equipment, the method comprising the following steps: obtaining a target topology relationship of a target legend and a to-be-recognized legend, the target topology relationship being the positional relationship between a first target shape category formed by target graphemes in the target legend and other graphemes in the target legend; recognizing the shape category formed by each grapheme in the to-be-recognized legend, and determining all shape categories in the to-be-recognized legend; searching for a second target shape category identical to the first target shape category in all the shape categories; determining a to-be-recognized topology relationship based on the second target shape category and the positional relationship of other graphemes in the to-be-recognized legend; and matching the target topology relationship with the to-be-recognized topology relationship to determine a legend recognition result. The shape category formed by graphemes is used as the basis for legend recognition, which guarantees the accuracy of legend recognition, and the shape category is used to reduce the matching range, thereby improving the efficiency of legend recognition.
Owner:GLODON CO LTD

LANGUAGE-INDEPENDENT DICTIONARY-TRAINED GRAPHEME-TO-PHONEME CONVERTER AND TEXT-TO-SPEAK MACHINE FOR IMPROVED SPEECH RECOGNITION

Techniques are described for recognizing a spoken wake word (WW) or a spoken command for a human-machine interface using a speech recognition system that does not require WW / command matching speech data for training. The technique trains a tokenizer and involves decomposing a word from a database into a multitude of combinations of unique subwords by splitting the word at a variety of different points for each combination. The word comprises one or more written units and one or more corresponding acoustic units. The technique maps the acoustic units that make up the word to the written units that make up the word to create an acoustic unit-to-written unit mapping.The technique involves assigning a subset of acoustic units to each of the unique subwords based on the acoustic unit-to-written unit mapping, in order to generate an acoustic unit-to-subword mapping for the word. The technique accumulates the acoustic unit-to-subword mappings for a large number of words from the database to create a subword probability dictionary.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

Language-independent dictionary-trained grapheme-to-phoneme converter and text-to-speech engine for improved speech recognition

A language-independent dictionary-trained grapheme-to-phoneme converter and text-to-speech engine for improved speech recognition is disclosed. Techniques are described for recognizing a spoken wakeup word (WW) or command using a speech recognition system that does not need to be trained with any speech data that matches the WW / command for a human machine interface. Techniques train a word segmentation device and include decomposing a word into a plurality of combinations by splitting the word from a database at a plurality of different points for each of the plurality of combinations of unique sub-words. The word includes one or more writing units and one or more corresponding acoustic units. Techniques map acoustic units constituting a word to writing units constituting the word to generate an acoustic unit-to-writing unit mapping. The technique includes assigning a subset of the acoustic units to each of the unique sub-words based on the acoustic unit-to-write unit map to generate an acoustic unit-to-sub-word assignment of the word. Techniques accumulate acoustic units to sub-word assignments for a plurality of words from a database to create a sub-word likelihood dictionary.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

DATA-FREE LANGUAGE RECOGNITION

Techniques are described for recognizing a spoken wake word (WW) or a spoken command for a human-machine interface using a speech recognition system that does not require WW or command-matching speech data for training. The system uses the text or grapheme representation of the WW or commands for pre-deployment training. The technique involves the system receiving a text representation of a target phrase in a target language. It includes training an acoustic model based on a speech database to distinguish speech signals according to the acoustic units of the target language. The training of the acoustic model is independent of the target phrase.It involves creating a recognition model based on the textual representation of the target phrase and the acoustic model to recognize the target phrase in speech, and processing speech from a speaker based on the acoustic model and the recognition model to detect the presence of the target phrase.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

A training sample generation method, model training method and device

ActiveCN115879002BQuality dataEngineering
Embodiments of the present application provide a training sample generation method and device, and a model training method, relating to the technical field of artificial intelligence, comprising: obtaining a plurality of original training samples; for each original training sample, obtaining video features and text features of the original training sample; determining a plurality of target image regions and a plurality of target graphemes in the original training sample based on the video features and the text features of the original training sample; calculating quality data of the original training sample based on the video features of the plurality of target image regions and the text features of the plurality of target graphemes; and selecting each original training sample based on a first number of original training samples with quality data less than a first preset threshold to obtain a target training sample. The embodiments can improve the training effect of the cross-modal model, save time and labor costs, improve the generation efficiency of the training sample, and further improve the training efficiency of the cross-modal model.
Owner:BEIJING IQIYI TECH CO LTD

System and method for reading development using word-to-graphic transformations

A system and method for reading development in children with dyslexia using word-to-graphic transformations are disclosed. The method includes presenting words to the learner and converting the words into corresponding graphical representations that visually depict the meaning of the words. The method further includes associating phonemes with parts of the graphical representation to help the learner connect speech sounds with written letters. Next, the method includes progressively simplifying the graphical representation of words as the learner's reading proficiency improves, transitioning them toward recognizing standard written words. The method also includes tracking the learner's progress through a proficiency tracking module, which dynamically adjusts the complexity of future word transformations based on performance data. Additionally, the method provides real-time auditory and visual feedback during the learning process to reinforce phoneme-grapheme associations. The system operates on an interactive platform, such as a tablet, allowing personalized learning experiences based on the child's individual progress.
Owner:MA KEVIN

Realistic Lip Synchronization for Artificial Intelligence-Powered Talking Avatars

PendingUS20260188297A1AnimationAutomatic speech
The system and method for generating realistic lip synchronization for an AI avatar. The lip synchronization process begins by receiving input data. The input data can either be text input or audio stream. If the input data is text input, a text-to-speech (TTS) module converts it into speech while generating word-level timestamps. If the input data is the audio stream, an automatic speech recognition (ASR) module transcribes the spoken content and provides word-level timestamps. The transcribed text is transformed into a sequence of phonemes using a grapheme-to-phoneme conversion system. The phonemes are mapped to visemes based on their corresponding mouth shapes using a predefined phoneme to viseme mapping table. Then the visemes are mapped to the blendshapes of the AI avatar using image similarity comparison. The selected blendshapes are animated to generate synchronized animation of the AI avatar, with transitions smoothed to ensure natural lip movements and facial expressions.
Owner:2HR LEARNING INC

HMM DECODING WITH ACOUSTIC MODEL COMPENSATION FOR PHONE MODELING AND SPEECH MODELING

Techniques are described for recognizing a spoken wake word (WW) or a spoken command for a human-machine interface using a speech recognition system that does not require WW or command-matching speech data for training. The system uses the text or grapheme representation of the WW or commands for pre-deployment training. The technique includes generating a statistical matrix that characterizes a phonetic model of an acoustic model that distinguishes speech signals according to a variety of acoustic units of a language. It further includes creating a speech recognition model based on the statistical matrix to recognize a target phrase in the speech. It also includes processing an audio signal based on the acoustic model and the speech recognition model to detect the presence of the target phrase.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

Speech synthesis method and device, electronic equipment and storage medium

The application provides a speech synthesis method and device, electronic equipment and storage medium, and relates to the technical field of speech synthesis. The method first acquires a text to be synthesized, then determines pronunciation features of the text to be synthesized according to loaded language rules, and then performs speech synthesis on the text to be synthesized according to the pronunciation features through a speech synthesis model to obtain target synthesized speech. The pronunciation features of the text to be synthesized can be determined from the dimensions of word formation characteristics, context dependence between graphemes of the text to be synthesized and the like according to the language rules, so that the pronunciation features of the text to be synthesized are more in line with language characteristics, thereby increasing the accuracy of the obtained pronunciation features of the text to be synthesized. The speech synthesis model performs speech synthesis on the text to be synthesized according to the pronunciation features with higher accuracy, thereby improving the quality of speech synthesis.
Owner:IFLYTEK CO LTD

Data-free speech recognition

The invention discloses data-free speech recognition. Techniques are described for recognizing a spoken wakeup word (WW) or command using a speech recognition system that does not need to be trained with any speech data that matches the WW or command for a human machine interface. The system is trained using a text or word representation of a WW or command prior to deployment. The technique includes receiving, by a system, a textual representation of a target phrase in a target language. The technique includes training an acoustic model based on a speech database to distinguish speech signals according to acoustic units of a target language. The training of the acoustic model is independent of the target phrase. The technique includes constructing a recognition model to recognize a target phrase in speech based on a textual representation of the target phrase and an acoustic model; and processing the speech from the speaker based on the acoustic model and the recognition model to detect the presence of the target phrase.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

Phoneme-based speech domain transfer method, system, and electronic device

Embodiments of the present application provide a phoneme-based speech domain migration method, system and electronic device. The method comprises: performing grapheme-to-phoneme conversion on target domain text to obtain a target domain phoneme sequence; converting the target domain phoneme sequence into a plurality of target domain phoneme N-gram sequences according to a phoneme N-gram dictionary; generating target domain speech segments with speaker diversity using the plurality of target domain phoneme N-gram sequences; and generating synthesized audio for the target domain based on the target domain speech segments. Embodiments of the present application use a dictionary constructed from speech segments of basic phoneme n-grams to generate speech for the text of the target domain, the splicing and synthesis method guided by phonemes has the ability to model the connected reading between words, and because the audio segments for splicing and synthesis come from a large amount of real speech, the synthesized speech has better speaker diversity, and the required computing resources are reduced, avoiding overfitting of the ASR model on the synthesized data.
Owner:AISPEECH CO LTD

Speech synthesis method and device, electronic equipment and storage medium

The invention provides a speech synthesis method and device, electronic equipment and a storage medium, and relates to the technical field of speech synthesis, and the method comprises the steps: firstly obtaining a to-be-synthesized text, then determining the pronunciation characteristics of the to-be-synthesized text according to a loaded language rule, and then carrying out the speech synthesis of the to-be-synthesized text according to the pronunciation characteristics through a speech synthesis model, according to the method, the pronunciation characteristics of the to-be-synthesized text can be determined from dimensions such as word formation characteristics and context dependence between the graphemes of the to-be-synthesized text according to language rules, so that the pronunciation characteristics of the to-be-synthesized text better conform to language characteristics, and the accuracy of the obtained pronunciation characteristics of the to-be-synthesized text is improved; the speech synthesis model carries out speech synthesis on the to-be-synthesized text according to the pronunciation features with higher accuracy, and the speech synthesis quality can be improved.
Owner:IFLYTEK CO LTD