Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20 results about "Viseme" patented technology

A viseme is any of several speech sounds that look the same, for example when lip reading (Fisher 1968). Visemes and phonemes do not share a one-to-one correspondence. Often several phonemes correspond to a single viseme, as several phonemes look the same on the face when produced, such as /k, ɡ, ŋ/, (viseme: /k/), /t͡ʃ, ʃ, d͡ʒ, ʒ/ (viseme: /ch/), /t, d, n, l/ (viseme: /t/), and /p, b, m/ (viseme: /p/). Thus words such as pet, bell, and men are difficult for lip-readers to distinguish, as all look like /pet/. However, there may be differences in timing and duration during actual speech in terms of the visual 'signature' of a given gesture that can not be captured with a single photograph. Conversely, some sounds which are hard to distinguish acoustically are clearly distinguished by the face (Chen 2001). For example, acoustically speaking English /l/ and /r/ can be quite similar (especially in clusters, such as 'grass' vs. 'glass'), yet visual information can show a clear contrast. This is demonstrated by the more frequent mishearing of words on the telephone than in person. Some linguists have argued that speech is best understood as bimodal (aural and visual), and comprehension can be compromised if one of these two domains is absent (McGurk and MacDonald 1976).

Intelligent question-answering method, device, equipment, and storage medium based on front and back nasal sounds

ActiveCN117636853BExcellent speech intent text outputIncrease the efficiency of intelligent question answeringSpeech recognitionFeature extractionRhinolalia
This invention relates to artificial intelligence technology and discloses an intelligent question-answering method, apparatus, device, and storage medium based on nasal sounds. The method includes: splitting user-input speech audio data to obtain a phoneme sequence; converting the phoneme sequence into initials and finals according to nasal sound conversion rules to obtain a phoneme arrangement matrix, and quantizing it to obtain a phoneme vector matrix; performing a fully connected operation on each phoneme vector in the phoneme vector matrix to obtain a set of combined phoneme sequences; performing feature extraction and word segmentation recognition operations on each combined phoneme sequence in the combined phoneme sequence set to obtain a word segmentation arrangement matrix corresponding to each combined phoneme sequence; performing word segmentation and arrangement on each word segmentation arrangement matrix, identifying the completeness of the intent in the word segmentation and arrangement results to obtain a complete intent statement; and filtering the complete intent statement according to current business scenario information to obtain the speech intent text. This invention can increase the accuracy of speech recognition and improve the efficiency of intelligent question answering.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

An input method, device and system based on number sequence and number sequence rime

PendingCN122633055ASymbol mappingWord list
The application discloses an input method, device and system based on number sequence and number sequence vowel definition and symbol mapping, which comprises the following steps: defining a number sequence and a number sequence vowel set, wherein the number sequence is a fixed sequence of numbers 1-9, 0 and an extended character, the number sequence vowel set comprises 11 groups of vowels arranged according to the number sequence, and each group of vowels corresponds to a unique number sequence coding bit; establishing a mapping relationship between the 11 groups of vowels and a number sequence input unit, wherein the number sequence input unit comprises at least 10 digital key positions and an extended character key position arranged according to the number sequence; and generating a candidate word list according to the defined number sequence and number sequence vowel set and the mapping relationship in response to a user input operation. The application solves the problem of inconsistent coding logic and multiple keystrokes of vowels in the prior art, significantly improves the input efficiency of Chinese pinyin, supports multiple input modes to share a unified coding system and reduces the learning cost of users.
Owner:梁晨

An input method, device and system based on number sequence finals

PendingCN122653452AWord listEngineering
The application discloses an input method, device and system based on number sequence final consonant definition and column direction mapping. The method comprises the following steps: according to the consistency principle of pronunciation tongue position, merging final consonants in the Chinese pinyin scheme into 11 groups of number sequence final consonants, establishing a column direction mapping relationship between the 11 groups of number sequence final consonants and 11 column character keys of a keyboard, so that the character final consonants on each column character key are merged into the corresponding number sequence final consonants through column direction compression; in response to the input operation of a user, converting the character final consonants into corresponding number sequence final consonant codes according to the mapping relationship, and generating a candidate word list. The application realizes single key triggering of final consonants by systematically merging final consonants into 11 groups and establishing a determined mapping with the column direction of the keyboard, solves the problems of multiple key strokes and inconsistent coding logic in the prior art, significantly improves the input efficiency of the Chinese pinyin, supports multiple input modes to share a unified coding system, and reduces the learning cost of users.
Owner:梁晨

Speech recognition method, system and terminal

The present disclosure provides a multi-functional speech recognition method, system and terminal. The present disclosure constructs an end-to-end speech recognition system with pinyin initial and final consonants as modeling units, and adds a modeling unit probability output module, effectively improving non-close sound character replacement errors, and relative to existing end-to-end speech recognition systems, adding various functions such as end-to-end speech recognition system acoustic recognition performance evaluation and pronunciation standard degree evaluation. The method comprises: receiving a speech to be recognized; performing acoustic feature extraction and encoding on the speech to be recognized; decoding the encoded acoustic features using a Chinese character decoder, wherein the Chinese character decoder uses pinyin initial and final consonants as modeling units, and maps the encoded acoustic feature sequence to a Chinese character sequence through initial and final consonants; and outputting a speech recognition result.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Method and system for converting natural intonation into melody, terminal and medium

The invention provides a method and a system for converting natural intonation into melody, a terminal and a medium, which are characterized in that fundamental frequency and tone length information of a voice sample are extracted, initial and final consonants are distinguished, fundamental frequency parameters of each final consonant are calculated, and the tone tuning range of each final consonant in a statement is determined. And matching music symbols by using the tuning ranges to generate a preliminary melody, and finally aligning the melody with the syllables of the voice sample to generate an audio file with a song melody. On the basis of the phonetic theory, the fundamental frequency, the tone length and the pause interval information of the voice sample are accurately extracted, and it is ensured that the generated melody can accurately reflect the characteristics and rhythm of natural intonations. The automatic method not only simplifies the operation process and improves the conversion efficiency, but also is suitable for a plurality of fields such as linguistics research, speech synthesis, language teaching, language training and pronunciation improvement.
Owner:GUANGMING HOSPITAL OF TRADITIONAL CHINESE MEDICINE PUDONG NEW AREA SHANGHAI

Sample set generation method, device and computer equipment for training speech recognition model

ActiveCN121438814BSpeech recognitionEngineeringCode conversion
The present application relates to the technical field of speech recognition, and aims to solve the problem of lack of training samples of heavy accent speech recognition model. A method, device and computer equipment for generating a sample set for training a speech recognition model are provided, wherein the method comprises: decoding a target command word into a toneless original pinyin sequence; based on a heavy accent rule library (heavy accent refers to non-standard pronunciation of initial / final change) constructed according to common non-standard pronunciation rules, a heavy accent pinyin sequence is generated through coding conversion; a heavy accent audio is generated through a text-to-speech audio generation tool; an audio input is input into a preset recognition model to obtain a recognition result and convert it into a recognized pinyin sequence; a screening rule is constructed based on the original and / or heavy accent pinyin, and the recognized pinyin sequence is compared to screen the audio; and the audio meeting the requirements is collected to form a training sample set. Through rule-based generation and accurate screening, high-quality heavy accent samples can be efficiently obtained, and the heavy accent recognition performance of the model can be improved.
Owner:深圳市友杰智新科技有限公司

Chinese speech signal segmentation method, device and equipment and storage medium

This disclosure provides a method, apparatus, device, and storage medium for segmenting Chinese speech signals. The method includes: sampling a target audio signal containing speech corresponding to a target Chinese text to obtain signal amplitudes corresponding to multiple sampling points; performing speech endpoint detection on the target audio data based on the signal amplitudes corresponding to the multiple sampling points to obtain multiple speech segments in the target audio data; determining the vowel position sequence corresponding to the target speech segment based on the formant energy of the speech signal in the target speech segment; determining syllable segmentation points and initial / final segmentation points of the target speech segment based on the signal amplitudes of sampling points between two adjacent vowel positions in the vowel position sequence and the initials and finals of the corresponding Chinese text segment; and segmenting the target speech segment based on the syllable segmentation points and initial / final segmentation points to obtain segmented speech primitives. This disclosure improves the accuracy of speech signal segmentation.
Owner:BEIJING INFORMATION TECH COLLEGE

System and method of modulating animation curves

A system and method of modulating animation curves based on audio input. The method including: identifying phonetic features for a plurality of visemes in the audio input; determining viseme animation curves based on parameters representing a spatial appearance of the plurality of visemes; modulating the viseme animation curves based on melodic accent, pitch sensitivity, or both, based on the phonetic features; and outputting the modulated animation curves.
Owner:JALI INC

A rhyme word recommendation method, device, equipment and storage medium

ActiveCN113850080BSemantic analysisEngineeringViseme
The present disclosure provides a rhyming word recommendation method and device, equipment and storage medium, relates to the technical field of computers, and particularly relates to the technical field of intelligent recommendation. The specific implementation scheme is as follows: obtaining a to-be-matched rhyming word and a rhyming configuration parameter; determining a target foot of a sentence to which the to-be-matched rhyming word belongs according to the to-be-matched rhyming word; determining a target rhyming word class corresponding to the target foot in a rhyming word class included in a rhyming word library, the feet of the sentences to which the rhyming words in each rhyming word class belong are identical in tone; determining a rhyming word matched with the to-be-matched rhyming word from the target rhyming word class according to the rhyming configuration parameter; and recommending the determined rhyming word. The scheme provided in the present disclosure can realize rhyming word recommendation.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Classical poem recitation assisting method fusing initial and final analysis

The invention discloses a classical poem recitation assisting method fusing initial and final analysis, and relates to the technical field of digital education, and the method comprises the steps: achieving the intelligent judgment and prompt of recitation rhythm and escort state based on the structural collection and analysis of recitation audio signals; the method comprises the following steps: acquiring and dividing a recitation audio into word segments through a voice recognition technology, and generating corresponding audio feature data information and text feature data information in combination with a poem structure; then judging whether the recitation duration is consistent with the poem structure or not through sentence level comparison, and providing a duration adjustment suggestion when the recitation duration is abnormal; on the basis that the recitation rhythm is normal, the vowel acoustic vector of each character is further extracted, and multi-level escort auxiliary suggestions are generated through initial and final difference analysis of the last character of adjacent sentences; according to the invention, through a process closed loop combining automatic acquisition, intelligent analysis and real-time feedback, the escort sensitivity and rhythm control ability of a learner are significantly improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Sample set generation method and device for training speech recognition model, and computer equipment

The invention relates to the technical field of speech recognition, and aims to solve the problem that training samples of a heavy accent speech recognition model are deficient. The invention provides a sample set generation method and apparatus for training a speech recognition model, and a computer device. The method comprises the steps of decoding a target command word into a silent original pinyin sequence; the method comprises the following steps of: performing code conversion on a duplicate accent rule base (duplicate accent refers to non-standard pronunciation of initial consonant / vowel change) constructed based on a common non-standard pronunciation rule to generate a duplicate accent pinyin sequence; generating a re-accent audio through a text-to-speech audio generation tool; inputting the audio into a preset recognition model to obtain a recognition result and converting the recognition result into a recognition pinyin sequence; constructing a screening rule by taking the original and / or double accent pinyin as a reference, and comparing and identifying the pinyin sequence to screen the audio; collecting audios meeting requirements to form a training sample set; according to the invention, through regularization generation and accurate screening, a high-quality heavy accent sample is efficiently obtained, and the heavy accent recognition performance of the model can be improved.
Owner:深圳市友杰智新科技有限公司

Tibetan written language character-voice conversion method based on voice structure characteristics

The invention, which belongs to the technical field of data processing, discloses a Tibetan language written language character-voice conversion method based on voice structure characteristics, comprising the following steps: S1, obtaining a Tibetan language text; s2, performing text segmentation processing on the Tibetan text by taking syllables as units, and extracting font structure elements of each syllable based on the font structure of the Tibetan to obtain a syllable font structure element set; s3, performing matching processing on initial consonant items and final consonant items of each syllable based on the syllable font structure element set to obtain a voice structure feature vector of each syllable; and S4, obtaining an initial consonant conversion coding table and a vowel conversion coding table, and mapping the voice structure feature vector of each syllable into a corresponding syllable phoneme sequence, thereby obtaining the Tibetan written language audio. According to the method, the problems that the basic characteristics of the Tibetan speech cannot be accurately described in a Tibetan language Latin transcription mode in the existing Tibetan speech synthesis technology and the like, and the calculation amount is too large and the character and speech conversion speed is low due to multilayer combination of Tibetan language phoneme sequences are solved.
Owner:TIBET UNIV

Realistic Lip Synchronization for Artificial Intelligence-Powered Talking Avatars

PendingUS20260188297A1AnimationAutomatic speech
The system and method for generating realistic lip synchronization for an AI avatar. The lip synchronization process begins by receiving input data. The input data can either be text input or audio stream. If the input data is text input, a text-to-speech (TTS) module converts it into speech while generating word-level timestamps. If the input data is the audio stream, an automatic speech recognition (ASR) module transcribes the spoken content and provides word-level timestamps. The transcribed text is transformed into a sequence of phonemes using a grapheme-to-phoneme conversion system. The phonemes are mapped to visemes based on their corresponding mouth shapes using a predefined phoneme to viseme mapping table. Then the visemes are mapped to the blendshapes of the AI avatar using image similarity comparison. The selected blendshapes are animated to generate synchronized animation of the AI avatar, with transitions smoothed to ensure natural lip movements and facial expressions.
Owner:2HR LEARNING INC

Chinese speech signal segmentation method and device, equipment and storage medium

ActiveCN121306099ASpeech recognitionSpeech synthesisFormantVowel
The embodiment of the invention discloses a Chinese speech signal segmentation method and device, equipment and a storage medium, and the method comprises the steps: carrying out the sampling of a target audio signal containing the speech corresponding to a target Chinese text, and obtaining signal amplitudes corresponding to a plurality of sampling points; performing voice endpoint detection on the target audio data based on the signal amplitudes corresponding to the plurality of sampling points to obtain a plurality of voice segments in the target audio data; determining a vowel position sequence corresponding to the target voice segment based on the formant energy of the voice signal in the target voice segment; determining syllable segmentation points and initial and final segmentation points of the target speech segment based on signal amplitudes of sampling points between two adjacent vowel positions in the vowel position sequence and initial and final consonants of the Chinese text segment corresponding to the target speech segment; and based on the syllable segmentation points and the initial and final segmentation points of the target voice segment, segmenting the target voice segment to obtain segmented voice elements. According to the embodiment of the invention, the accuracy of voice signal segmentation can be improved.
Owner:BEIJING INFORMATION TECH COLLEGE

High-simulation humanoid robot lip shape synchronous anthropomorphic control method and high-simulation humanoid robot lip shape synchronous anthropomorphic control system

The invention discloses a lip shape synchronous anthropomorphic control method and system for a high-simulation anthropomorphic robot. The method comprises the following steps: constructing a lip-shaped visual position mapping database of initial consonant phonemes and vowel phonemes; converting the pronunciation text into a pinyin sequence and identifying a pause symbol in the pinyin sequence; analyzing the pinyin into an initial consonant and vowel sequence, and extracting a corresponding lip action according to the mapping database; performing block processing on the action sequence according to the voice rhythm, and distributing the execution time of each action in combination with the audio duration; and a continuous lip shape control instruction sequence is generated by adopting a high-frequency interpolation mode, and is synchronously output in real time according to a time axis and audio playing. Through phoneme-level lip action driving, time fine distribution and continuous interpolation control, high consistency of voice content and lip actions in space and time dimensions is achieved, the anthropomorphic degree and the interaction reality sense of the humanoid robot in the language output process can be remarkably improved, and the user experience is improved. The lip-shaped driving system is suitable for various types of lip-shaped driving systems of humanoid robots.
Owner:WU XI WU JIE TAN SUO KE JI YOU XIAN GONG SI

Method for converting low-resource language voice into international phonetic symbol based on fused phonetic knowledge

The invention relates to a low-resource language voice-to-international phonetic symbol conversion method based on fused phonetic knowledge. According to the method, the technical threshold of translating extremely-low-resource voices into international phonetic symbols is reduced through tag injection, continuous embedding and phonetic system soft constraint decoding in a large model framework. By means of a lightweight adapter, the corpus can be migrated in about thirty hours, and the high cost of retraining the whole network is avoided; by means of an initial consonant-vowel-tone sign integrated label system and fusion of continuous fundamental frequency characteristics, the recognition precision of the system on phonemes and tones is synchronously improved. An interpretable result which conforms to the specification and is attached with confidence and a fundamental frequency curve is directly output, and the problems of illegal syllables and tone class confusion are thoroughly solved through joint decoding of initial consonant-vowel-tone sign transition probability and a language model. And a technical normal form with replicability is provided for digital continuation, cross-language information retrieval and cultural inheritance of endangered languages, and the method has important academic value and social significance.
Owner:INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES

Multi-syllable word disambiguation method and device, electronic equipment and readable storage medium

Embodiments of the present application provide a multi-sound character disambiguation method and device, electronic equipment and storage medium, including: obtaining attribute information of a target multi-sound character including mask information, word segmentation information, part-of-speech information and semantic information, inputting the attribute information into a Transformer encoder including an initial classifier, a final classifier and a tone classifier, splicing the output results to generate a first pinyin prediction result, and determining a final pinyin prediction result according to pinyin weight information of the target multi-sound character and the first pinyin prediction result. In the case of insufficient data or data imbalance, the initial classifier, the final classifier and the tone classifier are used to fully train the initial classifier, the final classifier and the tone classifier, so as to improve the multi-sound character prediction accuracy. At the same time, by increasing the pinyin weight information, the possible multi-sound character pronunciation can be limited in advance, so that the multi-sound character disambiguation prediction result is more accurate.
Owner:BEIJING SINOVOICE TECH CO LTD

Audio data processing methods, devices, and servers

This specification provides a method, apparatus, and server for processing audio data, applicable to the financial field. Based on this method, after receiving target audio data and obtaining the corresponding first data group through speech recognition, the first data group is first split according to a preset splitting rule to obtain multiple character matrices. Then, based on the pinyin data of the characters, the characters in the character matrices are mapped to corresponding letter character combinations to obtain the corresponding first letter character matrix. According to a preset exchange rule, letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined. The initial consonants and / or final vowels of the letter character combinations that satisfy the exchange conditions are then exchanged to obtain the corresponding second letter character matrix. Based on the second letter character matrix, the corresponding second data group is obtained. This method can efficiently hide relevant information carried in the audio data, protecting the data security of that information.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA