Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

34 results about "Viseme" patented technology

A viseme is any of several speech sounds that look the same, for example when lip reading (Fisher 1968). Visemes and phonemes do not share a one-to-one correspondence. Often several phonemes correspond to a single viseme, as several phonemes look the same on the face when produced, such as /k, ɡ, ŋ/, (viseme: /k/), /t͡ʃ, ʃ, d͡ʒ, ʒ/ (viseme: /ch/), /t, d, n, l/ (viseme: /t/), and /p, b, m/ (viseme: /p/). Thus words such as pet, bell, and men are difficult for lip-readers to distinguish, as all look like /pet/. However, there may be differences in timing and duration during actual speech in terms of the visual 'signature' of a given gesture that can not be captured with a single photograph. Conversely, some sounds which are hard to distinguish acoustically are clearly distinguished by the face (Chen 2001). For example, acoustically speaking English /l/ and /r/ can be quite similar (especially in clusters, such as 'grass' vs. 'glass'), yet visual information can show a clear contrast. This is demonstrated by the more frequent mishearing of words on the telephone than in person. Some linguists have argued that speech is best understood as bimodal (aural and visual), and comprehension can be compromised if one of these two domains is absent (McGurk and MacDonald 1976).

Audio-lip movement correlation measurement for dubbed content

Methods and apparatus are described for evaluating dubbing of media content. Phonemes in dubbed audio are extracted and mapped to visemes. Lip poses in video frames of the media content corresponding to the phonemes of the dubbed audio are compared to the visemes determined from the dubbed audio. A notification may be generated based on the comparison that indicates synchronization of the dubbed audio to lip poses of the video.
Owner:AMAZON TECH INC

Tone evaluation rehabilitation training device and system

A tone evaluation rehabilitation training device and system, the device comprises a flexible cap body and a bandage, the cap body is integrated with a 10-channel electrode plate to accurately cover language brain areas on both sides, electroencephalogram signals are transmitted in real time through Bluetooth 5.0, and a noise reduction earphone is used for providing sound output for a patient. The microphone is used for collecting audio of a patient and transmitting the audio to the processing module for analysis, the system dynamically regulates acoustic stimulation based on the neural activation feedback module, high-precision tone recognition is achieved in combination with audio preprocessing, SpecAugment data enhancement, a CNN-Transform mixed model and an attention mechanism, and training efficiency is optimized through pre-training model fine tuning and Warmup learning rate scheduling. An initial consonant-vowel-tone three-dimensional confusion probability table is introduced to divide interference intensity grades, high and low interference task paths are dynamically switched through a semantic error rate, and an evaluation module generates a four-tone accuracy rate, an F0 curve comparison graph and a personalized rehabilitation scheme. The recognition precision is improved, the training period of special crowds is shortened, and full-period self-adaptive rehabilitation is provided for preschool children to speech disorder patients.
Owner:林珍萍

Real-time extraction of 3D animation information from predicted speech

One or more inputs are processed with a machine-learned language model to obtain a prediction output. The inputs comprise speech information descriptive of one or more first words spoken by a user, and the prediction output comprises one or more second words predicted to follow the one or more first words. A sequence of visemes formed to produce the one or more second words is determined. Based on the sequence of visemes, facial animation information is generated descriptive of a facial animation that animates a three-dimensional representation of a mouth of the user forming the sequence of visemes to speak the one or more second words.
Owner:CHARTER COMM OPERATING LLC

Intelligent question-answering method, device, equipment, and storage medium based on front and back nasal sounds

ActiveCN117636853BExcellent speech intent text outputIncrease the efficiency of intelligent question answeringSpeech recognitionFeature extractionRhinolalia
This invention relates to artificial intelligence technology and discloses an intelligent question-answering method, apparatus, device, and storage medium based on nasal sounds. The method includes: splitting user-input speech audio data to obtain a phoneme sequence; converting the phoneme sequence into initials and finals according to nasal sound conversion rules to obtain a phoneme arrangement matrix, and quantizing it to obtain a phoneme vector matrix; performing a fully connected operation on each phoneme vector in the phoneme vector matrix to obtain a set of combined phoneme sequences; performing feature extraction and word segmentation recognition operations on each combined phoneme sequence in the combined phoneme sequence set to obtain a word segmentation arrangement matrix corresponding to each combined phoneme sequence; performing word segmentation and arrangement on each word segmentation arrangement matrix, identifying the completeness of the intent in the word segmentation and arrangement results to obtain a complete intent statement; and filtering the complete intent statement according to current business scenario information to obtain the speech intent text. This invention can increase the accuracy of speech recognition and improve the efficiency of intelligent question answering.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

Chinese character phrase and sentence decoding method based on stereotactic EEG signals

This invention discloses a method for decoding Chinese character phrases and sentences based on stereotactic electroencephalogram (sEEG) signals. This method uses stereotactic electroencephalogram (sEEG) signals as system input to decode Chinese phonemes, achieving greater generalization and adaptability than syllables. After identifying the initials, finals, and tones that make up Chinese Pinyin (Pinyin has 23 initials and no initials, 24 finals, and five tones), this method performs a two-way correction process using a language model modeled by Bayesian probabilistic modeling and a large language model. The decoded and corrected phrases or sentences are then displayed in real time on a user interface.
Owner:WESTLAKE UNIV

An input method, device and system based on number sequence and number sequence rime

PendingCN122633055ASymbol mappingWord list
The application discloses an input method, device and system based on number sequence and number sequence vowel definition and symbol mapping, which comprises the following steps: defining a number sequence and a number sequence vowel set, wherein the number sequence is a fixed sequence of numbers 1-9, 0 and an extended character, the number sequence vowel set comprises 11 groups of vowels arranged according to the number sequence, and each group of vowels corresponds to a unique number sequence coding bit; establishing a mapping relationship between the 11 groups of vowels and a number sequence input unit, wherein the number sequence input unit comprises at least 10 digital key positions and an extended character key position arranged according to the number sequence; and generating a candidate word list according to the defined number sequence and number sequence vowel set and the mapping relationship in response to a user input operation. The application solves the problem of inconsistent coding logic and multiple keystrokes of vowels in the prior art, significantly improves the input efficiency of Chinese pinyin, supports multiple input modes to share a unified coding system and reduces the learning cost of users.
Owner:梁晨

An input method, device and system based on number sequence finals

PendingCN122653452AWord listEngineering
The application discloses an input method, device and system based on number sequence final consonant definition and column direction mapping. The method comprises the following steps: according to the consistency principle of pronunciation tongue position, merging final consonants in the Chinese pinyin scheme into 11 groups of number sequence final consonants, establishing a column direction mapping relationship between the 11 groups of number sequence final consonants and 11 column character keys of a keyboard, so that the character final consonants on each column character key are merged into the corresponding number sequence final consonants through column direction compression; in response to the input operation of a user, converting the character final consonants into corresponding number sequence final consonant codes according to the mapping relationship, and generating a candidate word list. The application realizes single key triggering of final consonants by systematically merging final consonants into 11 groups and establishing a determined mapping with the column direction of the keyboard, solves the problems of multiple key strokes and inconsistent coding logic in the prior art, significantly improves the input efficiency of the Chinese pinyin, supports multiple input modes to share a unified coding system, and reduces the learning cost of users.
Owner:梁晨

Method, apparatus, and medium for recognizing a voice tone

ActiveCN120319275BSpeech recognitionAcousticsVowel
Provided are a method, apparatus, and medium for recognizing tones of speech. The method comprises: detecting the duration position of a vowel of a vowel in a speech to be recognized; detecting a tone core portion of the speech to be recognized based on the duration position of the vowel of the vowel; and identifying the tone category of the speech to be recognized based on the tone core portion. Thus, according to at least one embodiment of the present disclosure, the tone core portion can be more accurately detected within the duration position of the vowel of the vowel, thereby more accurately identifying the tone category.
Owner:NEW ORIENTAL EDUCATION & TECH GRP CO LTD

Mouth animation sequence generation method and apparatus

The present disclosure relates to a mouth shape animation sequence generation method and device. The method comprises: processing input information including input audio and / or input text to obtain a phoneme sequence corresponding to the input information; determining a viseme sequence corresponding to each phoneme in the phoneme sequence based on the phoneme sequence and a matched target mapping table; adjusting the weight of each viseme in the viseme sequence according to the viseme sequence and the corresponding target information to form an adjusted viseme sequence; determining a mouth shape animation corresponding to each viseme in the adjusted viseme sequence based on the target mapping table; and generating a mouth shape animation sequence according to the mouth shape animation. The target mapping table matched with the input information is used to realize the generation of the phoneme sequence based on the input information, the viseme sequence, and the mouth shape animation sequence. The input requirement is low, the mouth shape animation sequence is suitable for different accents and languages, the generation speed is fast, the time is short, the synchronization matching degree with the audio is good, the accuracy is high, and manual debugging is not required, thereby reducing the labor cost.
Owner:DIVINE VISION (SHENZHEN) CULTURE TECH CO LTD

Speech recognition method, system and terminal

The present disclosure provides a multi-functional speech recognition method, system and terminal. The present disclosure constructs an end-to-end speech recognition system with pinyin initial and final consonants as modeling units, and adds a modeling unit probability output module, effectively improving non-close sound character replacement errors, and relative to existing end-to-end speech recognition systems, adding various functions such as end-to-end speech recognition system acoustic recognition performance evaluation and pronunciation standard degree evaluation. The method comprises: receiving a speech to be recognized; performing acoustic feature extraction and encoding on the speech to be recognized; decoding the encoded acoustic features using a Chinese character decoder, wherein the Chinese character decoder uses pinyin initial and final consonants as modeling units, and maps the encoded acoustic feature sequence to a Chinese character sequence through initial and final consonants; and outputting a speech recognition result.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A semantic analysis method for speech in a crowded environment

The application discloses a kind of personnel dense environment under the semantic analysis method of voice, comprising: acquisition obtains the voice signal data of target user, lip movement video stream data;Extract lip movement feature sequence, map lip movement feature sequence into predicted voice feature vector;Input voiceprint separation model, separate the voice segment of target user from mixed voice signal, generate pure voice data characteristics;Pure voice data characteristics are carried out time domain segmentation, obtain the phonetic time domain waveform of single word;Single word phonetic time domain waveform is matched with initial, final time domain waveform library, obtains the corresponding pinyin expression of each word;Tone combination correlation analysis is carried out to the pinyin expression of continuous single word, obtains the meaning of the voice segment of target user.The application has the advantages that: by combining lip movement video stream and voice signal data, using deep learning and voiceprint separation technology, the voice of target user is effectively extracted, and the voice recognition accuracy in noisy environment is significantly improved.
Owner:SHENYANG LINKTECH INFORMATION TECH CO LTD

Method and system for converting natural intonation into melody, terminal and medium

The invention provides a method and a system for converting natural intonation into melody, a terminal and a medium, which are characterized in that fundamental frequency and tone length information of a voice sample are extracted, initial and final consonants are distinguished, fundamental frequency parameters of each final consonant are calculated, and the tone tuning range of each final consonant in a statement is determined. And matching music symbols by using the tuning ranges to generate a preliminary melody, and finally aligning the melody with the syllables of the voice sample to generate an audio file with a song melody. On the basis of the phonetic theory, the fundamental frequency, the tone length and the pause interval information of the voice sample are accurately extracted, and it is ensured that the generated melody can accurately reflect the characteristics and rhythm of natural intonations. The automatic method not only simplifies the operation process and improves the conversion efficiency, but also is suitable for a plurality of fields such as linguistics research, speech synthesis, language teaching, language training and pronunciation improvement.
Owner:GUANGMING HOSPITAL OF TRADITIONAL CHINESE MEDICINE PUDONG NEW AREA SHANGHAI

Phonetic and morphological code Chinese character input method of binary syllabification and initial consonants of head and tail components

The invention discloses a phonetic and morphological code Chinese character input method based on binary syllabification and initial consonants of head and tail components, belongs to the technical field of Chinese character input methods, and aims at solving the problems that an existing phonetic and morphological code input method is unreasonable in key position distribution, high in repeated code rate and too long in code length. The final key positions are mnemonic with the aid of final poems, are classified and distributed at 30 key positions according to final heads, and are combined with compound finals and the like; initial keys correspond to a standard English keyboard, V = zh, I = ch, U = sh, O is a zero initial, y and w are regarded as independent initial, and special keys reuse vowel keys. The initial consonants of radicals or strokes are taken by the head and tail parts, 201 radicals are adopted as the radicals, strokes are taken by non-radicals, the initial consonants of the radicals are taken by the head parts when a single character is the radicals, keys of part of the parts are adjusted, and auxiliary memorization is carried out by marking additional keys through visual elements. Common character codes are'initial consonants + final consonants + initial consonants of a first part + initial consonants of a tail part ', and the codes of words adopt different rules. The method has the advantages that no repeated code is added; the learning cost is low; the memory burden is relieved; the input efficiency is improved; and the repeated code rate is reduced.
Owner:王越泽

Electrical equipment control method and device, electrical equipment, medium and product

The invention relates to the technical field of smart home, and discloses an electrical equipment control method and device, electrical equipment, a medium and a product. Static initial consonant features and final features are obtained by performing feature extraction on dialect voice information of a user, and dynamic transition features of user pronunciation from initial consonants to final vowels are obtained; and capturing the difference of different dialects in the aspects of initial and final connection and tone change. Therefore, the accuracy of dialect recognition is improved based on the initial consonant features, the vowel features and the transition features. And then, the dialect instruction set corresponding to the dialect type is utilized to convert the dialect instruction of the user into a standard instruction, so that the user can control the electrical equipment through the dialect voice, and the use experience of the user is improved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI

Sample set generation method, device and computer equipment for training speech recognition model

The present application relates to the technical field of speech recognition, and aims to solve the problem of lack of training samples of heavy accent speech recognition model. A method, device and computer equipment for generating a sample set for training a speech recognition model are provided, wherein the method comprises: decoding a target command word into a toneless original pinyin sequence; based on a heavy accent rule library (heavy accent refers to non-standard pronunciation of initial / final change) constructed according to common non-standard pronunciation rules, a heavy accent pinyin sequence is generated through coding conversion; a heavy accent audio is generated through a text-to-speech audio generation tool; an audio input is input into a preset recognition model to obtain a recognition result and convert it into a recognized pinyin sequence; a screening rule is constructed based on the original and / or heavy accent pinyin, and the recognized pinyin sequence is compared to screen the audio; and the audio meeting the requirements is collected to form a training sample set. Through rule-based generation and accurate screening, high-quality heavy accent samples can be efficiently obtained, and the heavy accent recognition performance of the model can be improved.
Owner:深圳市友杰智新科技有限公司

Chinese speech signal segmentation method, device and equipment and storage medium

This disclosure provides a method, apparatus, device, and storage medium for segmenting Chinese speech signals. The method includes: sampling a target audio signal containing speech corresponding to a target Chinese text to obtain signal amplitudes corresponding to multiple sampling points; performing speech endpoint detection on the target audio data based on the signal amplitudes corresponding to the multiple sampling points to obtain multiple speech segments in the target audio data; determining the vowel position sequence corresponding to the target speech segment based on the formant energy of the speech signal in the target speech segment; determining syllable segmentation points and initial / final segmentation points of the target speech segment based on the signal amplitudes of sampling points between two adjacent vowel positions in the vowel position sequence and the initials and finals of the corresponding Chinese text segment; and segmenting the target speech segment based on the syllable segmentation points and initial / final segmentation points to obtain segmented speech primitives. This disclosure improves the accuracy of speech signal segmentation.
Owner:BEIJING INFORMATION TECH COLLEGE

Gnn-lstm method for chinese lip speech classification based on node multi-association graph information fusion

The application belongs to the technical field of lip language mouth shape analysis, and discloses a GNN-LSTM Chinese lip language classification method based on node multi-association graph information fusion. Lip key points are represented by three structures of adjacency graph, symmetric graph and upper and lower lip relationship graph. High-dimensional space-time features are extracted under the synergistic effect of graph convolutional neural network and long short-term memory network, the space-time global correlation between lip key points is effectively captured, and the mouth shape class is divided based on initial and final vowels. The multi-level collaborative relationship between initial and final vowels is considered, the influence of mouth shape similarity and visual ambiguity on model performance is reduced, a mouth shape library is established by using the extracted high-dimensional space-time features, the lip shape and corresponding pinyin are more discriminatively mapped and induced, the influence of mouth shape similarity and visual ambiguity on model performance is reduced, the method can adapt to complex changes and many-to-one mapping phenomena in actual pronunciation processes, and effectively enhances the accuracy and robustness of subsequent mouth shape classification.
Owner:XIANGJIANG LAB

System and method of modulating animation curves

A system and method of modulating animation curves based on audio input. The method including: identifying phonetic features for a plurality of visemes in the audio input; determining viseme animation curves based on parameters representing a spatial appearance of the plurality of visemes; modulating the viseme animation curves based on melodic accent, pitch sensitivity, or both, based on the phonetic features; and outputting the modulated animation curves.
Owner:JALI INC

A rhyme word recommendation method, device, equipment and storage medium

ActiveCN113850080BSemantic analysisEngineeringViseme
The present disclosure provides a rhyming word recommendation method and device, equipment and storage medium, relates to the technical field of computers, and particularly relates to the technical field of intelligent recommendation. The specific implementation scheme is as follows: obtaining a to-be-matched rhyming word and a rhyming configuration parameter; determining a target foot of a sentence to which the to-be-matched rhyming word belongs according to the to-be-matched rhyming word; determining a target rhyming word class corresponding to the target foot in a rhyming word class included in a rhyming word library, the feet of the sentences to which the rhyming words in each rhyming word class belong are identical in tone; determining a rhyming word matched with the to-be-matched rhyming word from the target rhyming word class according to the rhyming configuration parameter; and recommending the determined rhyming word. The scheme provided in the present disclosure can realize rhyming word recommendation.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Speech speed adjusting method and system, computer device and storage medium

The application discloses a speech speed adjusting method and system, computer equipment and a storage medium, relates to the technical field of information processing, and comprises the following steps: collecting a plurality of to-be-synthesized texts during speech synthesis, converting the to-be-synthesized texts into corresponding initial-final sequences, and recording the artificial marking pronunciation duration of the initial-final sequences obtained by an artificial marking method; vectorizing the initial-final sequences into phoneme vectors, and after the phoneme vectors pass through n layers of FFT Block and one-dimensional convolution, a predicted duration of the initial-final sequences is predicted; a difference between the predicted duration and the artificial marking pronunciation duration is calculated by using an L1 loss function, so as to train an initial-final duration prediction model during speech synthesis; and a target to-be-synthesized text is input into the initial-final duration prediction model, and corresponding target speech after speech speed adjustment is output according to a prediction result. Through the method of the application, speech speed adjustment of different speech during speech synthesis can be realized, and speech speed diversity can be realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Classical poem recitation assisting method fusing initial and final analysis

The invention discloses a classical poem recitation assisting method fusing initial and final analysis, and relates to the technical field of digital education, and the method comprises the steps: achieving the intelligent judgment and prompt of recitation rhythm and escort state based on the structural collection and analysis of recitation audio signals; the method comprises the following steps: acquiring and dividing a recitation audio into word segments through a voice recognition technology, and generating corresponding audio feature data information and text feature data information in combination with a poem structure; then judging whether the recitation duration is consistent with the poem structure or not through sentence level comparison, and providing a duration adjustment suggestion when the recitation duration is abnormal; on the basis that the recitation rhythm is normal, the vowel acoustic vector of each character is further extracted, and multi-level escort auxiliary suggestions are generated through initial and final difference analysis of the last character of adjacent sentences; according to the invention, through a process closed loop combining automatic acquisition, intelligent analysis and real-time feedback, the escort sensitivity and rhythm control ability of a learner are significantly improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Sample set generation method and device for training speech recognition model, and computer equipment

The invention relates to the technical field of speech recognition, and aims to solve the problem that training samples of a heavy accent speech recognition model are deficient. The invention provides a sample set generation method and apparatus for training a speech recognition model, and a computer device. The method comprises the steps of decoding a target command word into a silent original pinyin sequence; the method comprises the following steps of: performing code conversion on a duplicate accent rule base (duplicate accent refers to non-standard pronunciation of initial consonant / vowel change) constructed based on a common non-standard pronunciation rule to generate a duplicate accent pinyin sequence; generating a re-accent audio through a text-to-speech audio generation tool; inputting the audio into a preset recognition model to obtain a recognition result and converting the recognition result into a recognition pinyin sequence; constructing a screening rule by taking the original and / or double accent pinyin as a reference, and comparing and identifying the pinyin sequence to screen the audio; collecting audios meeting requirements to form a training sample set; according to the invention, through regularization generation and accurate screening, a high-quality heavy accent sample is efficiently obtained, and the heavy accent recognition performance of the model can be improved.
Owner:深圳市友杰智新科技有限公司

Tibetan written language character-voice conversion method based on voice structure characteristics

The invention, which belongs to the technical field of data processing, discloses a Tibetan language written language character-voice conversion method based on voice structure characteristics, comprising the following steps: S1, obtaining a Tibetan language text; s2, performing text segmentation processing on the Tibetan text by taking syllables as units, and extracting font structure elements of each syllable based on the font structure of the Tibetan to obtain a syllable font structure element set; s3, performing matching processing on initial consonant items and final consonant items of each syllable based on the syllable font structure element set to obtain a voice structure feature vector of each syllable; and S4, obtaining an initial consonant conversion coding table and a vowel conversion coding table, and mapping the voice structure feature vector of each syllable into a corresponding syllable phoneme sequence, thereby obtaining the Tibetan written language audio. According to the method, the problems that the basic characteristics of the Tibetan speech cannot be accurately described in a Tibetan language Latin transcription mode in the existing Tibetan speech synthesis technology and the like, and the calculation amount is too large and the character and speech conversion speed is low due to multilayer combination of Tibetan language phoneme sequences are solved.
Owner:TIBET UNIV

Realistic Lip Synchronization for Artificial Intelligence-Powered Talking Avatars

PendingUS20260188297A1AnimationAutomatic speech
The system and method for generating realistic lip synchronization for an AI avatar. The lip synchronization process begins by receiving input data. The input data can either be text input or audio stream. If the input data is text input, a text-to-speech (TTS) module converts it into speech while generating word-level timestamps. If the input data is the audio stream, an automatic speech recognition (ASR) module transcribes the spoken content and provides word-level timestamps. The transcribed text is transformed into a sequence of phonemes using a grapheme-to-phoneme conversion system. The phonemes are mapped to visemes based on their corresponding mouth shapes using a predefined phoneme to viseme mapping table. Then the visemes are mapped to the blendshapes of the AI avatar using image similarity comparison. The selected blendshapes are animated to generate synchronized animation of the AI avatar, with transitions smoothed to ensure natural lip movements and facial expressions.
Owner:2HR LEARNING INC

Chinese speech signal segmentation method and device, equipment and storage medium

ActiveCN121306099ASpeech recognitionSpeech synthesisFormantVowel
The embodiment of the invention discloses a Chinese speech signal segmentation method and device, equipment and a storage medium, and the method comprises the steps: carrying out the sampling of a target audio signal containing the speech corresponding to a target Chinese text, and obtaining signal amplitudes corresponding to a plurality of sampling points; performing voice endpoint detection on the target audio data based on the signal amplitudes corresponding to the plurality of sampling points to obtain a plurality of voice segments in the target audio data; determining a vowel position sequence corresponding to the target voice segment based on the formant energy of the voice signal in the target voice segment; determining syllable segmentation points and initial and final segmentation points of the target speech segment based on signal amplitudes of sampling points between two adjacent vowel positions in the vowel position sequence and initial and final consonants of the Chinese text segment corresponding to the target speech segment; and based on the syllable segmentation points and the initial and final segmentation points of the target voice segment, segmenting the target voice segment to obtain segmented voice elements. According to the embodiment of the invention, the accuracy of voice signal segmentation can be improved.
Owner:BEIJING INFORMATION TECH COLLEGE

High-simulation humanoid robot lip shape synchronous anthropomorphic control method and high-simulation humanoid robot lip shape synchronous anthropomorphic control system

The invention discloses a lip shape synchronous anthropomorphic control method and system for a high-simulation anthropomorphic robot. The method comprises the following steps: constructing a lip-shaped visual position mapping database of initial consonant phonemes and vowel phonemes; converting the pronunciation text into a pinyin sequence and identifying a pause symbol in the pinyin sequence; analyzing the pinyin into an initial consonant and vowel sequence, and extracting a corresponding lip action according to the mapping database; performing block processing on the action sequence according to the voice rhythm, and distributing the execution time of each action in combination with the audio duration; and a continuous lip shape control instruction sequence is generated by adopting a high-frequency interpolation mode, and is synchronously output in real time according to a time axis and audio playing. Through phoneme-level lip action driving, time fine distribution and continuous interpolation control, high consistency of voice content and lip actions in space and time dimensions is achieved, the anthropomorphic degree and the interaction reality sense of the humanoid robot in the language output process can be remarkably improved, and the user experience is improved. The lip-shaped driving system is suitable for various types of lip-shaped driving systems of humanoid robots.
Owner:WU XI WU JIE TAN SUO KE JI YOU XIAN GONG SI

Method for converting low-resource language voice into international phonetic symbol based on fused phonetic knowledge

The invention relates to a low-resource language voice-to-international phonetic symbol conversion method based on fused phonetic knowledge. According to the method, the technical threshold of translating extremely-low-resource voices into international phonetic symbols is reduced through tag injection, continuous embedding and phonetic system soft constraint decoding in a large model framework. By means of a lightweight adapter, the corpus can be migrated in about thirty hours, and the high cost of retraining the whole network is avoided; by means of an initial consonant-vowel-tone sign integrated label system and fusion of continuous fundamental frequency characteristics, the recognition precision of the system on phonemes and tones is synchronously improved. An interpretable result which conforms to the specification and is attached with confidence and a fundamental frequency curve is directly output, and the problems of illegal syllables and tone class confusion are thoroughly solved through joint decoding of initial consonant-vowel-tone sign transition probability and a language model. And a technical normal form with replicability is provided for digital continuation, cross-language information retrieval and cultural inheritance of endangered languages, and the method has important academic value and social significance.
Owner:INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES