Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

116 results about "Syllable" patented technology

A syllable is a unit of organization for a sequence of speech sounds. It is typically made up of a syllable nucleus (most often a vowel) with optional initial and final margins (typically, consonants). Syllables are often considered the phonological "building blocks" of words. They can influence the rhythm of a language, its prosody, its poetic metre and its stress patterns. Speech can usually be divided up into a whole number of syllables: for example, the word ignite is composed of two syllables: ig and nite.

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Natural statement decoding method and device based on high-density electrocorticogram

The invention discloses a natural statement decoding method and device based on high-density electrocorticogram. The method comprises the following steps: acquiring an electroencephalogram signal acquired based on the high-density electrocorticogram; different frequency band signals are extracted from the electroencephalogram signals; the signals of different frequency bands comprise high gamma frequency band signals and at least one frequency band signal with the frequency lower than that of the high gamma frequency band signals; acquiring a voice starting point corresponding to each frequency band signal, and determining a target voice starting point based on the voice starting point corresponding to each frequency band signal; after the target voice starting point, acquiring a syllable classification result and a tone decoding result corresponding to each frequency band signal, determining a target syllable classification result based on the syllable classification result corresponding to each frequency band signal, and determining a target tone decoding result based on the tone decoding result corresponding to each frequency band signal; and determining a target tone language corresponding to the electroencephalogram signal based on the target syllable classification result and the target tone decoding result. According to the scheme, the natural statement decoding accuracy can be improved.
Owner:SHANGHAI TECH UNIV +1

Tibetan word learning method based on image recognition and speech synthesis

The invention discloses a Tibetan word learning method based on image recognition and speech synthesis, and the method comprises the following steps: carrying out the standardization processing of an image containing Tibetan words, then recognizing Tibetan characters in the image, and generating a corresponding character sequence. And constructing data containing syllables and pronunciation information through phoneme segmentation, and generating input content for controlling speech synthesis. And generating a voice signal matched with the original image content by using the improved voice synthesis model in combination with the syllable labels, the rhythm factors and the voice features. The images, the recognized characters and the voice are integrally displayed in the learning terminal, linkage feedback from the characters to the voice is achieved, and the interaction effect and efficiency of Tibetan learning are improved. According to the invention, the end-to-end feedback of the learning process is realized in a mode of combining image recognition and speech synthesis.
Owner:TIBET MIRAN EDUCATION TECH CO LTD

Short video copywriting tone automatic adjusting method driven by hierarchical rhythm mapping

The invention discloses a hierarchical rhythm mapping-driven short video copywriting mood automatic adjustment method, and relates to the technical field of video processing, and the method comprises the steps: 1, receiving a text character string and a language type identifier, and building an occupation column for bearing a tone mark, an accent mark and a duration mark at each level; 2, dividing each sentence into phrase segments based on the hierarchical index table, freezing boundaries by taking the phrase segments as units, presetting sentence end termination styles according to punctuations, determining kernel phrases according to semantic anchor points, initializing trends of the kernel phrases, and performing time sequence elastic alignment and hierarchical backfilling to obtain a sentence end termination pattern; and finally outputting a triple sequence which covers all syllables and is composed of tone marks, accent marks and duration marks as a target rhythm control sequence. And step 3, performing audio generation based on the target rhythm control sequence to obtain new dubbing. According to the method, the tone accuracy and expressive force of short video dubbing are improved, and the time and cost of manual adjustment are remarkably reduced.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Authentication system and authentication method

An authentication system includes an acquisition unit configured to acquire a voice signal of an utterance voice of a speaker; a detection unit configured to detect a first utterance section during which the speaker is uttering from the voice signal and a second utterance section during which the speaker is uttering from voice signals of a plurality of speakers registered in a database; a determination unit configured to collate a first voice signal of the first utterance section with a second voice signal of the second utterance section and determine an authentication condition for authentication using the first voice signal based on a length of the second voice signal of the second utterance section or the number of syllables included in the second utterance section; and an authentication unit configured to authenticate the speaker based on the authentication condition.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Language decoding method and device based on electroencephalogram signals and electronic equipment

The invention relates to a language decoding method and device based on electroencephalogram signals and electronic equipment, and the method comprises the steps: collecting the electroencephalogram signals of a subject, and carrying out the preprocessing of the electroencephalogram signals, and obtaining a neural feature time sequence; the neural feature time sequence is input into a parallel decoding architecture, and the parallel decoding architecture comprises at least two decoding branches and is used for decoding subunits of different orthogonal dimensions of a language from the neural feature time sequence; obtaining a subunit probability sequence output by each decoding branch through a parallel decoding architecture; performing time sequence probability accumulation and fusion on the sub-unit probability sequence of each decoding branch to obtain a stabilized sub-unit sequence; performing legality verification and combination on the stabilized subunit sequence according to a phonetic system rule to generate a legal syllable sequence; and disambiguating the syllable sequence based on the semantic context, and outputting a continuous natural language text. According to the method, effective modeling and generalization decoding are carried out on massive syllable categories under the condition of limited training data.
Owner:AFFILIATED HUSN HOSPITAL OF FUDAN UNIV +1

English vocabulary follow-up pronunciation correction method based on English teaching

The invention discloses an English vocabulary follow-up pronunciation correction method based on English teaching, and relates to the technical field of voice processing, and the method comprises the following steps: S1, constructing a standard vocabulary pronunciation unit; s2, carrying out phoneme difference positioning by using a standard vocabulary pronunciation unit; s3, performing syllable dynamic comparison by using the phoneme deviation distribution of the learner; s4, performing vocabulary level comparison by using the learner phoneme deviation distribution and the syllable deviation mapping; and S5, performing difference mapping output by using the vocabulary pronunciation difference set. By setting the difference comparison between the actual pronunciation of the learner and the standard pronunciation, an accurate alignment relationship can be established at the phoneme level, and the difference between the actual pronunciation and the standard pronunciation in the aspects of acoustic characteristics, pronunciation duration, phoneme boundaries and the like is quantized, so that compared with the prior art, only the overall similarity or fuzziness judgment can be given, and the accuracy of pronunciation is improved. And a difference comparison mechanism can provide a clearer error correction reference, so that the learner can clearly recognize the position and degree of the problem when receiving the result.
Owner:GUANGZHOU COLLEGE OF COMMERCE

Intelligent microphone starting method based on voice recognition

InactiveCN121053971ASpeech recognitionSyllableRAPID SPEECH
The invention discloses an intelligent microphone starting method based on voice recognition, and relates to the technical field of voice recognition. The method comprises the following steps: acquiring a user voice signal and surrounding environment audio data to obtain a mixed audio signal stream; identifying an abnormal waveform and determining a frequency feature vector; determining a waveform fluctuation amplitude and obtaining an enhanced waveform stable representation; recognizing syllable interval duration and analyzing interval shortening degree; analyzing the urgent speech speed characteristics to determine a time domain stretching compensation value, and adjusting the syllable interval duration to obtain a syllable sequence; identifying spectrum features of the environment background audio and determining candidate wake-up words; performing similarity matching with an emergency wake-up word template to determine an intention recognition result; and analyzing the confidence score to obtain an emergency response instruction execution signal. Through multi-level signal processing and intelligent feature extraction, high-precision voice intention recognition in a complex environment is realized, the robustness and reliability of an emergency response system are remarkably improved, and rapid help seeking of a user in a crisis scene is guaranteed.
Owner:GUANGZHOU AOYUAN ELECTRONICS CO LTD

Chinese pinyin sliding input system and method based on multidirectional sliding gestures

The invention discloses a multidirectional sliding Chinese character pinyin input method and a matrix type keyboard layout system, and belongs to the technical field of human-computer interaction of touch screen equipment (IPC classification number: G06F3 / 0488, G06F3 / 0323G06F 40 / 274). Aiming at the problems of complex vowel operation redundancy, initial and final combination conflict and the like in the existing sliding input method, the invention provides the following schemes: 1, the multidirectional sliding input method comprises the following steps of: starting sliding from any initial consonant or vowel (containing a zero initial consonant 0 / er key) on the basis of a matrix keyboard layout; different vowels or vowel combinations are mapped every time sliding is performed by one key in different directions or continuous sliding is performed. 2, multi-direction sliding input logic: defining an 8-direction sliding mapping er to be rightward-or-generated from e in a layout E; the method has the technical effects that the input efficiency is improved by more than 60% (high-frequency syllable steps are reduced by 70%), the false touch rate is lower than 5%, the national standard pinyin library is covered by 100%, and the method is obviously superior to a traditional touch screen input scheme.
Owner:江远

Pinyin teaching equipment

The utility model relates to the technical field of teaching equipment, in particular to Chinese pinyin teaching equipment which comprises a showing stand, the middle of the showing stand is in a grid shape and forms a plurality of storage spaces, syllable blocks are clamped in the storage spaces, and a sliding frame capable of sliding up and down is arranged in front of the showing stand. The display rack has the advantages that the middle of the display rack is in the grid shape to form a plurality of storage spaces, the syllable blocks are placed in the storage spaces, the arch-shaped elastic plates are fixedly connected to the two sides of the outer surfaces of the syllable blocks, the syllable blocks can be fixedly installed in the storage spaces through bending of the arch-shaped elastic plates, and installation and disassembly are convenient. The front of the showing stand is connected with the sliding frame in a sliding mode, the sliding frame is a rectangular frame, the syllable blocks can also be clamped in the sliding frame, a teacher can conveniently combine and identify different syllables, the sliding frame can shield other syllables, and interference to the identification process of children is avoided.
Owner:杨小均

Method, device and storage medium for screening least pronounced text

The present disclosure provides a screening method, device and equipment of minimum pronunciation text and storage medium. Firstly, the to-be-selected corpus and the predefined syllable bag are acquired. Then, all similar syllables corresponding to each to-be-selected text are acquired respectively. Then, each to-be-selected text is traversed. The pronunciation text with the maximum number of similar syllables contained in the predefined syllable bag is selected from each to-be-selected text. When one pronunciation text is selected each time, the pronunciation syllable corresponding to each similar syllable of the pronunciation text is removed from the predefined syllable bag until the number of pronunciation syllables remaining in the predefined syllable bag is less than the number threshold. Finally, each selected pronunciation text is used as the training data for the current TTS training. According to the present disclosure, the to-be-selected texts are screened in descending order according to the coverage of all similar syllables of each to-be-selected text in the predefined syllable bag, so that the pronunciation text with the highest pronunciation coverage and the minimum number can be quickly screened.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Method for learning to read using specialized text

A method of providing an instructional scaffold to person learning to read English language comprises displaying printed matter to the learner in the form of lists. Preferably text stories. The normal spacing of the letters in words is kept intact; in a first part of the method the rime portions of monosyllable words are bolded and made larger than the onset portions; and in a second part of the method, which also may be used independently, the sequential syllables of multisyllable words are emphasized by alternating plain font with bolded font and the long vowels are identified, for example by underscore, again with the words of any text being kept intact.
Owner:NOAH TEXT LLC

A method and device for screening phonemes, an electronic device and a readable storage medium

The application provides a phoneme screening method and device, electronic equipment and readable storage medium. A to-be-detected speech stream is acquired, and the phoneme category of each phoneme contained in each syllable in the to-be-detected speech stream is determined according to phoneme classification information corresponding to a target language to which the to-be-detected speech stream belongs. Acoustic information of each phoneme is extracted, and the acoustic information of each phoneme is subjected to dimension reduction processing and clustering processing to determine a plurality of clustering clusters corresponding to each phoneme category. For each phoneme category, a phoneme stable interval corresponding to the phoneme category is determined based on the plurality of clustering clusters corresponding to the phoneme category. The to-be-detected speech stream is screened based on the phoneme stable interval corresponding to each phoneme category, and a plurality of target phonemes located in the phoneme stable interval in the to-be-detected speech stream are determined. The to-be-detected speech stream is subjected to voiceprint comparison and analysis based on the plurality of determined target phonemes. In this way, the accuracy of voiceprint recognition and matching can be improved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Chinese pronunciation defect recognition method and system based on speech recognition

The invention provides a Chinese pronunciation defect recognition method and system based on speech recognition, and the method comprises the steps: obtaining a Chinese speech signal, generating an initial time-frequency spectrum through short-time Fourier transform, extracting a frequency feature sequence from the initial time-frequency spectrum, carrying out the segmentation processing of the speech signal through a time-domain signal segmentation technology, and carrying out the recognition of a Chinese pronunciation defect. Preliminary segmentation fragments are obtained; for the abnormal tone distribution diagram and the abnormal airflow distribution diagram, fusing tone duration and breathing rhythm mode features to generate a comprehensive feature matrix, performing boundary optimization on an abnormal interval through tone smoothing processing, and determining feature distribution after smoothing; and positioning a specific pronunciation defect syllable through the defect candidate interval in combination with conjoint analysis of a tone change mode and an airflow amplitude threshold value, tracing correlation characteristics of a frequency abrupt change point location and signal energy distribution, and judging a defect cause combination to obtain a final defect positioning result.
Owner:PUYANG VOCATIONAL & TECHN COLLEGE

A method for embedding a voice watermark and a related device

PendingCN122511266ASyllableSpeech sound
This application discloses a method and related apparatus for embedding speech watermarks, which can be used in the field of speech signal processing. The method involves: first, acquiring a syllable sequence; then, predicting the duration probability distribution of each syllable in the syllable sequence using a pre-trained speech watermarking model; next, sampling the initial duration of each syllable from the duration probability distribution; then, performing parity editing on the initial duration of each syllable based on the target binary watermark sequence to obtain a duration sequence; then, generating speech features based on the duration sequence; and finally, synthesizing target speech with the target watermark based on the duration sequence and speech features. Thus, by performing parity editing on the initial duration of syllables to achieve watermark embedding, the speech watermark is elevated from the signal level to the information level, effectively resisting generative attacks such as neural codecs and neural vocoders, and significantly improving the robustness of speech watermarks in synthesized speech.
Owner:UNIV OF SCI & TECH OF CHINA

Automatic debugging method and system, electronic equipment and storage medium

The invention provides an automatic debugging method and system, electronic equipment and a storage medium. The method is used for a sound sensing device and comprises the steps that a test audio is played to a user, and the test audio is any one of words, syllables, music, voice, vocal music and pure tone; receiving response information of the user; determining whether the response information is matched with the test audio; when the response information is not matched with the test audio, storing the test audio as a reference optimized audio; determining a target distinctive speech feature corresponding to the reference optimized audio; determining a target operation parameter corresponding to the target distinctive speech feature according to a corresponding relationship between the distinctive speech feature and the operation parameter of the sound sensing device; and adjusting the target operation parameters to optimize the sound sensing equipment. The method is favorable for reducing the debugging cost and simplifying the debugging process.
Owner:SHANGHAI LISTENT MEDICAL TECH CO LTD

Method and system for automatic detection of tone sandhi in continuous speech in Chinese

The application discloses a Chinese continuous speech tone change automatic detection method and system, and belongs to the technical field of speech signal processing.The method comprises the following steps: a speech forced alignment step is used to acquire syllable time boundaries; a fundamental frequency transition trajectory extraction step is used to extract a fundamental frequency feature vector in an adjacent syllable connection area; a tone change rule matching step is used to acquire an expected mode from a knowledge base containing necessary and variable rules; a tone change mode detection step is used to calculate a matching score and determine whether the tone change is correct, missing or excessive; and a feedback generation step is used to generate a pitch curve labeling graph and rule explanation.The application realizes accurate tone change detection by focusing on the fundamental frequency transition features of the connection area, provides reasonable evaluation by using a hierarchical rule knowledge base, and helps learners to improve pronunciation by visual feedback.
Owner:SICHUAN NORMAL UNIV

Human-computer interaction method and device, computer equipment and storage medium

The invention relates to the technical field of man-machine interaction, and discloses a man-machine interaction method and device, computer equipment and a storage medium, the man-machine interaction method comprises the following steps: obtaining audio data to be played, the audio data being generated based on a text; determining a target syllable corresponding to the audio data according to the spectrum feature corresponding to the audio data; based on a syllable mouth shape mapping table, determining a target mouth shape corresponding to the target syllable, the syllable mouth shape mapping table being used for storing a corresponding relationship between the syllable and the mouth shape; and playing the audio data, and synchronizing the mouth shape of the three-dimensional face displayed on the target screen as a target mouth shape. Through the technical scheme of the invention, the problem that the mouth shape of the three-dimensional face is not synchronous with the audio in the related technology is solved, the mouth shape of the three-dimensional face is synchronous with the audio in man-machine interaction, and the user experience is improved.
Owner:BEIJING SIWEI ZHILIAN TECH CO LTD

Syllable identification method and related equipment

The embodiment of the invention discloses a syllable recognition method, device and equipment and a computer readable storage medium, which are used for performing syllable recognition under the condition of improving the syllable recognition accuracy. The method provided by the embodiment of the invention comprises the following steps: acquiring to-be-processed electroencephalogram signal characterization, and performing layering processing of different time scales corresponding to pronunciation time points on the to-be-processed electroencephalogram signal characterization by a time scale layering module to obtain electroencephalogram signal characterization respectively corresponding to a plurality of time scale layers, a representation enhancement module corresponding to the time scale hierarchy performs representation enhancement processing of the time scale hierarchy on the electroencephalogram signal representation corresponding to the time scale hierarchy to obtain a target enhanced electroencephalogram signal representation of the time scale hierarchy; and the syllable recognition module performs characteristic analysis of integrating each time scale and each space scale on the target enhanced electroencephalogram signal representation of each time scale level to obtain a syllable recognition result corresponding to the electroencephalogram signal representation to be processed output by the multi-resolution syllable recognition model.
Owner:SHENZHEN READLINE BIOTECH CO LTD

Systems and methods for intelligent playback

Systems and methods for intelligent playback of media content may include an intelligent media playback system that, in response to determining the speech tempo in audio content by measuring syllable density of speech in the audio content, automatically adjusts a playback speed of the audio content as the audio content is being played based on the determined speech tempo. In some embodiments, the system may automatically and dynamically adjust the playback speed to result in a desired target speech tempo. In addition, the system may determine whether to automatically adjust playback speed of the audio content, as the media is being played, based on the detected speech tempo of the speech in the audio content and the determined type of content of media. Such automatic adjustments in playback speed result in more efficient playback of the audio content.
Owner:DISH NETWORK TECHNOLOGIES INDIA PTE LTD +1

Deep learning-based harmonic generation method and terminal

The invention discloses a deep learning-based harmonic generation method and terminal, and the method comprises the steps: obtaining a source language text of a to-be-converted language, and extracting a phoneme stream of the source language text; segmenting the phoneme stream according to a language rule of a target language to obtain a plurality of first syllables; obtaining a pronunciation library corresponding to the target language, matching a second syllable corresponding to each first syllable in the pronunciation library, and obtaining a font corresponding to the second syllable; establishing an association relationship between the source language text and all fonts, and training the initial model through the association relationship to obtain a target model; and when a to-be-converted text corresponding to the to-be-converted language is received, obtaining a homophonic font corresponding to the to-be-converted text through the target model. According to the method, the nonlinear mapping rule between the phoneme combination of the source language text and the target language pronunciation is obtained based on deep learning, so that when facing uncommon vocabularies and new words never appearing in a training set, the system still can generate accurate and reasonable homophonic phonetic notation through reasoning according to the learned general pronunciation logic.
Owner:FUJIAN STAR NET EVIDEO INFORMATION SYST CO LTD

Intelligent voice pronunciation trainer with display screen

1.The name of the design product: smart voice pronunciation practice device with display screen. 2.The use of the design product: for the correction of standard phonetic symbols and word pronunciation. Collect the audio of students' pronunciation, compare the standard pronunciation data through the built-in voice recognition algorithm, and intuitively present the tone curve, syllable duration and other parameter differences on the display screen. At the same time, it is accompanied by a text annotated pronunciation part diagram to help students quickly locate the pronunciation errors such as tongue position and lip shape. It can also preset the word and poem reading tasks synchronized with the teaching materials, automatically generate pronunciation accuracy reports, and facilitate teachers to carry out targeted layered teaching to consolidate students' language pronunciation foundation. 3.The design points of the design product: the combination of shape, pattern and color. 4.The picture or photo that best indicates the design points: perspective view. 5.The design protected contains color.
Owner:HUNAN MECHANICAL & ELECTRICAL POLYTECHNIC

Domain name detection method and device, communication equipment, readable storage medium and program product

The invention relates to a domain name detection method and device, communication equipment, a readable storage medium and a program product, relates to the technical field of network information security, and can improve the accuracy of a domain name detection result. The method comprises the following steps: acquiring a domain name main body character string corresponding to a domain name to be detected, and segmenting the domain name main body character string according to a pronunciation unit to obtain a plurality of syllable units; merging the plurality of syllable units according to a preset syllable unit combination configuration, obtaining a plurality of syllable blocks according to a merging result, and determining a plurality of syllable features according to the plurality of syllable blocks; determining the pronunciation naturalness of the domain name main body character string according to the plurality of syllable features; and determining an anomaly detection result of the domain name to be detected according to the pronunciation naturalness.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

A pronunciation annotation method, system and program product for english natural phonics

This application discloses a method, system, and computer program product for phonetic annotation of English phonics. The method includes: obtaining a target word and its corresponding syllable separator and International Phonetic Alphabet (IPA); determining whether the word is a polysyllabic word based on the syllable separator; rewriting the IPA for monosyllabic and polysyllabic words according to the correspondence between phonics phonetic annotation and IPA; rewriting monosyllabic words using the first font or first font size, and rewriting polysyllabic words by using the first font or first font size for syllables other than weak syllables, and rewriting weak syllables using the second font or second font size; adding spaces to the corresponding positions of the rewritten string according to the position of the syllable separator in the target word to generate English phonics phonetic annotation. This application uses IPA rewriting for English phonics phonetic annotation, applicable to 100% of words, with simple rules, simplified annotation, and easy to learn and master.
Owner:BEIJING LINGEMA TECHNOLOGY CO LTD

Multi-language audio inter-translation method and device

The embodiment of the invention discloses a multi-language audio inter-translation method and device, and the method comprises the following steps: carrying out the voice recognition of a source language audio, obtaining a source language text, and extracting a content feature, a tone feature and an emotional rhythm feature from the source language audio; preprocessing the source language text, and calculating the number of syllables of the source language text; performing controlled constraint on a target language text generation process based on the syllable number, and generating a target language text which is consistent with the source language text in semantics and is matched with the syllable number; and according to the target language text, the content features, the timbre features and the emotion rhythm features, generating a target language audio consistent with the source language audio in timbre, emotion and duration. The target language audio generated by the embodiment of the invention is accurate in semantics and matched in duration, and the unique voice characteristics of a source language audio speaker can be highly restored.
Owner:BEIJING AISHU WISDOM TECH CO LTD

Method and system for generating a unified script code (USC) representation for multilingual text processing

PCT designated stageWO2026074589A1Natural language translationSyllableDevanagari
The present invention provides a computer-implemented method and system for generating a Unified Script Code (USC) representation of multilingual text to enable consistent, script-neutral, and phonemically accurate encoding across diverse languages and writing systems The system converts Unicode-encoded text from one or more Indic scripts into a Devanagari-based intermediate form using predefined or bitwise mapping. It then normalizes the text by inserting inherent vowels, converting dependent vowel signs (matras) into independent vowels, and removing halant characters. The resulting USC representation explicitly encodes consonant-vowel sequences, reduces script specific variation, and preserves phonetic integrity. This approach improves tokenization efficiency, enhances performance in natural language processing and machine learning tasks, and supports reversible conversion to original scripts.
Owner:SINGH NAU NIHAL

A language decoding method and device, electronic equipment and storage medium

ActiveCN122220477BDecoding methodsSyllable
The application discloses a language decoding method and device, electronic equipment and storage medium. The method comprises the following steps: obtaining an electroencephalogram decoding result of a current time step, wherein the electroencephalogram decoding result comprises electroencephalogram decoding scores corresponding to each preset syllable; expanding a historical candidate sentence of a previous time step based on candidate words associated with the preset syllable to obtain a corresponding expanded sentence set; for each expanded sentence, determining an integrated language decoding score of the expanded sentence based on a historical integrated language decoding score of the previous time step, a first language probability score of the expanded sentence and the corresponding electroencephalogram decoding score; determining a candidate sentence sequence of the current time step based on the integrated language decoding scores of the expanded sentences; performing sentence end detection based on the candidate sentence sequence of the current time step, and determining language content to be output at the current time step according to a detection result of the sentence end detection. The application improves the accuracy, intelligibility and naturalness of interaction of the output language content.
Owner:SHANGHAI NEURO XESS TECH CO LTD

A method of obtaining training data for a speech synthesis model

The application discloses a method for obtaining training data of a speech synthesis model, comprising the following steps: screening part of corpus with high frequency syllables from a corpus of a target language; recording and sampling the screened corpus and marking the corpus to obtain an initial training set, and training the speech synthesis model through the initial training set; synthesizing speech of unmarked corpus in the corpus according to the speech synthesis model to obtain quality information of the speech synthesis; determining decision attributes of the unmarked corpus according to the quality information, taking the occurrence times of basic letters in the unmarked corpus as conditional attributes, constructing a decision table, and obtaining a dominant rough set. The application effectively solves the problems of few training data samples, high professional requirement, high cost and heavy workload in the recording and marking of training data samples in the target language synthesis task.
Owner:GARZE TIBETAN AUTONOMOUS PREFECTURE INST OF SCI & TECH INFORMATION