Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Syllable" patented technology

A syllable is a unit of organization for a sequence of speech sounds. It is typically made up of a syllable nucleus (most often a vowel) with optional initial and final margins (typically, consonants). Syllables are often considered the phonological "building blocks" of words. They can influence the rhythm of a language, its prosody, its poetic metre and its stress patterns. Speech can usually be divided up into a whole number of syllables: for example, the word ignite is composed of two syllables: ig and nite.

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Systems and methods for intelligent playback

PendingUS20260155154A1Speech analysisUsing detectable carrier informationSpeech rateSyllable
Systems and methods for intelligent playback of media content may include an intelligent media playback system that, in response to determining the speech tempo in audio content by measuring syllable density of speech in the audio content, automatically adjusts a playback speed of the audio content as the audio content is being played based on the determined speech tempo. In some embodiments, the system may automatically and dynamically adjust the playback speed to result in a desired target speech tempo. In addition, the system may determine whether to automatically adjust playback speed of the audio content, as the media is being played, based on the detected speech tempo of the speech in the audio content and the determined type of content of media. Such automatic adjustments in playback speed result in more efficient playback of the audio content.
Owner:DISH NETWORK TECHNOLOGIES INDIA PTE LTD +1

A pronunciation annotation method, system and program product for english natural phonics

PendingCN122334178ANatural language processingSyllable
This application discloses a method, system, and computer program product for phonetic annotation of English phonics. The method includes: obtaining a target word and its corresponding syllable separator and International Phonetic Alphabet (IPA); determining whether the word is a polysyllabic word based on the syllable separator; rewriting the IPA for monosyllabic and polysyllabic words according to the correspondence between phonics phonetic annotation and IPA; rewriting monosyllabic words using the first font or first font size, and rewriting polysyllabic words by using the first font or first font size for syllables other than weak syllables, and rewriting weak syllables using the second font or second font size; adding spaces to the corresponding positions of the rewritten string according to the position of the syllable separator in the target word to generate English phonics phonetic annotation. This application uses IPA rewriting for English phonics phonetic annotation, applicable to 100% of words, with simple rules, simplified annotation, and easy to learn and master.
Owner:BEIJING LINGEMA TECHNOLOGY CO LTD

Speech recognition method and device based on hotword enhancement

PendingCN122454980ASyllableSpeech sound
The application provides a hotword enhancement-based speech recognition method and device, and relates to the technical field of artificial intelligence. The method comprises the following steps: obtaining first pinyin sequences of each hotword to obtain first pinyin sequences, and obtaining second pinyin sequences converted from speech data of a user; adding confused pinyin syllables to pinyin syllables in each first pinyin sequence and / or second pinyin sequence; matching the current first pinyin sequence with the current second pinyin sequence to obtain matched hotwords; and obtaining a speech recognition result by using a multimodal large language model based on the matched hotwords. By adding confused pinyin syllables to pinyin syllables, the obtained pinyin sequences cover pinyin syllables caused by acoustic deviations such as accents, unclear pronunciation, and non-standard pronunciation, greatly improving the recall rate of hotwords in complex real scenarios, improving the accuracy of hotword matching, and thus increasing the accuracy of the speech recognition result.
Owner:TSINGHUA UNIVERSITY +1

Gaming machine

PendingJP2026103242ARoulette gamesSyllableFlip-flop
The aim is to improve the gameplay experience by adding variety to the progression of the game. [Solution] The triggers for the performance control means to cause the sound output unit to output sound include a first operation trigger and a second operation trigger, and the sounds that the performance control means causes the sound output unit to output include a first voice and a second voice, the performance control means causes the first voice to output based on the first operation trigger and the second voice to output based on the second operation trigger, the length of the last syllable group of the second voice is longer than the length of the last syllable group of the first voice, or the volume of the last syllable group of the second voice is greater than the volume of the last syllable group of the first voice.
Owner:OLYMPIA KK

Text language automatic detection method and system based on syllable and affix features

PendingCN122366422ASyllableFeature vector
The application discloses a text language automatic detection method and system based on syllable and affix features, relates to the field of natural language processing, and comprises the following steps: text preprocessing and basic language unit extraction are performed on a text to be detected to obtain a text syllable sequence and an affix list; based on the affix list, a multidimensional pragmatic scene distribution vector in the context of the text syllable sequence is extracted, and core connectivity in a pre-constructed morphological paradigm knowledge graph is acquired; the function load value of the affix is matched from an affix function load database, the affix function load value is normalized, and a weighted affix feature vector is generated; syntax dependency relationship analysis and preliminary structure construction are performed on the text to be detected, and a dependency syntax tree of the text to be detected is generated. The application improves the detection accuracy and result interpretability of highly similar languages and mixed code texts by explicitly modeling and quantitatively evaluating deep morphological-syntactic rule features of languages.
Owner:CHANGJI UNIV

An environment perception intelligent assistance method and device based on speech recognition and a medium

PendingCN122369439ASyllableEnvironmental perception
This invention discloses an environmental perception intelligent assistance method, device, and medium based on speech recognition, relating to the field of intelligent speech interaction technology. The method includes: acquiring continuous speech signals and obtaining speech segments through framing, spectral denoising, and silence detection; dividing syllable boundaries; extracting syllable sequences based on the syllable boundaries and converting them into pinyin; generating a candidate path set by combining a vocabulary dictionary; reading the candidate path set; matching the word texts in the candidate paths with the syllable sequences to obtain candidate word segments; grouping and filtering the candidate word segments to generate stable segments; sorting the stable segments in chronological order to obtain a stable segment set; determining a state word group splicing threshold based on the stable segment set; splicing and summarizing the stable segments to obtain a state word group set; determining a state sequence based on the continuity relationship between state word groups in the current speech segment and the previous speech segment; and sorting the state sequence in chronological order to obtain a state sequence set. This invention improves the accuracy of semantic information extraction.
Owner:BOYIN HEARING EQUIP (SUZHOU) CO LTD

A device wake-up method and apparatus, a computer device, and a storage medium

ActiveCN115527532Bavoid false wakeupimprove accuracyBootstrappingSpeech recognitionSyllableEngineering
This application provides a device wake-up method, apparatus, computer device, and storage medium, belonging to the field of speech recognition technology. The method includes: classifying multiple speech frames in an acquired speech signal to obtain multiple classification information, the classification information indicating the probability that each syllable, character, or word from a target phrase is included in the speech frame; determining the end point of the target phrase based on the multiple classification information, the end point indicating the time when the playback of the target phrase in the speech signal ends; and waking up the target device based on the end point of the target phrase. This technical solution can determine the time when the playback of the target phrase ends in the speech signal and finally wake up the target device at the time when the playback of the target phrase ends, ensuring that the target device is only woken up when the target phrase is completely detected, avoiding false wake-ups and improving the accuracy of wake-up.
Owner:SOUNDAI TECH CO LTD

Vietnamese spelling correction method fusing graph convolution and large model

PendingCN122311192ASyllableConvolution
This invention relates to a Vietnamese spelling error correction method integrating graph convolution and a large model, belonging to the field of natural language processing technology. The invention includes the following steps: constructing a Vietnamese syllable confusion set and establishing a syllable confusion graph based on it; encoding the confusion graph using a graph convolutional network to obtain a graph-enhanced syllable representation; segmenting the input text into words and locating syllable segments that match nodes in the confusion graph; feeding the input into a large model to obtain an output prediction distribution; calculating the similarity weights of candidate confused syllables based on the graph-enhanced syllable representation, and performing graph constraint smoothing on the logits output by the large model, thereby improving the error correction capability for fine-grained confusion errors related to tones and consonants; and outputting the corrected Vietnamese text. This invention improves the accuracy and robustness of Vietnamese spelling error correction by explicitly injecting a priori syllable confusion structure.
Owner:KUNMING UNIV OF SCI & TECH

Speech discrimination test system and device

PendingUS20260151054A1Medical communicationAudiometeringSyllableSpoken language
The invention relates to a test to measure the ability of a subject to discriminate between speech sounds. The speech sounds may be selected from the world's most widely spoken languages to enable the test to be carried out irrespective of the language spoken by the subject. Speech sounds may be presented in sequences, such as triplets. The subject would then be required to detect which speech sound in each sequence is different from the others. The test may be provided on a computer or similar device, in an embodiment a tablet computer, enabling the subject to conduct a self-assessment. The test may be performed on hearing aid users to determine the likelihood of a hearing aid user obtaining better test results after receiving one or two cochlear implants.
Owner:COCHLEAR LIMITED

Language decoding method and device based on electroencephalogram signals and electronic equipment

This invention relates to a language decoding method, apparatus, and electronic device based on electroencephalogram (EEG) signals. The method includes: acquiring EEG signals from a subject and preprocessing them to obtain a neural feature time series; inputting the neural feature time series into a parallel decoding architecture, which includes at least two decoding branches for decoding sub-units of language with different orthogonal dimensions from the neural feature time series; obtaining the sub-unit probability sequence output by each decoding branch through the parallel decoding architecture; performing temporal probability accumulation and fusion on the sub-unit probability sequences of each decoding branch to obtain a stable sub-unit sequence; performing legality verification and combination of the stable sub-unit sequences according to phonological rules to generate a legal syllable sequence; and performing disambiguation processing on the syllable sequence based on semantic context to output continuous natural language text. This invention achieves effective modeling and generalized decoding of massive syllable categories under limited training data conditions.
Owner:AFFILIATED HUSN HOSPITAL OF FUDAN UNIV +1

Audio data processing method and device for acoustic communication and storage medium

The application provides an audio data processing method and device for sound wave communication and a storage medium, which comprises the following steps: obtaining initial communication text to be sent, and performing encoding processing on the initial communication text to determine the encoding data corresponding to the initial communication text; the encoding data comprises at least one of check bit data and error correction code data corresponding to the initial communication text; performing frequency conversion on the encoding data to determine the frequency data corresponding to the encoding data; performing audio generation processing on the frequency data by using a pulse code modulation mode and a sine wave function to determine the audio data corresponding to the frequency data; the audio data comprises multiple syllables, and the volume change trend of each syllable from the beginning to the end is first large and then small; the audio data is sent to a receiving end device; and the audio data is used to enable the receiving end device to decode the audio data and then perform sound wave communication. The technical scheme can improve the success rate of sound wave communication.
Owner:ZHEJIANG UNIVIEW TECH CO LTD

A method for extracting and analyzing regional style characteristics of minnan tai folk songs

PendingCN122454957ASyllableTime domain
The present application relates to the technical field of audio classification, in particular to a method for extracting and analyzing regional style features of Min-Tai folk songs, which generates non-ideographic syllable time domain positioning interval, effective performance interval of accompanying character, regional exclusive accompanying character sound and rhyme feature anchoring sequence, and accompanying character-real character sound and rhyme linkage change parameter in sequence, and finally obtains a regional style feature identification matrix of Min-Tai folk songs. The present application collects original audio of Min-Tai folk song performance, completes accurate positioning of non-ideographic syllable time domain, screens effective performance interval of accompanying character through fundamental frequency continuity determination, anchors regional exclusive accompanying character sound and rhyme features through frame-by-frame spectrum comparison, extracts accompanying character-real character sound and rhyme linkage change parameter, finally binds features and parameter matching regional dimension, and generates a style feature identification matrix which can accurately distinguish regional attributes of Min-Tai folk songs.
Owner:FUJIAN NORMAL UNIV

System and method for producing syllable-unit three-dimensional finger gesture data for finger gesture recognition

PendingUS20260154988A1Character and pattern recognitionSyllableData set
Provided are a system and a method for producing syllable-unit 3D finger language posture data for finer language gesture recognition. The data set production method according to an embodiment may estimate 3D finger language postures from phoneme-unit 2D finger language videos, may pseudo-label phoneme ground truths of finger language gestures, and may produce, as a training data set, syllable-unit 3D finger language postures and syllable ground truths of the finger language gestures by combining the estimated 3D finger language postures and the phoneme ground truths. Accordingly, insufficient training data sets may be secured through augmentation without the time and cost burden.
Owner:KOREA ELECTRONICS TECH INST

A multi-scene intelligent assistant system and interaction method based on dialect voice wake-up

InactiveCN122116909ADigital data information retrievalSpeech recognitionSyllableSpeech segmentation
The application relates to the technical field of intelligent voice interaction, in particular to a multi-scene intelligent assistant system based on dialect voice wake-up and an interaction method, which comprises a voice segmentation module, a scene screening module, a terminal access module, an intra-domain wake-up module and an entry determination module. The application establishes a continuous processing relationship by surrounding the voice segment boundary, the scene label residence state, the terminal release state and the dialect wake-up word structure, narrows the scene range of the current voice segment participating in identification according to the scene label residence state, checks the tone fluctuation, the syllable sequence and the tailing convergence in the limited range, judges the interaction entry attribution in combination with the terminal access state and the activation confirmation information, keeps the scene constraint and the dialect wake-up discrimination synchronous, keeps the terminal access and the entry confirmation connected, controls the cross-scene repeated matching, the near-sound trigger confusion and the terminal access conflict, and makes the dialect voice wake-up and the subsequent interaction more coherent.
Owner:XIAMEN LUJIANG TECHNOLOGY CO LTD

A speech therapy treatment instrument abnormal voice detection method and system

PendingCN122290639ASyllableAbnormal voice
This invention relates to the field of pediatric therapeutic instrument technology, specifically to a method and system for detecting abnormal speech in a speech therapy instrument. The method involves acquiring voltage signals and filtering discrete speech frame sequences, extracting formant trajectories to analyze frequency change trends, comparing the articulation direction distribution of children with standard pronunciation, identifying directional deviation segments, and using formant neighborhood envelope peak comparison to pinpoint envelope offset segments. It also compares spectral peak and valley distribution characteristics, performs multidimensional temporal overlap comparison, and finally outputs the abnormal speech detection results. This invention utilizes amplitude cohesion to filter stable speech frame sequences, extracts structured formant trajectories, correlates syllable segment change directions to enhance dynamic trend discrimination, analyzes neighborhood envelope peak offsets to refine spectral characterization, cross-validates peak and valley distribution and temporal overlap segments, strengthens multidimensional consistency, achieves hierarchical identification and precise localization of pronunciation abnormalities, and improves the stability and distinguishability of detection results.
Owner:WOMEN & CHILDRENS MEDICAL CENTER AFFILIATED WITH GUANGZHOU MEDICAL UNIVERSITY

An AI-based automatic music composition and lyrics generation system and method

PendingCN122313929AEmotion perceptionSyllable
This invention discloses an artificial intelligence-based automatic music composition and lyrics generation system and method, specifically relating to the field of automatic music composition and lyrics generation technology. It models the attenuation characteristics of melody pitch distribution and dynamic changes in a target acoustic environment and extracts the main emotion-carrying frequency band. It calculates an emotion transmission attenuation coefficient based on the degree of spectral attenuation, couples the emotion transmission attenuation coefficient with the temporal distribution relationship of the lyrics semantics, constructs a temporal evolution sequence of emotion perception shift, analyzes the alignment relationship between lyric syllable accents and melody beats, calculates a rhythmic clarity degradation index, quantifies the degree of rhythmic ambiguity, and unifies the emotion transmission attenuation coefficient and rhythmic clarity degradation index into an acoustic adaptation degradation state vector. Based on this vector, it jointly adjusts the melody pitch distribution, note value configuration, and lyric syllable density, adaptively optimizing the emotion expression path and rhythmic structure for different playback terminals and acoustic environments.
Owner:SHENZHEN KUAIGE INTELLIGENT CO LTD

An artificial intelligence-based english oral english intelligent correction method and system

PendingCN122157694ASpeech analysisSyllableSpoken language
The application relates to an English oral English intelligent correction method based on artificial intelligence, characterized in that the pronunciation accuracy of a user to each test syllable is measured; the standard pronunciation of a training example sentence is played, and the user is reminded to read for the first time; the proficiency of each word of the training example sentence is predicted, the pronunciation weight, volume and pronunciation time length of the test syllable in the target word are corrected according to the predicted proficiency, and the modified pronunciation data of the word after individual processing is obtained; the user is played the correction emphasized pronunciation restored based on the word modified pronunciation data, and the user is reminded to read for the second time. The application can correct the pronunciation of the test syllable contained in the word with low proficiency, the test syllable is read with exaggeration, different words present different emphasized hearing, the user can be better reminded and corrected, the impression of the user on the reading correction of the unskilled word is deepened, and the effect of oral pronunciation correction is better.
Owner:WUHAN POLYTECHNIC

A system and method for converting a chinese dialect phonetic transcription into international phonetic alphabet transcription

PendingCN122454958ASyllableNatural language processing
The application discloses a Chinese dialect phonetic international phonetic alphabet transcription method and system. The application comprises: 1. A pre-trained self-supervised acoustic model is used as a feature extractor to convert the input dialect audio into implicit speech representation; 2. A configurable IPA syllable list of dialects is constructed: a configuration mechanism of the IPA syllable list required by the target dialect is provided according to the requirement; 3. The acoustic model is retrained through a small sample fine-tuning mechanism: the dialect corpus within a specified time is labeled, and the labeled target dialect corpus is used to continue training part of the parameters of the acoustic model to complete fine-tuning adaptation, so that the model can better capture the acoustic characteristics of the target dialect; 4. The frame-by-frame syllable output by the small sample fine-tuned acoustic model is mapped into the final IPA syllable sequence by using an IPA decoding module. The application can realize accurate transcription of dialect phonetics under a small amount of labeled corpus, and solve the technical problem of standardized phonetic transcription caused by the lack of large-scale training data of Chinese dialects.
Owner:HANGZHOU NORMAL UNIVERSITY