Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

165 results about "Syllable" patented technology

A syllable is a unit of organization for a sequence of speech sounds. It is typically made up of a syllable nucleus (most often a vowel) with optional initial and final margins (typically, consonants). Syllables are often considered the phonological "building blocks" of words. They can influence the rhythm of a language, its prosody, its poetic metre and its stress patterns. Speech can usually be divided up into a whole number of syllables: for example, the word ignite is composed of two syllables: ig and nite.

Multi-language cross-culture communication auxiliary method and system based on large model

The invention provides a multi-language cross-culture communication assisting method and system based on a large model. The method comprises the following steps: receiving a source language audio stream during a call, calling a multi-language sound frequency harmonic modulation feature library to extract fundamental frequency harmonic intensity distribution and tone turning features, and generating a cultural acoustic fingerprint vector; based on the vector, controlling a microphone array phase difference, directionally enhancing a fundamental frequency harmonic component of a speaker and suppressing noise, and outputting a high signal-to-noise ratio spectrogram; analyzing the pronunciation rhythm and tone turning characteristics of the spectrogram, capturing the pitch jump and duration of the syllable boundary, and generating an acoustic culture label; associating the spectrogram with a target semantic library, matching harmonic distribution and a cultural context rule based on a large model, and outputting a cultural interpretation prompt containing an ambiguity resolution suggestion; and generating a calibration result according to the acoustic tag and the semantic prompt, and overlapping the dynamic floating subtitles to the face area of the speaker in the video conference picture. According to the invention, cultural tone ambiguity in multi-language communication is eliminated.
Owner:LUSTER LIGHTWAVE CO LTD

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Speech cloning system and method fusing rhythm characteristics

The invention discloses a voice cloning system and method fusing rhythm characteristics, belongs to the technical field of voice synthesis and natural language, and is applied to the aspect of fine-grained rhythm control in zero-sample voice synthesis. The implementation method comprises the following steps of: 1, extracting rhythm features and audio features of an audio file, and further respectively acquiring pause features, speed features and tone features in the rhythm features by sequentially adopting transcriptional text inverse coding, syllable-level speed registration quantization and pitch sequence feature splicing modes; 2, fusing the features of the audio files in a manner of discarding feature screening without guidance of a classifier; 3, generating a target Mel spectrogram based on conditional flow matching; generating a target audio file from a to-be-cloned audio file through the trained voice cloning model controlled by the fusion rhythm; compared with the prior art, fine-grained rhythm control of tone and rhythm feature decoupling is realized in zero-sample speech synthesis, so that intonation accuracy based on a context scene is improved.
Owner:BEIJING INST OF TECH

Natural statement decoding method and device based on high-density electrocorticogram

The invention discloses a natural statement decoding method and device based on high-density electrocorticogram. The method comprises the following steps: acquiring an electroencephalogram signal acquired based on the high-density electrocorticogram; different frequency band signals are extracted from the electroencephalogram signals; the signals of different frequency bands comprise high gamma frequency band signals and at least one frequency band signal with the frequency lower than that of the high gamma frequency band signals; acquiring a voice starting point corresponding to each frequency band signal, and determining a target voice starting point based on the voice starting point corresponding to each frequency band signal; after the target voice starting point, acquiring a syllable classification result and a tone decoding result corresponding to each frequency band signal, determining a target syllable classification result based on the syllable classification result corresponding to each frequency band signal, and determining a target tone decoding result based on the tone decoding result corresponding to each frequency band signal; and determining a target tone language corresponding to the electroencephalogram signal based on the target syllable classification result and the target tone decoding result. According to the scheme, the natural statement decoding accuracy can be improved.
Owner:SHANGHAI TECH UNIV +1

Anti-fraud outbound call identification method for cross-number-segment voiceprint tracking

The invention relates to an anti-fraud outbound recognition method for cross-number-segment voiceprint tracking, and the method comprises the steps: introducing a time domain voiceprint adversarial extraction mechanism, and adding an adversarial voice discriminator, an active inhibition device, a channel, and the interference of non-identity factors of language emotion in voiceprint modeling; constructing a speaker constant feature residual channel: taking stable syllable fragments including vowels and gutto vowels in call content as a reference extraction window, and filtering emotion or speech speed driving components; outputting a steady-state voiceprint contour flow as a unique basis for subsequent cross-number identity fusion; designing a dynamic number fusion graph based on the voiceprint steady-state flow; each time of voiceprint appearance is regarded as a single time point node, and all similar historical voiceprints form a serial number merging path; a number time sequence edge weight function is set; the occurrence time interval, the use frequency and the regional jump of the new number and the old number are synthesized, and whether the number belongs to the same user or not is evaluated; and establishing atlas connection between the voiceprint main body node and all number nodes thereof.
Owner:SHIJIAZHUANG LINGYUE TECHNOLOGY CO LTD

Terminal for executing korean reading program and operation method therefor

A terminal device for executing a Korean reading program, according to an embodiment of the present invention, comprises: a word search and identification unit that provides a word search and identification interface that supports a field-specific category and a translated word of a word of interest specified by a user from among English words belonging to the field-specific category and supports the translated word to be flickered and displayed, or a color of the translated word to vary within a rectangular border; a phoneme combination unit that interworks with the word search and identification unit, and provides a phoneme combination interface that supports information regarding a combination step (order) of phonemes for syllables of the translated word and pronunciation information and phonetic symbols regarding the phonemes and syllables, so as to be output; and a combination exercise unit that generates syllables by combining phoneme key information input via an interface provided by the phoneme combination unit, and combines the generated syllable information into a word.
Owner:KACEN CO LTD

Apparatus, systems, methods, and improvements formachine processing of language, for example such using sound / meaning associations

Systems and methods evaluate the meanings of whole words bound within the meanings of their parts by analyzing and identifying features and components of words (such as letters and syllables), using sound information, for example phoneme class structure, to find sound correlation and using synonymity to find meaning correlation between words (or word parts) and a data set of known / identifiable sound / meaning associations called morphemes to find morpheme‑word‑parts that optimally describe an input word. Various means are described of compiling morphemes into a data set. The systems and methods further include analyzing the semantic functions of word parts and their modifying interrelationships, and using identifications made by one or more analyses. The systems and methods can evaluate or improve understanding of word meanings. The systems and methods are useful in researching sound / meaning associations in language generally and for improving language processing generally by the incorporation of more semantic information.
Owner:JOHNSON MOLLY ANNE

Model training method, voice recognition method, device, equipment, medium and product

The invention provides a model training method, a voice recognition method, a voice recognition device, electronic equipment, a computer readable storage medium and a computer program product, and relates to the technical field of voice processing, the training method comprises the following steps: obtaining a target sequence, the target sequence being an audio feature sequence of sample voice data with noise; inputting the target sequence into a self-attention layer of a speech recognition model, and controlling the self-attention layer to perform attention processing on local sub-features in the target sequence and context features of the local sub-features to obtain sub-features after attention processing; obtaining a text recognition result of the sample voice data based on each sub-feature after attention processing; and training a speech recognition model based on the difference between the text recognition result and the labeled text carried by the sample speech data. According to the invention, the accuracy and response speed of the speech recognition model for recognizing the keyword syllables in the sound signals can be improved, and the man-machine interaction experience is improved.
Owner:MIDEA GRP (SHANGHAI) CO LTD +1

Tibetan word learning method based on image recognition and speech synthesis

The invention discloses a Tibetan word learning method based on image recognition and speech synthesis, and the method comprises the following steps: carrying out the standardization processing of an image containing Tibetan words, then recognizing Tibetan characters in the image, and generating a corresponding character sequence. And constructing data containing syllables and pronunciation information through phoneme segmentation, and generating input content for controlling speech synthesis. And generating a voice signal matched with the original image content by using the improved voice synthesis model in combination with the syllable labels, the rhythm factors and the voice features. The images, the recognized characters and the voice are integrally displayed in the learning terminal, linkage feedback from the characters to the voice is achieved, and the interaction effect and efficiency of Tibetan learning are improved. According to the invention, the end-to-end feedback of the learning process is realized in a mode of combining image recognition and speech synthesis.
Owner:TIBET MIRAN EDUCATION TECH CO LTD

Short video copywriting tone automatic adjusting method driven by hierarchical rhythm mapping

The invention discloses a hierarchical rhythm mapping-driven short video copywriting mood automatic adjustment method, and relates to the technical field of video processing, and the method comprises the steps: 1, receiving a text character string and a language type identifier, and building an occupation column for bearing a tone mark, an accent mark and a duration mark at each level; 2, dividing each sentence into phrase segments based on the hierarchical index table, freezing boundaries by taking the phrase segments as units, presetting sentence end termination styles according to punctuations, determining kernel phrases according to semantic anchor points, initializing trends of the kernel phrases, and performing time sequence elastic alignment and hierarchical backfilling to obtain a sentence end termination pattern; and finally outputting a triple sequence which covers all syllables and is composed of tone marks, accent marks and duration marks as a target rhythm control sequence. And step 3, performing audio generation based on the target rhythm control sequence to obtain new dubbing. According to the method, the tone accuracy and expressive force of short video dubbing are improved, and the time and cost of manual adjustment are remarkably reduced.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Sensitive word dynamic monitoring method and device in Tibetan complex environment speech recognition

The invention provides a sensitive word dynamic monitoring method and device in Tibetan complex environment speech recognition, and the method comprises the steps: carrying out the segment segmentation of a to-be-monitored Tibetan speech input stream in a complex environment, and obtaining the framed speech signals of the Tibetan speech in the complex environment; performing multi-dimensional perceptual fusion on the phoneme-level score, the semantic-level score, the scene-level score and the emotion-level score of each candidate sensitive word in the framing voice signals to obtain perceptual fusion features of the context corresponding to each candidate sensitive word; determining the decoding confidence coefficient of the Tibetan syllable sequence through the path stability of a decoding path in speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence; and determining the context sensitivity of each candidate sensitive word based on the decoding confidence and each perceptual fusion feature, and carrying out graded early warning on the Tibetan speech input stream in the complex environment through the context sensitivity. Based on the scheme, multi-dimensional fusion scoring of sensitive words in Tibetan complex environment speech recognition can be realized.
Owner:BEIJING WANGZHI TIANYUAN BIG DATA TECH CO LTD +1

Interaction method and system of intelligent glasses and translation machine

The embodiment of the invention belongs to the field of intelligent interaction, and relates to an interaction method of intelligent glasses and a translator, which comprises the following steps: the intelligent glasses synchronously acquire voice information of a user through a plurality of microphone matrixes; the voice information is transmitted to a translator in real time through a wireless communication protocol, a timestamp is added to each frame of voice information, and the translator detects transmission delay according to the timestamps and dynamically adjusts the transmission rate of the voice information; the translation machine divides the voice information into different subunits for processing according to syllables, intonations and grammar by adopting a preset acoustic model segmentation processing algorithm; the translation machine translates the voice information by using a vocabulary selection algorithm based on a language environment; and the translation machine feeds back a translation result to the intelligent glasses through an incremental feedback algorithm. The invention further provides an interaction system of the intelligent glasses and the translation machine. The objective of the invention is to realize low-delay and high-accuracy translation result feedback while ensuring high-precision speech recognition and rapid translation.
Owner:深圳目渡科技有限公司

Authentication system and authentication method

An authentication system includes an acquisition unit configured to acquire a voice signal of an utterance voice of a speaker; a detection unit configured to detect a first utterance section during which the speaker is uttering from the voice signal and a second utterance section during which the speaker is uttering from voice signals of a plurality of speakers registered in a database; a determination unit configured to collate a first voice signal of the first utterance section with a second voice signal of the second utterance section and determine an authentication condition for authentication using the first voice signal based on a length of the second voice signal of the second utterance section or the number of syllables included in the second utterance section; and an authentication unit configured to authenticate the speaker based on the authentication condition.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Language decoding method and device based on electroencephalogram signals and electronic equipment

The invention relates to a language decoding method and device based on electroencephalogram signals and electronic equipment, and the method comprises the steps: collecting the electroencephalogram signals of a subject, and carrying out the preprocessing of the electroencephalogram signals, and obtaining a neural feature time sequence; the neural feature time sequence is input into a parallel decoding architecture, and the parallel decoding architecture comprises at least two decoding branches and is used for decoding subunits of different orthogonal dimensions of a language from the neural feature time sequence; obtaining a subunit probability sequence output by each decoding branch through a parallel decoding architecture; performing time sequence probability accumulation and fusion on the sub-unit probability sequence of each decoding branch to obtain a stabilized sub-unit sequence; performing legality verification and combination on the stabilized subunit sequence according to a phonetic system rule to generate a legal syllable sequence; and disambiguating the syllable sequence based on the semantic context, and outputting a continuous natural language text. According to the method, effective modeling and generalization decoding are carried out on massive syllable categories under the condition of limited training data.
Owner:AFFILIATED HUSN HOSPITAL OF FUDAN UNIV +1

English vocabulary follow-up pronunciation correction method based on English teaching

The invention discloses an English vocabulary follow-up pronunciation correction method based on English teaching, and relates to the technical field of voice processing, and the method comprises the following steps: S1, constructing a standard vocabulary pronunciation unit; s2, carrying out phoneme difference positioning by using a standard vocabulary pronunciation unit; s3, performing syllable dynamic comparison by using the phoneme deviation distribution of the learner; s4, performing vocabulary level comparison by using the learner phoneme deviation distribution and the syllable deviation mapping; and S5, performing difference mapping output by using the vocabulary pronunciation difference set. By setting the difference comparison between the actual pronunciation of the learner and the standard pronunciation, an accurate alignment relationship can be established at the phoneme level, and the difference between the actual pronunciation and the standard pronunciation in the aspects of acoustic characteristics, pronunciation duration, phoneme boundaries and the like is quantized, so that compared with the prior art, only the overall similarity or fuzziness judgment can be given, and the accuracy of pronunciation is improved. And a difference comparison mechanism can provide a clearer error correction reference, so that the learner can clearly recognize the position and degree of the problem when receiving the result.
Owner:GUANGZHOU COLLEGE OF COMMERCE

Intelligent microphone starting method based on voice recognition

InactiveCN121053971ASpeech recognitionSyllableRAPID SPEECH
The invention discloses an intelligent microphone starting method based on voice recognition, and relates to the technical field of voice recognition. The method comprises the following steps: acquiring a user voice signal and surrounding environment audio data to obtain a mixed audio signal stream; identifying an abnormal waveform and determining a frequency feature vector; determining a waveform fluctuation amplitude and obtaining an enhanced waveform stable representation; recognizing syllable interval duration and analyzing interval shortening degree; analyzing the urgent speech speed characteristics to determine a time domain stretching compensation value, and adjusting the syllable interval duration to obtain a syllable sequence; identifying spectrum features of the environment background audio and determining candidate wake-up words; performing similarity matching with an emergency wake-up word template to determine an intention recognition result; and analyzing the confidence score to obtain an emergency response instruction execution signal. Through multi-level signal processing and intelligent feature extraction, high-precision voice intention recognition in a complex environment is realized, the robustness and reliability of an emergency response system are remarkably improved, and rapid help seeking of a user in a crisis scene is guaranteed.
Owner:GUANGZHOU AOYUAN ELECTRONICS CO LTD

Speech synthesis method and device

The invention relates to a speech synthesis method and device, and the method comprises the steps: obtaining a to-be-synthesized text; performing phoneme conversion on the to-be-synthesized text to obtain a first phoneme sequence, the first phoneme sequence comprising at least one phoneme and a tone of a syllable where the at least one phoneme is located; decoupling phonemes and tones in the first phoneme sequence, and extracting phoneme features of the to-be-synthesized text based on a decoupled second phoneme sequence; wherein the second phoneme sequence comprises the at least one phoneme; performing text coding on the to-be-synthesized text in a syllable dimension, and extracting semantic features of the to-be-synthesized text; and performing fusion processing on the semantic features and the phoneme features, and generating voice based on the fused features. According to the invention, the synthesized voice is accurate in pronunciation and natural in rhythm, and it is ensured that the synthesized voice achieves expected effects in the aspects of naturalness, rhythm and semantic consistency.
Owner:YOUKU CULTURE TECH (BEIJING) CO LTD

Storage format for chinese language and related processing method and apparatus

A storage format of Chinese language (“Readable Hanyu Expression” or “RHE”) and the related processing methods and systems. Unlike current Chinese processing methods which directly code Chinese characters into fonts for display, RHE takes an indirect approach by storing Chinese language in the RHE storage format that can be mapped to several display forms including simplified and traditional Chinese characters, Hanyu Pinyin, etc. In the RHE storage format, each Chinese word is stored as an RHE storage element having the format (Syllable+Tone)n+Mark, where n is the number of syllables (Chinese characters) in the word, Syllable represents the pronunciation (without the tone) of the character, Tone represents the tone of the pronunciation, and Mark is a value that differentiates different words having the same pronunciations and tones. Various mapping tables are used to map RHE storage elements to standard Chinese character codes (such as Unicode) and Pinyin expressions.
Owner:YANG MINGWEI

Chinese pinyin sliding input system and method based on multidirectional sliding gestures

The invention discloses a multidirectional sliding Chinese character pinyin input method and a matrix type keyboard layout system, and belongs to the technical field of human-computer interaction of touch screen equipment (IPC classification number: G06F3 / 0488, G06F3 / 0323G06F 40 / 274). Aiming at the problems of complex vowel operation redundancy, initial and final combination conflict and the like in the existing sliding input method, the invention provides the following schemes: 1, the multidirectional sliding input method comprises the following steps of: starting sliding from any initial consonant or vowel (containing a zero initial consonant 0 / er key) on the basis of a matrix keyboard layout; different vowels or vowel combinations are mapped every time sliding is performed by one key in different directions or continuous sliding is performed. 2, multi-direction sliding input logic: defining an 8-direction sliding mapping er to be rightward-or-generated from e in a layout E; the method has the technical effects that the input efficiency is improved by more than 60% (high-frequency syllable steps are reduced by 70%), the false touch rate is lower than 5%, the national standard pinyin library is covered by 100%, and the method is obviously superior to a traditional touch screen input scheme.
Owner:江远

Pinyin teaching equipment

The utility model relates to the technical field of teaching equipment, in particular to Chinese pinyin teaching equipment which comprises a showing stand, the middle of the showing stand is in a grid shape and forms a plurality of storage spaces, syllable blocks are clamped in the storage spaces, and a sliding frame capable of sliding up and down is arranged in front of the showing stand. The display rack has the advantages that the middle of the display rack is in the grid shape to form a plurality of storage spaces, the syllable blocks are placed in the storage spaces, the arch-shaped elastic plates are fixedly connected to the two sides of the outer surfaces of the syllable blocks, the syllable blocks can be fixedly installed in the storage spaces through bending of the arch-shaped elastic plates, and installation and disassembly are convenient. The front of the showing stand is connected with the sliding frame in a sliding mode, the sliding frame is a rectangular frame, the syllable blocks can also be clamped in the sliding frame, a teacher can conveniently combine and identify different syllables, the sliding frame can shield other syllables, and interference to the identification process of children is avoided.
Owner:杨小均

Storage format for Chinese language and related processing method and apparatus

A storage format of Chinese language (“Readable Hanyu Expression” or “RHE”) and the related processing methods and systems. Unlike current Chinese processing methods which directly code Chinese characters into fonts for display, RHE takes an indirect approach by storing Chinese language in the RHE storage format that can be mapped to several display forms including simplified and traditional Chinese characters, Hanyu Pinyin, etc. In the RHE storage format, each Chinese word is stored as an RHE storage element having the format (Syllable+Tone)n+Mark, where n is the number of syllables (Chinese characters) in the word, Syllable represents the pronunciation (without the tone) of the character, Tone represents the tone of the pronunciation, and Mark is a value that differentiates different words having the same pronunciations and tones. Various mapping tables are used to map RHE storage elements to standard Chinese character codes (such as Unicode) and Pinyin expressions.
Owner:YANG MINGWEI

Method, device and storage medium for screening least pronounced text

The present disclosure provides a screening method, device and equipment of minimum pronunciation text and storage medium. Firstly, the to-be-selected corpus and the predefined syllable bag are acquired. Then, all similar syllables corresponding to each to-be-selected text are acquired respectively. Then, each to-be-selected text is traversed. The pronunciation text with the maximum number of similar syllables contained in the predefined syllable bag is selected from each to-be-selected text. When one pronunciation text is selected each time, the pronunciation syllable corresponding to each similar syllable of the pronunciation text is removed from the predefined syllable bag until the number of pronunciation syllables remaining in the predefined syllable bag is less than the number threshold. Finally, each selected pronunciation text is used as the training data for the current TTS training. According to the present disclosure, the to-be-selected texts are screened in descending order according to the coverage of all similar syllables of each to-be-selected text in the predefined syllable bag, so that the pronunciation text with the highest pronunciation coverage and the minimum number can be quickly screened.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Method for learning to read using specialized text

A method of providing an instructional scaffold to person learning to read English language comprises displaying printed matter to the learner in the form of lists. Preferably text stories. The normal spacing of the letters in words is kept intact; in a first part of the method the rime portions of monosyllable words are bolded and made larger than the onset portions; and in a second part of the method, which also may be used independently, the sequential syllables of multisyllable words are emphasized by alternating plain font with bolded font and the long vowels are identified, for example by underscore, again with the words of any text being kept intact.
Owner:NOAH TEXT LLC

A method and device for screening phonemes, an electronic device and a readable storage medium

The application provides a phoneme screening method and device, electronic equipment and readable storage medium. A to-be-detected speech stream is acquired, and the phoneme category of each phoneme contained in each syllable in the to-be-detected speech stream is determined according to phoneme classification information corresponding to a target language to which the to-be-detected speech stream belongs. Acoustic information of each phoneme is extracted, and the acoustic information of each phoneme is subjected to dimension reduction processing and clustering processing to determine a plurality of clustering clusters corresponding to each phoneme category. For each phoneme category, a phoneme stable interval corresponding to the phoneme category is determined based on the plurality of clustering clusters corresponding to the phoneme category. The to-be-detected speech stream is screened based on the phoneme stable interval corresponding to each phoneme category, and a plurality of target phonemes located in the phoneme stable interval in the to-be-detected speech stream are determined. The to-be-detected speech stream is subjected to voiceprint comparison and analysis based on the plurality of determined target phonemes. In this way, the accuracy of voiceprint recognition and matching can be improved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Input method word frequency adjustment method and device

The present application discloses an input method word frequency adjustment method and device, which are used to solve the technical problem of poor word frequency adjustment effect of input method phrases. An input method word frequency adjustment method includes the following steps: obtaining corpus data; segmenting the corpus data through a word segmentation model to generate a number of word segmentation units; annotating the word segmentation units through a phonetic recognition model to generate word segmentation unit syllables; saving word segmentation units with the same syllables to the same syllable vocabulary; counting the occurrence probability of the first word segmentation unit in the same syllable vocabulary; comparing the occurrence probability of the first word segmentation unit with a preset threshold to obtain a comparison result; adjusting the word frequency of the first word segmentation unit according to the comparison result; arranging the word segmentation unit order of the syllable vocabulary where the first word segmentation unit is located in a preset order according to the adjusted word frequency of the first word segmentation unit, and updating the syllable vocabulary. By dynamically adjusting the word frequency of phrases in the same syllable vocabulary, the accuracy of input is improved.
Owner:BEIJING THUNISOFT INFORMATION TECH

Chinese pronunciation defect recognition method and system based on speech recognition

The invention provides a Chinese pronunciation defect recognition method and system based on speech recognition, and the method comprises the steps: obtaining a Chinese speech signal, generating an initial time-frequency spectrum through short-time Fourier transform, extracting a frequency feature sequence from the initial time-frequency spectrum, carrying out the segmentation processing of the speech signal through a time-domain signal segmentation technology, and carrying out the recognition of a Chinese pronunciation defect. Preliminary segmentation fragments are obtained; for the abnormal tone distribution diagram and the abnormal airflow distribution diagram, fusing tone duration and breathing rhythm mode features to generate a comprehensive feature matrix, performing boundary optimization on an abnormal interval through tone smoothing processing, and determining feature distribution after smoothing; and positioning a specific pronunciation defect syllable through the defect candidate interval in combination with conjoint analysis of a tone change mode and an airflow amplitude threshold value, tracing correlation characteristics of a frequency abrupt change point location and signal energy distribution, and judging a defect cause combination to obtain a final defect positioning result.
Owner:PUYANG VOCATIONAL & TECHN COLLEGE

Chinese character phrase and sentence decoding method based on stereotactic EEG signals

This invention discloses a method for decoding Chinese character phrases and sentences based on stereotactic electroencephalogram (sEEG) signals. This method uses stereotactic electroencephalogram (sEEG) signals as system input to decode Chinese phonemes, achieving greater generalization and adaptability than syllables. After identifying the initials, finals, and tones that make up Chinese Pinyin (Pinyin has 23 initials and no initials, 24 finals, and five tones), this method performs a two-way correction process using a language model modeled by Bayesian probabilistic modeling and a large language model. The decoded and corrected phrases or sentences are then displayed in real time on a user interface.
Owner:WESTLAKE UNIV

A method for embedding a voice watermark and a related device

PendingCN122511266ASyllableSpeech sound
This application discloses a method and related apparatus for embedding speech watermarks, which can be used in the field of speech signal processing. The method involves: first, acquiring a syllable sequence; then, predicting the duration probability distribution of each syllable in the syllable sequence using a pre-trained speech watermarking model; next, sampling the initial duration of each syllable from the duration probability distribution; then, performing parity editing on the initial duration of each syllable based on the target binary watermark sequence to obtain a duration sequence; then, generating speech features based on the duration sequence; and finally, synthesizing target speech with the target watermark based on the duration sequence and speech features. Thus, by performing parity editing on the initial duration of syllables to achieve watermark embedding, the speech watermark is elevated from the signal level to the information level, effectively resisting generative attacks such as neural codecs and neural vocoders, and significantly improving the robustness of speech watermarks in synthesized speech.
Owner:UNIV OF SCI & TECH OF CHINA

Automatic debugging method and system, electronic equipment and storage medium

The invention provides an automatic debugging method and system, electronic equipment and a storage medium. The method is used for a sound sensing device and comprises the steps that a test audio is played to a user, and the test audio is any one of words, syllables, music, voice, vocal music and pure tone; receiving response information of the user; determining whether the response information is matched with the test audio; when the response information is not matched with the test audio, storing the test audio as a reference optimized audio; determining a target distinctive speech feature corresponding to the reference optimized audio; determining a target operation parameter corresponding to the target distinctive speech feature according to a corresponding relationship between the distinctive speech feature and the operation parameter of the sound sensing device; and adjusting the target operation parameters to optimize the sound sensing equipment. The method is favorable for reducing the debugging cost and simplifying the debugging process.
Owner:SHANGHAI LISTENT MEDICAL TECH CO LTD