Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "Vocal pitch" patented technology

Multi-language cross-culture communication auxiliary method and system based on large model

The invention provides a multi-language cross-culture communication assisting method and system based on a large model. The method comprises the following steps: receiving a source language audio stream during a call, calling a multi-language sound frequency harmonic modulation feature library to extract fundamental frequency harmonic intensity distribution and tone turning features, and generating a cultural acoustic fingerprint vector; based on the vector, controlling a microphone array phase difference, directionally enhancing a fundamental frequency harmonic component of a speaker and suppressing noise, and outputting a high signal-to-noise ratio spectrogram; analyzing the pronunciation rhythm and tone turning characteristics of the spectrogram, capturing the pitch jump and duration of the syllable boundary, and generating an acoustic culture label; associating the spectrogram with a target semantic library, matching harmonic distribution and a cultural context rule based on a large model, and outputting a cultural interpretation prompt containing an ambiguity resolution suggestion; and generating a calibration result according to the acoustic tag and the semantic prompt, and overlapping the dynamic floating subtitles to the face area of the speaker in the video conference picture. According to the invention, cultural tone ambiguity in multi-language communication is eliminated.
Owner:LUSTER LIGHTWAVE CO LTD

Natural statement decoding method and device based on high-density electrocorticogram

The invention discloses a natural statement decoding method and device based on high-density electrocorticogram. The method comprises the following steps: acquiring an electroencephalogram signal acquired based on the high-density electrocorticogram; different frequency band signals are extracted from the electroencephalogram signals; the signals of different frequency bands comprise high gamma frequency band signals and at least one frequency band signal with the frequency lower than that of the high gamma frequency band signals; acquiring a voice starting point corresponding to each frequency band signal, and determining a target voice starting point based on the voice starting point corresponding to each frequency band signal; after the target voice starting point, acquiring a syllable classification result and a tone decoding result corresponding to each frequency band signal, determining a target syllable classification result based on the syllable classification result corresponding to each frequency band signal, and determining a target tone decoding result based on the tone decoding result corresponding to each frequency band signal; and determining a target tone language corresponding to the electroencephalogram signal based on the target syllable classification result and the target tone decoding result. According to the scheme, the natural statement decoding accuracy can be improved.
Owner:SHANGHAI TECH UNIV +1

Suzhou dialect medical voice electronic medical record conversion system and method

The invention provides a Suzhou dialect medical voice electronic medical record conversion system and method, and the system comprises an acoustic feature extraction module which is used for extracting the acoustic features of an input voice signal; the dialect tone recognition module is used for recognizing a multi-tone system of Suzhou dialects; the voiced sound processing module is used for detecting voiced sound initial consonants in Suzhou dialects and performing acoustic feature mapping; the medical term mapping module comprises a corresponding relation library of dialect medical vocabularies and standard medical terms and a context-based ambiguity resolution unit; the speech recognition engine comprises an acoustic model, a pronunciation dictionary and a language model; and the medical record generation module is used for converting the identification result into a structured electronic medical record. According to the method, a sliding window processing strategy is adopted, continuous voice input of a doctor can be effectively processed, a long-time voice input scene is supported, and various requirements in actual clinical application are met.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Chinese learner-oriented tone evaluation and improvement method

PendingCN121415801ASpeech recognitionTime domainVocal tract
The invention relates to the technical field of speech recognition, in particular to a Chinese learner-oriented tone evaluation and improvement method, which comprises the following steps of: extracting a fundamental frequency F0 curve and sound channel parameters of a speech signal, constructing a turbid and clear adaptive fusion model, and fusing to generate fundamental frequency related characteristics and fundamental frequency unrelated characteristics; constructing a tone error corpus, carrying out tone classification labeling, generating a multi-level tone feature set through a turbid and clear adaptive fusion model, aligning a voice signal with a reference voice time domain, and extracting and decoupling a tone shape feature vector and a tone domain feature vector; and based on the decoupled feature vectors, establishing a dual-channel evaluation path, hierarchically calculating the tone type distance and the tone domain distance of the tones, carrying out weighted fusion, generating a final evaluation score, and finally generating feedback information for the Chinese learner. According to the method, a complete method from feature decoupling to two-channel evaluation is constructed, the pronunciation problem of the learner is quantitatively diagnosed, and the pertinence and efficiency of Chinese tone learning are improved.
Owner:GUANGXI UNIV

Tone evaluation rehabilitation training device and system

A tone evaluation rehabilitation training device and system, the device comprises a flexible cap body and a bandage, the cap body is integrated with a 10-channel electrode plate to accurately cover language brain areas on both sides, electroencephalogram signals are transmitted in real time through Bluetooth 5.0, and a noise reduction earphone is used for providing sound output for a patient. The microphone is used for collecting audio of a patient and transmitting the audio to the processing module for analysis, the system dynamically regulates acoustic stimulation based on the neural activation feedback module, high-precision tone recognition is achieved in combination with audio preprocessing, SpecAugment data enhancement, a CNN-Transform mixed model and an attention mechanism, and training efficiency is optimized through pre-training model fine tuning and Warmup learning rate scheduling. An initial consonant-vowel-tone three-dimensional confusion probability table is introduced to divide interference intensity grades, high and low interference task paths are dynamically switched through a semantic error rate, and an evaluation module generates a four-tone accuracy rate, an F0 curve comparison graph and a personalized rehabilitation scheme. The recognition precision is improved, the training period of special crowds is shortened, and full-period self-adaptive rehabilitation is provided for preschool children to speech disorder patients.
Owner:林珍萍

Suzhou dialect speech recognition system and method based on tone track neural field

The invention provides a Suzhou dialect speech recognition system and method based on a tone track neural field, and the system comprises a tone track neural field module which is used for modeling the tone change of Suzhou dialects into a continuous space-time neural field; the bidirectional semantic memory network module comprises a forward prediction memory bank and a backward correction memory bank; the phoneme-font coupling error corrector is used for realizing polyphone disambiguation and homonym error correction by establishing association mapping between a phoneme sequence and a font sequence; the semantic entropy calculation module is used for evaluating the uncertainty of the recognition result; and the self-adaptive fusion decision module is used for generating a final recognition text. According to the invention, through an online learning mechanism, the system can continuously accumulate experience from actual use, automatically discover a new language mode and update an identification strategy. The self-improvement capability enables the system to adapt to the dynamic change of languages, and the performance is continuously improved along with the increase of the use time. Each use of the user helps the system to become more intelligent and accurate.
Owner:程思民

Cross-language speech recognition method, system and device and storage medium

The invention discloses a cross-language speech recognition method, system and device and a storage medium, and the method comprises the steps: carrying out the preprocessing of training speech, and obtaining a training speech frame sequence; extracting content characterization, speaker characterization and pitch characterization from the training voice frame sequence; performing voice reconstruction according to the content representation, the speaker representation and the pitch representation to obtain a target language voice; based on the training voice, constructing a target language voice recognition model according to the target language voice; in response to the target language recognition instruction, obtaining a target voice; and inputting the target voice into the target language voice recognition model to obtain a recognition result output by the target language voice recognition model, so that accurate modeling of dialect acoustic characteristics can be realized under the condition of extremely low annotation data through cross-language feature decoupling and a self-supervised migration mechanism. The recognition robustness of the complex tones and the characteristic vocabularies of the Guiwili is remarkably improved, and efficient generalization application in a dialect scene is achieved.
Owner:GUANGXI COMM IND SERVICE CO LTD +1

Melody note sequence generation method and system based on tone features and machine learning, electronic device, and storage medium

The application belongs to the technical field of audio processing, and specifically provides a melody note sequence generation method and system based on tone features and machine learning, an electronic device, and a storage medium. The method comprises obtaining a tone sequence based on an original Chinese text; constructing a rich context feature vector; pre-training a gradient boosting model; inputting the context feature vector into the pre-trained gradient boosting model to obtain the probability distribution of each category under the melody mode; concatenating the probability distribution of each category under the melody mode in order to generate a melody note sequence separated by spaces; and obtaining slow, clear, and rhythm-exaggerated singing audio suitable for aphasia MIT rehabilitation based on the melody note sequence and the Chinese text. The application automatically extracts Chinese tone features and predicts the melody mode with context awareness, thereby realizing high-quality generation from Chinese text to natural melody note sequence.
Owner:CHANGSHA LIANYU TECHNOLOGY CO LTD

A smart accompaniment singing method and system

ActiveCN121122218BEngineeringAudio frequency
This invention belongs to the field of audio processing technology and provides an intelligent accompaniment singing method and system. The method can collect the user's vocal pitch and voiceprint information, then divide the song selected based on the song selection command into vocal data and accompaniment data. The pitch of the accompaniment data is adjusted by comparing the vocal data with the user's vocal pitch, and the audio is played based on the adjusted accompaniment data. The user's singing data is continuously acquired and monitored, and when the user's singing data meets preset conditions, accompaniment vocals generated using the voiceprint information and vocal data are added to the audio. This proposed method can tailor the song's pitch to the user's individual needs, adjusting the accompaniment pitch to a suitable range for the user, ensuring easier singing. It can also selectively add accompaniment vocals based on the user's singing performance, thereby guiding the user to adjust their singing style, improve their singing level, and enhance their singing experience.
Owner:SHENZHEN WANSHENG CULTURE TECH CO LTD

An AI-generated text detection method, system, device and medium

The application discloses an AI generated text detection method, system, device and medium, and relates to the technical field of text processing. The application converts a Chinese text sentence to be detected into a tone category sequence composed of tone categories. In this process, the categories of Chinese tones are used as a quantitative index of phonological structure, and text analysis is converted from a high-dimensional and high-cost semantic space to a low-dimensional and high-efficiency phonological feature space. The application is not dependent on training data of a specific model and is adaptable to text detection of multiple models. Then, N-gram analysis is performed on the generated tone category sequence, and the occurrence frequency of each N-gram combination is calculated to form a feature vector. Finally, an AI generated text is judged through a pre-trained classification model. In this process, the calculation type in processing is mainly string processing and frequency statistics, and the inconsistency problem of the model in multiple model invocations is avoided, so that the AI generated text can be accurately detected.
Owner:FOSHAN UNIVERSITY

Video speech recognition method and system for illegal short video

The invention relates to the technical field of voice data processing, in particular to a video voice recognition method and system for illegal short videos, and the method comprises the steps: obtaining a prohibited lexicon and a to-be-detected video, and converting Chinese characters in the prohibited lexicon and the to-be-detected video into pinyin character strings which do not contain tones; the step of comparing each to-be-detected Chinese character with each forbidden word in the forbidden word bank is as follows: determining the similar weight of the forbidden word of each to-be-detected Chinese character; analyzing the closeness degree of serial numbers of the most similar Chinese characters in the forbidden words of the to-be-detected Chinese characters and the adjacent Chinese characters based on the pinyin character strings, and determining forbidden word matching values of the to-be-detected Chinese characters in combination with the forbidden word similar weights; judging whether each to-be-detected Chinese character is a prohibited Chinese character or not based on the distribution condition of the prohibited word matching value of each to-be-detected Chinese character and the adjacent Chinese character; and taking the to-be-detected video with the prohibited Chinese characters as a violation video. The invention aims to improve the detection accuracy and efficiency of illegal short videos.
Owner:CHANGAN COMM SCI & TECH CO LTD

Chinese pronunciation defect recognition method and system based on speech recognition

The invention provides a Chinese pronunciation defect recognition method and system based on speech recognition, and the method comprises the steps: obtaining a Chinese speech signal, generating an initial time-frequency spectrum through short-time Fourier transform, extracting a frequency feature sequence from the initial time-frequency spectrum, carrying out the segmentation processing of the speech signal through a time-domain signal segmentation technology, and carrying out the recognition of a Chinese pronunciation defect. Preliminary segmentation fragments are obtained; for the abnormal tone distribution diagram and the abnormal airflow distribution diagram, fusing tone duration and breathing rhythm mode features to generate a comprehensive feature matrix, performing boundary optimization on an abnormal interval through tone smoothing processing, and determining feature distribution after smoothing; and positioning a specific pronunciation defect syllable through the defect candidate interval in combination with conjoint analysis of a tone change mode and an airflow amplitude threshold value, tracing correlation characteristics of a frequency abrupt change point location and signal energy distribution, and judging a defect cause combination to obtain a final defect positioning result.
Owner:PUYANG VOCATIONAL & TECHN COLLEGE

An unsupervised speech recognition modeling method based on phoneme segment level representation discretization

ActiveCN119446136BSpeech recognitionCluster algorithmLaotian language
The present application relates to an unsupervised speech recognition modeling method based on phoneme segment level representation discretization, belonging to the field of speech recognition. The present application uses an IFMF model to extract phoneme features from the original audio through a speech feature discretization module, then trains a K-Means clustering algorithm to cluster the audio representation, and obtains the clustering index as a discrete label; through an adversarial learning module, a generator network and a discriminator are trained; through a speech discrete representation decoder module, a language model is used to decode the phoneme discretization-based model output obtained by unsupervised training. The present application fuses multiple speech features, adopts a phoneme segment level representation discretization method, considers the influence of tone information on the Lao language, and as much as possible reduces the influence of information redundancy caused by fine-grained features on cross-modal modeling. The method of the present application achieves competitive results compared with traditional speech recognition methods.
Owner:KUNMING UNIV OF SCI & TECH

Speech recognition method and system based on big data

The invention relates to the technical field of speech recognition, in particular to a speech recognition method and system based on big data, and the method comprises the steps: collecting each speech signal frame and the fundamental frequency thereof in a target region, obtaining a first matching index of each voice signal frame with a positive tone, a second matching index of each voice signal frame with a negative tone, a third matching index of each voice signal frame with a negative tone, and a fourth matching index of each voice signal frame with a positive tone; taking the maximum value of the first matching index, the second matching index, the third matching index and the fourth matching index corresponding to each voice signal frame as a characteristic value of each voice signal frame, obtaining a label of each characteristic value, and inputting each voice signal frame and the label of the characteristic value into a trained machine learning model, and obtaining the voice content of each voice signal frame. According to the invention, the problem of low voice content recognition accuracy is solved.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

A national vocal dialect prosody intelligent correction method and system

This invention proposes an intelligent error correction method and system for ethnic vocal music dialect prosody. The method includes: collecting multimodal data, extracting features to form a training dataset, constructing and updating a dynamic dialect prosody map, and performing cross-modal comparative analysis to obtain a joint representation of cross-modal error features. A neural vocoder direct-connect correction model is constructed, inputting the joint representation to generate a corrected Mel spectrum. A cross-modal generative adversarial network is constructed, including a generator and three discriminators. The generator generates a joint latent representation based on high-fidelity corrected audio and three features. The three discriminators respectively judge the naturalness of the audio, the compliance of the text tone, and the rhythmic coordination of the musical score. This invention solves the problem of manual error correction by constructing a prosody map, overcomes the bottleneck of manual detection by utilizing cross-modal comparative analysis, improves the generalization ability of the model by combining meta-learning algorithms, and optimizes the correction results with the help of adversarial networks, thereby reducing manual costs and improving the error correction efficiency of ethnic vocal music works.
Owner:GUIZHOU RADIO & TV UNIV

Tibetan speech recognition method fusing tone perception hybrid expert and search correction

The application discloses a Tibetan speech recognition method fusing tone perception hybrid experts and retrieval error correction, and belongs to the technical field of signal processing in the electronic industry. The specific steps of the recognition method are as follows: I: Tibetan speech signals are acquired, and corresponding log-mel spectrogram features and fundamental frequency contour features are extracted; II: the tone gating weight is calculated based on the fundamental frequency contour features, and is dynamically routed to the corresponding hybrid expert network to acquire dialect-independent acoustic hidden layer features. The application effectively solves the model interference problem caused by the presence or absence of tones among multiple dialects, avoids the parameter conflict between tone dialects and non-tone dialects, significantly improves the recognition performance of multi-dialect hybrid training, corrects the homonym heterograph error commonly existing in Tibetan, improves the performance of multi-dialect Tibetan speech recognition, can be used for converting Tibetan speech into characters, and is helpful for protecting and mining Tibetan culture.
Owner:CHINA UNIVERSITY OF POLITICAL SCIENCE AND LAW

A method for evaluating and correcting spoken pronunciation

ActiveCN121862161BImprove the pertinence of correctionsImprove efficiencySpeech analysisSpoken languageSpeech sound
This invention relates to the field of speech recognition technology, specifically a spoken pronunciation correction and evaluation method. The method includes: constructing an audio time-series sequence containing multiple phonemes based on the standard text spoken by the user; comparing the audio features corresponding to different phonemes in the audio time-series sequence based on the standard pronunciation of each phoneme to locate the pronunciation deviation position corresponding to each phoneme; if the current phoneme has a pronunciation deviation position, performing contextual scene analysis on the pronunciation deviation position by combining the duration, interval, and tone of the preceding and following phonemes to determine the deviation type in the current phoneme's context; identifying the user's abnormal pronunciation patterns; and configuring correction weights for each abnormal pronunciation pattern based on the difference feature values ​​between any two abnormal pronunciation patterns, and determining the correction direction corresponding to each phoneme according to the correction weights of the abnormal pronunciation patterns. This achieves targeted and efficient spoken pronunciation correction.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Multi-Host Redundant Audio Interface System with Real-Time Vocal Pitch Correction and Enhanced Signal Processing

An audio interface comprises a CPU (central processing unit) that includes a DSP (digital signal processor), a MIDI (musical instrument digital interface) processor, and two Voc FX (vocal effects) modules; MIDI IN, MIDI OUT, and MIDI HOST ports connected to the MIDI processor; first and second microphone ports connected to the first and second Voc FX modules, respectively; a plurality of OUTPUT ports connected to the DSP; and first and second HOST ports connected to the DSP; wherein the audio interface provides multi-host redundancy, real-time signal processing and MIDI control,
Owner:APOGEE ELECTRONICS CORP

Method, apparatus, and medium for recognizing a voice tone

ActiveCN120319275BSpeech recognitionAcousticsVowel
Provided are a method, apparatus, and medium for recognizing tones of speech. The method comprises: detecting the duration position of a vowel of a vowel in a speech to be recognized; detecting a tone core portion of the speech to be recognized based on the duration position of the vowel of the vowel; and identifying the tone category of the speech to be recognized based on the tone core portion. Thus, according to at least one embodiment of the present disclosure, the tone core portion can be more accurately detected within the duration position of the vowel of the vowel, thereby more accurately identifying the tone category.
Owner:NEW ORIENTAL EDUCATION & TECH GRP CO LTD

Speech synthesis method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, and provides a speech synthesis method, device and equipment and a medium, which are applied to the fields of medical treatment, finance and the like, and the method comprises the following steps: integrating initial speech data to obtain target speech data; extracting the target voice data to obtain target phoneme features; fusing the target voice data and the target phoneme features to obtain target fusion features; and synthesizing the target fusion feature based on a synthesis strategy to obtain a target synthesized speech. According to the embodiment of the invention, the method achieves the integration processing of the initial voice data to obtain the target voice data, carries out the extraction and fusion processing of the target voice data to obtain the target fusion features, and carries out the synthesis processing of the target fusion features based on the synthesis strategy to obtain the target synthesis voice. The tone accuracy and the phoneme alignment accuracy of the synthesized voice are improved, the method has good field adaptability and rhythm continuity, high fidelity of small sample voice cloning is ensured, and therefore the synthesis efficiency is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Spoken language pronunciation correction evaluation method based on deep learning

The invention relates to the technical field of speech recognition, in particular to a spoken language pronunciation correction evaluation method based on deep learning, which comprises the following steps: constructing an audio time sequence containing a plurality of phonemes on the basis of a standard text when a user pronunciates; on the basis of the standard pronunciation of each phoneme, comparing audio features corresponding to different phonemes in the audio time sequence, and positioning a pronunciation deviation position corresponding to each phoneme; if the current phoneme has the pronunciation deviation position, performing context scene analysis on the pronunciation deviation position by combining the duration, interval and tone of the previous and next phonemes, and determining the deviation type of the current phoneme in the scene; identifying an abnormal pronunciation mode of the user; and based on the difference characteristic value between any two abnormal pronunciation modes, the correction weight of each abnormal pronunciation mode is configured, and the correction direction corresponding to each phoneme is determined according to the correction weight of the abnormal pronunciation mode. The pertinence and the processing efficiency of spoken language pronunciation correction are realized.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Song identification method, electronic equipment, storage medium and product

The invention discloses a song recognition method, electronic equipment, a storage medium and a product. The method comprises the following steps: acquiring first audio data; determining first syllable information and first tone information based on the first audio data; determining a first syllable duration based on the first syllable information; determining second audio data corresponding to the first audio data based on the first tone information, the first syllable duration and a database; the database comprises at least one piece of second audio data and second tone information and second syllable duration corresponding to the second audio data; the first syllable duration represents the duration of each syllable in the first audio data, and the second syllable duration represents the duration of each syllable in the second audio data. The song is searched by using the tone in the audio and the duration of each syllable, so that the song identification flexibility is improved, and the identification accuracy is improved.
Owner:MIGU CO LTD +1

Dialect speech recognition method, system and model

The present application belongs to the technical field of speech recognition, and particularly relates to a dialect speech recognition method, system and model. The dialect speech recognition method comprises: S1, generating pseudo labels and screening by a plurality of teacher models; S2, sub-model training; S3, data fusion; S4, student model training, repeating steps S3 to S4 until round fusion training is performed, and student models of each region are obtained after training. The present application combines multi-teacher model knowledge distillation and K nearest neighbor dynamic data fusion training to construct an end-to-end speech transcription method suitable for a multi-dialect scene, and is particularly suitable for a less resourceful dialect environment with a complex tone system, rich speech variants and scarce data.
Owner:GUIZHOU UNIV +1

Han-Zhuang translation and speech synthesis system based on Tacotron2 and HiFi-GAN

The invention discloses a Chinese-Zhuang translation and speech synthesis system based on Tacotron2 and HiFi-GAN, the system comprises text preprocessing, Chinese-Zhuang mapping and speech synthesis module realization, the text preprocessing comprises text standardization, text legality verification and word segmentation and part-of-speech tagging, the Chinese-Zhuang mapping is used for utilizing international phonetic symbol correspondence, and the Chinese-Zhuang translation and speech synthesis module is used for realizing Chinese-Zhuang translation and speech synthesis. The tone characteristics of Chinese can be reasonably converted into the tone mode of Zhuang language, and the implementation of the speech synthesis module comprises Tacotron2 model processing and a vocoder. Compared with the prior art, the method has the advantages that a seamless conversion system platform from Chinese text to Zhuang language speech synthesis is provided, the minority language is protected and inherited, convenient speech interaction experience is provided for Zhuang language users, the accessibility of information is enhanced, and the user experience is improved. The method is especially suitable for visually impaired people or scenes requiring hand-eye liberation.
Owner:GUANGXI UNIV FOR NATITIES +1

Courseware making method and device based on artificial intelligence and storage medium

PendingCN122002103AImprove participationEnhance sense of accomplishmentSelective content distributionElectrical appliancesMachine learningData science
The invention relates to the technical field of artificial intelligence, in particular to a courseware making method and device based on artificial intelligence and a storage medium, and the method comprises the steps: constructing a tone problem distribution table based on tone accuracy scores and machine bias labels; determining a target training group based on the tone question distribution table, further determining a courseware generation parameter, extracting a recent continuous practice sequence based on the target training group, and further generating a training intensity adjustment parameter to adjust the courseware generation parameter; constructing an individual cognitive bias error model based on listen-and-read entries in a user historical period, and generating a bias error pair list; generating training content based on the courseware generation parameters, the target training group, the bias error pair list and the tone problem distribution table; and arranging a teaching process based on the training content, and generating a courseware script. According to the invention, the pronunciation data, the personal bias mode and the linguistic knowledge base of the learner can be dynamically combined, and an accurate closed loop from data diagnosis to targeted teaching is realized.
Owner:BEIJING QIXING INTERACTIVE EDUCATION TECHNOLOGY CO LTD

Multi-language voice test method and device, computer equipment and storage medium

The embodiment of the invention discloses a multi-language voice testing method and device, computer equipment and a storage medium, and relates to the technical field of voice testing. The method comprises the following steps: acquiring standard pronunciation data and accent feature data; extracting frequency spectrum features and tone features of the accent feature data; determining an accent model according to the spectrum features and the tone features; generating an accent voice test instruction based on the accent model and the standard pronunciation data through a preset neural voice synthesizer; and testing a voice assistant based on the accent voice test instruction to obtain a test result. According to the automatic voice test process constructed by the invention, the systematic bottleneck that the traditional voice test depends on manual recording is fundamentally solved through multi-stage technology collaboration. The core value of the method is that the voice test is converted into an extensible digital generation process from labor-intensive operation.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Multi-mode Chinese tone mouth shape auxiliary teaching system based on multiple attention mechanisms

The invention relates to the technical field of language teaching, and provides a multi-modal Chinese tone mouth shape auxiliary teaching system based on multiple attention mechanisms, and the system comprises a plurality of modules: a self-adaptive vocal cord vibration signal extraction module which is used for collecting and purifying a voice signal, extracting a fundamental frequency and scoring; the multi-scale tone feature fusion module receives the fundamental frequency signals and scores, extracts features and constructs vectors; the hierarchical tone recognition module processes the feature vectors through a neural network and an attention mechanism, recognizes tones and extracts voiceprint features; the personalized vocal organ motion modeling module selects and generates a personalized vocal organ motion model based on vocal print features; the mouth shape and tongue position dynamic generation module creates a three-dimensional visual mouth shape animation; the pronunciation feedback guidance module compares the pronunciation of the user with the standard pronunciation and provides multi-modal feedback; the tone recognition accuracy is improved through the adaptive signal extraction technology, and an effective auxiliary tool is provided for Chinese teaching.
Owner:KUNMING UNIV OF SCI & TECH

A semantic analysis method for speech in a crowded environment

The application discloses a kind of personnel dense environment under the semantic analysis method of voice, comprising: acquisition obtains the voice signal data of target user, lip movement video stream data;Extract lip movement feature sequence, map lip movement feature sequence into predicted voice feature vector;Input voiceprint separation model, separate the voice segment of target user from mixed voice signal, generate pure voice data characteristics;Pure voice data characteristics are carried out time domain segmentation, obtain the phonetic time domain waveform of single word;Single word phonetic time domain waveform is matched with initial, final time domain waveform library, obtains the corresponding pinyin expression of each word;Tone combination correlation analysis is carried out to the pinyin expression of continuous single word, obtains the meaning of the voice segment of target user.The application has the advantages that: by combining lip movement video stream and voice signal data, using deep learning and voiceprint separation technology, the voice of target user is effectively extracted, and the voice recognition accuracy in noisy environment is significantly improved.
Owner:SHENYANG LINKTECH INFORMATION TECH CO LTD

Dialect speech recognition method, system and model

The invention belongs to the technical field of speech recognition, and particularly relates to a dialect speech recognition method, system and model. The dialect speech recognition method comprises the following steps: S1, a plurality of teacher models generate pseudo tags and screen the pseudo tags; s2, training a sub-model; s3, data fusion; and S4, performing student model training, repeating the steps S3 to S4 until round fusion training is performed, and obtaining the student model which is trained in each region. According to the invention, through combination of multi-teacher model knowledge distillation and K-nearest neighbor dynamic data fusion training, an end-to-end voice transcription method suitable for a multi-dialect scene is constructed, and the method is especially suitable for a low-resource dialect environment with a complex tone system, rich voice variants and scarce data.
Owner:GUIZHOU UNIV +1

Speech rehabilitation training and dynamic feedback system and method based on tone-gesture mapping

The invention relates to a speech rehabilitation training and dynamic feedback system and method based on tone-gesture mapping, and the system comprises a processor, a display screen, an audio input device, an audio output device, an audio recognition module, a tone standard deviation analysis module, a detection result output module, and a gesture coding module. The audio recognition module comprises a target word access module, an audio acquisition module and an audio analysis module, and the target word access module is used for storing and calling standard tones of target words; the audio acquisition module acquires user audio information; the audio analysis module is used for extracting acoustic features of a user and aligning a user tone with a standard tone by using a dynamic time warping algorithm; the tone standard degree deviation analysis module performs normalization processing, and calculates the deviation distance between the user tone and the standard tone, the standard deviation degree and the user pronunciation accuracy; and the gesture coding module is used for gesture guidance action display regulation and control. The method is high in efficiency, high in adaptability and easy to identify and expand.
Owner:HANGZHOU NORMAL UNIVERSITY