Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Speech disturbances" patented technology

Speech disorders or speech impediments are a type of communication disorder where 'normal' speech is disrupted. This can mean stuttering, lisps, etc. Someone who is unable to speak due to a speech disorder is considered mute.

Voice disorder detection method, device and equipment and readable storage medium

The invention relates to a voice disorder detection method, device and equipment and a readable storage medium. The method comprises the following steps: acquiring a standard audio-visual material, responding to a pronunciation operation of a to-be-detected object on the standard audio-visual material, acquiring multi-modal pronunciation data, extracting audio acoustic features based on the pronunciation audio, extracting video visual features based on videos of facial and oral activities, and displaying the extracted audio acoustic features and the extracted video visual features. And performing multi-modal feature fusion on the audio acoustic features, the video visual features and the demographic information coding data to obtain a fusion feature vector, and based on the fusion feature vector and a pre-trained prediction model, obtaining a voice obstacle detection result of the to-be-detected object. Compared with the prior art, according to the embodiment of the invention, through multi-modal feature fusion, the accuracy and comprehensiveness of voice disorder detection can be improved, the diagnosis efficiency can be improved, the dependence on professionals can be reduced, the burden of medical resources can be reduced, the specific type of pronunciation problems can be clarified, and a scientific basis can be provided for subsequent personalized intervention treatment.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Chinese learner-oriented tone evaluation and improvement method

PendingCN121415801ASpeech recognitionTime domainVocal tract
The invention relates to the technical field of speech recognition, in particular to a Chinese learner-oriented tone evaluation and improvement method, which comprises the following steps of: extracting a fundamental frequency F0 curve and sound channel parameters of a speech signal, constructing a turbid and clear adaptive fusion model, and fusing to generate fundamental frequency related characteristics and fundamental frequency unrelated characteristics; constructing a tone error corpus, carrying out tone classification labeling, generating a multi-level tone feature set through a turbid and clear adaptive fusion model, aligning a voice signal with a reference voice time domain, and extracting and decoupling a tone shape feature vector and a tone domain feature vector; and based on the decoupled feature vectors, establishing a dual-channel evaluation path, hierarchically calculating the tone type distance and the tone domain distance of the tones, carrying out weighted fusion, generating a final evaluation score, and finally generating feedback information for the Chinese learner. According to the method, a complete method from feature decoupling to two-channel evaluation is constructed, the pronunciation problem of the learner is quantitatively diagnosed, and the pertinence and efficiency of Chinese tone learning are improved.
Owner:GUANGXI UNIV

Artificial Intelligence based Speech and Language Therapy and Language Learning which utilizes Facial Recognition, Voice Recognition and Character Avatars

The invention provides an AI-powered speech, language therapy, and language learning system that utilizes Natural Language Processing (NLP), Convolutional Neural Networks (CNNs), and Generative Adversarial Networks (GANs) to deliver personalized, real-time therapy and learning for individuals with speech disorders or those seeking to improve language proficiency. The system analyzes user speech, language comprehension, and facial expressions, providing immediate feedback on pronunciation, fluency, articulation, and sentence structure. A GAN-generated avatar interacts with the user, mimicking human expressions and offering dynamic, engaging sessions. The platform adapts exercises based on user performance using personalized algorithms to ensure continuous progress. Additionally, it securely stores user data in compliance with privacy regulations, making it accessible through web and mobile platforms. This invention improves upon existing speech therapy and language learning solutions by integrating real-time visual and auditory feedback with AI-driven personalization, offering a more immersive and effective experience.
Owner:ALI SYED AHAD

Speech recognition and training method for speech disorder rehabilitation patient

ActiveCN120612926BSpeech recognitionSpeech disorderSpeech sound
The present application relates to the technical field of speech recognition, in particular to a speech recognition and training method for speech disorder rehabilitation patients. The present application first extracts the text in the speech data, and obtains the problem words and the standard scores of the text of the patient's pronunciation; further constructs the feature vector of each text, and classifies the standard pronunciation text and the text of the patient's pronunciation; further obtains the pronunciation disorder degree of each problem word of the patient according to the text difference and the feature vector difference between the pronunciation word class cluster and the corresponding standard word class cluster, in combination with the standard score of the text of the patient and the proportion of the problem word; finally, according to the distribution of the problem word in the sentence, in combination with the pronunciation disorder degree and the appearance frequency of the problem word, the pronunciation difficulty of each problem word is obtained, the personalized pronunciation disorder of the patient is accurately identified, and personalized training schemes are formulated for relevant personnel, thereby providing support for improving the intervention effect.
Owner:PLA ARMY 82ND GRP MILITARY HOSPITAL

Vocal music training method and system based on artificial intelligence

The invention relates to the technical field of voice analysis, and particularly discloses a vocal music training method and system based on artificial intelligence, and the method comprises the steps: collecting and analyzing the body posture parameters of a target vocal music training person, carrying out the preprocessing, detecting the posture abnormality, forming a first abnormal training set, capturing and analyzing the pronunciation feature parameters, and recognizing the pronunciation abnormality, thereby obtaining a vocal music training result; and integrating the first abnormal training set and the second abnormal training set to generate a total abnormal parameter training set, matching a correction set, and visually displaying the body posture abnormal degree and the abnormal reason and the pronunciation abnormal degree and the reason, so that the target vocal music training personnel can visually see own posture and pronunciation problems, and self-adjustment and correction are facilitated.
Owner:SHANGLUO UNIV

Tone evaluation rehabilitation training device and system

A tone evaluation rehabilitation training device and system, the device comprises a flexible cap body and a bandage, the cap body is integrated with a 10-channel electrode plate to accurately cover language brain areas on both sides, electroencephalogram signals are transmitted in real time through Bluetooth 5.0, and a noise reduction earphone is used for providing sound output for a patient. The microphone is used for collecting audio of a patient and transmitting the audio to the processing module for analysis, the system dynamically regulates acoustic stimulation based on the neural activation feedback module, high-precision tone recognition is achieved in combination with audio preprocessing, SpecAugment data enhancement, a CNN-Transform mixed model and an attention mechanism, and training efficiency is optimized through pre-training model fine tuning and Warmup learning rate scheduling. An initial consonant-vowel-tone three-dimensional confusion probability table is introduced to divide interference intensity grades, high and low interference task paths are dynamically switched through a semantic error rate, and an evaluation module generates a four-tone accuracy rate, an F0 curve comparison graph and a personalized rehabilitation scheme. The recognition precision is improved, the training period of special crowds is shortened, and full-period self-adaptive rehabilitation is provided for preschool children to speech disorder patients.
Owner:林珍萍

Text adversarial sample semantic improvement method based on context preservation

The invention discloses a text adversarial sample semantic improvement method based on context preservation. The method comprises the following steps: constructing a keyword space and a part-of-speech disturbance space on an unlabeled public data set; determining a semantic disturbance position, and generating a knowledge label; carrying out reasonable sorting on the knowledge labels after error correction check; combining the sorted knowledge labels and the public data set through special tokens to obtain annotated data; performing word segmentation on each piece of labeling data through a predefined mapping function to obtain a token sequence, and performing text coding filling on the token sequence to obtain a filled token sequence; and completing multi-task semantic adaptive training by using the filled token sequence, and performing text adversarial sample generation based on the mask language model by using the mask language model after semantic training. According to the method, the semantic consistency and the attack efficiency of the text adversarial sample generated based on the mask language model can be effectively improved under the condition that the quality of the generated text is not influenced.
Owner:SOUTH CHINA UNIV OF TECH

English pronunciation practice auxiliary device

The invention relates to the technical field of language learning, and discloses an English pronunciation practice auxiliary device, which comprises a storage plate, the surface of the storage plate is provided with a radio assembly and a camera assembly, the upper surface of the storage plate is provided with an entertainment area, and the entertainment area is internally provided with an entertainment mechanism which is used for simulating a cartoon situation. The children's attention is attracted through dynamic actions, the entertainment mechanism comprises a vibration assembly and a moving assembly, and the vibration assembly is used for achieving ejector rod vibration through electromagnetic driving and generating tactile feedback of different frequencies and intensities. By integrating the functions of audio processing, dynamic visual analysis and entertainment interaction, pronunciation data and dynamic facial features of a practicer can be collected in real time, accurate pronunciation problem recognition is achieved through a scientific algorithm, the learning interest and pronunciation accuracy of the practicer are improved through multi-sensory interaction, and the system is particularly suitable for the language learning requirements of children.
Owner:HENAN UNIV OF URBAN CONSTR

Apparatus for the self-administration of therapies for the treatment of speech disorders

PCT designated stageWO2026028021A1Psychotechnic devicesSensorsAuditory stimuliSound sources
The present invention relates to an apparatus (10) for the self-administration of therapies for the treatment of speech disorders comprising a stimulus delivery device (11,12) with a horizontal development substantially curved along an angular portion so as to outline a portion of a circle, comprising at least one panel (11) for the delivery of visual stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); at least one acoustic source (12) for the delivery of acoustic stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); and a local electronic processing unit configured to control the generation of visual and acoustic stimuli by means of the at least one panel (11) and the at least one acoustic source (12), and characterized in that it comprises at least one first acoustic detector (15) for detecting a reproduction of a verbal element produced by the patient, wherein the at least one first acoustic detector (15) is carried by the at least one panel (11) and positioned so as to detect acoustic signals generated in a space inside the portion of circle outlined by the stimulus delivery device (11,12).
Owner:LINARI MEDICAL SRL

A brain-computer interface system for recognizing the intention of Chinese oral language based on a sound-meaning integration double model

ActiveCN121560160BSpoken languageStereotaxis
The application provides a Chinese spoken language intention recognition brain-computer interface system based on a sound-meaning integration double model, belongs to the technical field of biomedical engineering, and relates to language brain-computer interface technology. Taking sound-meaning integration as the core, the stereotactic intracranial electroencephalogram (sEEG) technology is adopted to collect neural signals of the brain articulatory motor coding area and the semantic concept organization coding area. The system comprises a voice initiation decoder, a speech decoder, a semantic decoder and a Chinese word speech-semantic fusion synthesizer, the target decoder is constructed by extracting high gamma band features of key brain areas of the frontal lobe (left inferior frontal gyrus, premotor cortex, etc.), the temporal lobe (anterior temporal lobe, dorsolateral temporal lobe, etc.). At the same time, a visual and auditory induction training paradigm is matched, three tasks of listening to sound to group words, looking at words to group words and word association are set, and the subjects are supported to generate words independently. The system effectively solves the homonym and near homonym word ambiguity problem in Chinese spoken language recognition, and provides a precise interactive tool for ALS and other speech disorder patients.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models

The invention provides a Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models, belongs to the technical field of biomedical engineering, and relates to a language brain-computer interface technology. Sound-sense integration is taken as a core, and a stereotactic intracranial electroencephalogram (sEEG) technology is adopted to acquire neural signals of a brain phonetic motion coding region and a semantic concept organization coding region. The system comprises a sound production starting decoder, a voice decoder, a semantic decoder and a Chinese word voice-semantic fusion synthesizer, and a target decoder is constructed by extracting high gamma wave band characteristics of key brain regions of frontal lobe (left subfrontal gyrus, anterior cortex of motion and the like) and temporal lobe (anterior temporal lobe, dorsal lateral temporal lobe and the like). Meanwhile, an audio-visual induction training normal form is matched, three tasks of listening word combination, character reading word combination and vocabulary association are set, and subjects are supported to autonomously generate vocabularies. The system effectively solves the problem of ambiguity of homophonous and near-phonetic words in spoken Chinese recognition, and provides an accurate interaction tool for speech disorder patients such as ALS and the like.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Speech disorder detection method, device, equipment and readable storage medium

The present disclosure relates to a speech disorder detection method, device, equipment and readable storage medium. By acquiring standard audio-visual materials, in response to the pronunciation operation of the to-be-detected object to the standard audio-visual materials, multi-modal pronunciation data is collected, audio acoustic features are extracted based on the pronunciation audio, video visual features are extracted based on the video of the face and oral cavity activity, the audio acoustic features, the video visual features and the demographic information coding data are fused to obtain a fusion feature vector, and based on the fusion feature vector and a pre-trained prediction model, a speech disorder detection result of the to-be-detected object is obtained. Compared with the prior art, the embodiment of the present disclosure can improve the accuracy and comprehensiveness of speech disorder detection, improve the diagnosis efficiency, reduce the dependence on professionals, reduce the burden of medical resources, and clearly determine the specific type of pronunciation problem, thereby providing a scientific basis for subsequent individualized intervention treatment.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Voice interaction method and device, equipment and storage medium

The invention discloses a voice interaction method and device, equipment and a storage medium, and relates to the technical field of voice processing, and the method comprises the steps: determining an optimal voice collection mode of a voice signal under the condition that the voice signal of a user is detected; acquiring a real-time voice signal of the user through the optimal voice acquisition mode; detecting whether abnormal voice signals exist in the real-time voice signals or not, wherein the abnormal voice signals are voice signals with pronunciation problems; if yes, repairing the real-time voice signal; and generating and executing a voice control instruction according to the restored voice signal. According to the invention, when the real-time voice signal of the user has the pronunciation problem, the real-time voice signal can be repaired, and the voice control instruction is generated and executed according to the repaired voice signal, so that the problem that the existing far-field voice recognition system cannot accurately recognize the voice instruction with the voice problem is solved, and the voice recognition efficiency is improved. And the voice interaction is limited.
Owner:SHENZHEN SKYWORTH DISPLAY TECH CO LTD

A method and device for generating speech synthesis information

The present invention provides a method and apparatus for generating speech synthesis information, which relates to the technical field of speech data processing and can be used in the financial field or other technical fields. The method includes: obtaining a target video, and extracting human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech disorders; fusing the human joint point features and the lip language features to obtain a fused feature, and performing text recognition on the fused feature to obtain text vocabulary information; performing speech synthesis on the text vocabulary information to obtain speech synthesis information. The apparatus executes the above method. The method and apparatus for generating speech synthesis information provided by the embodiments of the present invention can accurately and efficiently identify the speech information that customers with speech disorders want to express, and improve the business handling efficiency.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Language disorder adjuvant therapy system based on speech recognition

InactiveCN120932641ASpeech recognitionVocal organDimensional simulation
The invention belongs to the field of intelligent medical rehabilitation engineering, particularly relates to a speech recognition-based speech disorder adjuvant therapy system, and solves the problems of high pathological speech misjudgment rate, large dialect interference and single feedback in the prior art. A phoneme fault-tolerant threshold value is dynamically adjusted by constructing an adaptive speech recognition engine, regional features and pathological features are separated by using a dialect adaptation module, and three-dimensional simulation, AR real-time correction and touch graded vibration of vocal organs are realized in combination with a multi-modal feedback device; the pathological speech recognition precision is improved; and the neural remodeling efficiency is optimized.
Owner:JIANGXI DINGJI MINGCHUANG TECHNOLOGY CO LTD

Captioned telephone service system for user with speech disorder

ActiveUS12592993B2Special service for subscribersSpeech recognitionSpeech ProcessorSpeech disturbances
A captioned telephone service (CTS) system provides a transcription service and a speech-to-text and text-to-speech converting service for the deaf or hard-of-hearing user with a speech disorder during a phone call between the user and the peer. The CTS system transcribes the peer's voice into text to be displayed on the user's device. The CTS system transcribes the user's voice into text based on the database storing user's spoken audios and corresponding texts, and converts the text into a clear and articulate speech to be sent to the peer's device instead of the user's voice in order to help the peer better understand what the user said. The CTS system includes a speech-to-text handler and a text-to-speech handler. The speech-to-text handler transcribes user's voice into text using the database, and the text-to-speech handler converts the text into a clear and articulate speech.
Owner:MEZMO CORP

Speech data processing method and system based on large language model

The application discloses a speech data processing method and system based on a large language model, relates to the technical field of speech data processing, and comprises the following steps: performing multi-channel feature decomposition on a received original speech signal, and constructing an acoustic state representation tensor; constructing a semantic candidate distribution space, generating multiple sets of semantic hypothesis vectors, and constructing a semantic evolution path graph; generating a semantic uncertainty function representing semantic ambiguity and speech disturbance sensitivity; dynamically constructing a reasoning depth control parameter and inputting the same to a multi-layer reasoning path scheduling unit of the large language model, constructing an intention structure vector, and mapping the intention structure vector into a structured semantic output. The technical problems that in the prior art, under a complex acoustic environment, it is difficult to accurately and effectively separate acoustic features, leading to low speech understanding accuracy of high ambiguity, and lacking dynamic reasoning ability to cope with semantic uncertainty risks are solved, and the technical effects of improving semantic understanding precision, ambiguity resolution ability of speech interaction, and reducing business misjudgment rate and risk are achieved.
Owner:GUANGDONG JINWAN INFORMATION TECH CO LTD

Audio data processing methods, devices, media and equipment

This disclosure relates to an audio data processing method, apparatus, medium, and device, comprising: acquiring an audio to be evaluated and a text to be read aloud; extracting the speech to be evaluated from the audio; scoring one or more segments of speech content corresponding to each word in the text to be evaluated, to obtain one or more scoring results; determining the text reading evaluation corresponding to the text to be read aloud based on the one or more scoring results, including positive evaluation, negative evaluation, and neutral evaluation; determining an evaluation template based on the text reading evaluation corresponding to the text to be read aloud and generating feedback text; and synthesizing the feedback text into feedback speech and sending it to the user corresponding to the audio to be evaluated. In this way, all of the user's pronunciations can be scored and comprehensively evaluated, resulting in a more objective, reasonable, and comprehensive text reading evaluation. Furthermore, the feedback evaluation speech can be automatically generated based on the evaluation, allowing users to more intuitively understand their pronunciation problems and reducing the influence of subjective human factors.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Computer programs for terminal devices, computer programs for speech recognition servers, and communication systems

To provide a technology that can display appropriate text strings in situations where both the voice of a user with a speech impediment and the voice of a user without a speech impediment may be input. [Solution] The computer program causes the computer of the terminal device to function as follows: a supply unit that supplies voice data to a speech recognition engine; an acquisition unit that obtains a string set including first string data and second string data from the speech recognition engine; and a display control unit that causes the display unit to display a string screen containing a string corresponding to at least one of the first string data and the second string data. The first string data is data generated based on a first speech recognition model for a user with a speech disorder, and the second string data is data generated based on a second speech recognition model for a user without a speech disorder.
Owner:BROTHER KOGYO KK

System for establishing pronunciation disorder vocal cord vibration model

The invention belongs to the technical field of pronunciation disorder vocal cord vibration models, and particularly relates to a pronunciation disorder vocal cord vibration model establishing system which comprises an image acquisition module used for acquiring vocal cord vibration images by using a high-definition electronic stroboscopic laryngoscope; the image preprocessing module is used for performing enhancement, noise reduction and segmentation processing on the image acquired by the image acquisition module; and the vocal cord vibration feature extraction module is used for extracting vocal cord features according to the image processed by the image preprocessing module. The vocal cord region can be accurately identified and segmented by using a semantic segmentation algorithm based on deep learning, interference of complex backgrounds such as surrounding tissues and secretions can be effectively eliminated, the vocal cord tissues and the backgrounds can be accurately distinguished even if mucus adheres to the periphery of the vocal cord, accurate image data can be provided for subsequent feature extraction, and the vocal cord recognition accuracy is improved. And the reliability and accuracy of feature extraction are improved.
Owner:THE SECOND HOSPITAL OF TIANJIN MEDICAL UNIV

Voice error detection method and device, electronic equipment and storage medium

ActiveCN116052720BSpeech recognitionSpeech errorSpeech input
The present disclosure provides a speech error detection method and device, electronic equipment and storage medium, and relates to the technical field of speech processing. The method comprises obtaining a to-be-detected speech of a target title and a plurality of pronunciation data of the target title, the plurality of pronunciation data comprising target pronunciation data corresponding to the to-be-detected speech; inputting the to-be-detected speech into a pre-trained speech error detection model, and combining the plurality of pronunciation data to obtain a speech error detection result of the to-be-detected speech relative to the target pronunciation data. The method provided in the present disclosure can effectively handle the multi-pronunciation problem in speech error detection, improve the speech error detection effect in the multi-pronunciation scene, and avoid the problem of inaccurate results caused by relying on a single pronunciation mode for speech error detection of a multi-pronunciation title.
Owner:IFLYTEK CO LTD

Speech disorder assessment method and system based on multi-modal large model

The application discloses a speech disorder evaluation method and system based on a multi-modal large model. The system comprises: acquiring corresponding multi-modal data for a target, the multi-modal data including audio signals, lip videos and tongue ultrasonic images; performing cross-modal feature extraction and semantic alignment on the multi-modal data to map the extracted modal features to a unified semantic space aligned with a large model text embedding space, obtaining a semantic vector sequence; using the semantic vector sequence as input, simulating clinical multi-level reasoning logic using the large model to obtain a preliminary evaluation result of articulation disorder, the preliminary evaluation result including severity level and disorder type; using key information in the preliminary evaluation result to set a retrieval query strategy to guide the large model to generate an evaluation report and rehabilitation suggestions. The application improves the accuracy, real-time performance and robustness of speech disorder evaluation.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Captioned telephone service system for user with speech disorder

A captioned telephone service system is provided for assisting users with speech disorders in communicating with peers. The system comprises a CTS server, a user device, and a user application configured with a sentence refinement unit and a text-to-speech handler. The user inputs one or more words, which are processed by the sentence refinement unit to correct errors, insert missing grammar, and generate grammatically complete sentences. A contextual adaptation module analyzes conversation and peer information to select the most relevant sentence. The user confirms, regenerates, or edits the sentence, which is then converted to natural speech by the text-to-speech handler and transmitted to the peer's device. The system further employs user and peer databases to learn from prior interactions, thereby improving accuracy and facilitating effective communication for speech-impaired individuals.
Owner:MEZMO CORP

Speech recognition based sentence correction method and apparatus, device, and storage medium

The present application relates to the technical field of sentence correction, and particularly relates to a sentence correction method and device based on speech recognition, equipment and a storage medium, wherein a comprehensive bundle search technology is adopted to recognize speech data, and a first candidate sentence with the highest bundle search score and a plurality of candidate sentences are extracted in terms of semantic features and pinyin features, and feature fusion is performed; the first candidate sentence is corrected by using the fused features after feature fusion, thereby reducing the negative influence caused by pronunciation problems of users and multi-pronunciation character problems of texts, and improving the accuracy and efficiency of sentence correction.
Owner:GUANGZHOU YANLI NETWORK TECH CO LTD +3

Artificial intelligence-based vocal music training method and system

The present application relates to the technical field of voice analysis, and particularly discloses a vocal music training method and system based on artificial intelligence, which collects and analyzes body posture parameters of a target vocal music training person, detects posture abnormalities after preprocessing, forms a first abnormal training set, simultaneously captures and analyzes pronunciation feature parameters, identifies pronunciation abnormalities, forms a second abnormal training set, comprehensively generates an abnormal parameter training total set, and matches a correction set, and the abnormal degree and causes of body posture and the abnormal degree and causes of pronunciation are visually displayed, so that the target vocal music training person can directly see his / her own posture and pronunciation problems, and is facilitated to make self-adjustment and correction.
Owner:SHANGLUO UNIV

Personalized synthesis and recognition enhancement of dysarthric speech

ActiveCN120412540BSpeech synthesisDysarthric speechSpeech disorder
This invention discloses a personalized synthesis and recognition enhancement method for speech disorders. The speech disorder synthesis model includes a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module. The input of the speech disorder synthesis model includes samples, and the output includes synthesized speech disorder speech. The samples are speech disorder text sequences. The input of the long-range dependent feature encoding module includes samples, and the output is an alignment vector z. The input of the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The input of the decoding module is the final embedding representation, and the output is the synthesized speech disorder speech. The speech disorder synthesis model of this invention improves the ability to extract personalized features of speech disorder speech, enhances speech synthesis performance, and improves the fine-grained expression of speech disorder speech features.
Owner:TIANJIN UNIV

Tool for assisting people with speech disorder

ActiveUS12361951B2SensorsDiagnostic recording/measuringSpeech disorderSpeech disability
Various tools are disclosed for providing assistive or augmentative means to enhance the fluency and accuracy of persons having speech disabilities. These technologies may automatically ascertain and dynamically improve the accuracy with which automatic speech recognition (ASR) systems recognize utterances of persons having impaired speech conditions. In an embodiment, digitized audio information about a speaker's utterance is processed to determine a set of candidate words matching the utterance. From these candidate words, a set of concepts is determined using a finite state machine model. A pictogram representing each concept is identified and presented to the speaker so that the speaker may select the pictogram corresponding to the best match of his or her intended meaning associated with the utterance. An action corresponding to speaker's selection then may be performed. For example, displaying or synthesizing speech from textual information describing the selected concept.
Owner:CERNER INNOVATION INC

A method for evaluating oral conversation quality in college English learning

The present invention discloses a method for evaluating the quality of oral dialogues for college English learning, which relates to the field of data processing technology. The method comprises the following steps: constructing a dialogue knowledge database; segmenting the spoken dialogue speech into actual dialogue boxes; identifying dialogue application scenarios; obtaining semantic misunderstanding parts; obtaining a misunderstanding degree coefficient; analyzing the degree of pronunciation problems and the speed of initiating responses in the actual dialogue boxes; comprehensively obtaining a dialogue coherence index; obtaining a quality assessment value; obtaining a quality assessment value range using sample dialogue content; calculating the proportion of parts whose quality assessment values ​​exceed the quality assessment value range as a recognition ratio; and the greater the recognition ratio, the lower the quality of the spoken dialogue. By identifying the dialogue application scenarios, obtaining the semantic misunderstanding parts, obtaining the misunderstanding degree coefficient, and analyzing the degree of pronunciation problems and the speed of initiating responses, oral dialogues can be more accurately evaluated from multiple aspects.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY