Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Speech disturbances" patented technology

Speech disorders or speech impediments are a type of communication disorder where 'normal' speech is disrupted. This can mean stuttering, lisps, etc. Someone who is unable to speak due to a speech disorder is considered mute.

Chinese learner-oriented tone evaluation and improvement method

PendingCN121415801ASpeech recognitionTime domainVocal tract
The invention relates to the technical field of speech recognition, in particular to a Chinese learner-oriented tone evaluation and improvement method, which comprises the following steps of: extracting a fundamental frequency F0 curve and sound channel parameters of a speech signal, constructing a turbid and clear adaptive fusion model, and fusing to generate fundamental frequency related characteristics and fundamental frequency unrelated characteristics; constructing a tone error corpus, carrying out tone classification labeling, generating a multi-level tone feature set through a turbid and clear adaptive fusion model, aligning a voice signal with a reference voice time domain, and extracting and decoupling a tone shape feature vector and a tone domain feature vector; and based on the decoupled feature vectors, establishing a dual-channel evaluation path, hierarchically calculating the tone type distance and the tone domain distance of the tones, carrying out weighted fusion, generating a final evaluation score, and finally generating feedback information for the Chinese learner. According to the method, a complete method from feature decoupling to two-channel evaluation is constructed, the pronunciation problem of the learner is quantitatively diagnosed, and the pertinence and efficiency of Chinese tone learning are improved.
Owner:GUANGXI UNIV

Apparatus for the self-administration of therapies for the treatment of speech disorders

PCT designated stageWO2026028021A1Psychotechnic devicesSensorsAuditory stimuliSound sources
The present invention relates to an apparatus (10) for the self-administration of therapies for the treatment of speech disorders comprising a stimulus delivery device (11,12) with a horizontal development substantially curved along an angular portion so as to outline a portion of a circle, comprising at least one panel (11) for the delivery of visual stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); at least one acoustic source (12) for the delivery of acoustic stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); and a local electronic processing unit configured to control the generation of visual and acoustic stimuli by means of the at least one panel (11) and the at least one acoustic source (12), and characterized in that it comprises at least one first acoustic detector (15) for detecting a reproduction of a verbal element produced by the patient, wherein the at least one first acoustic detector (15) is carried by the at least one panel (11) and positioned so as to detect acoustic signals generated in a space inside the portion of circle outlined by the stimulus delivery device (11,12).
Owner:LINARI MEDICAL SRL

A brain-computer interface system for recognizing the intention of Chinese oral language based on a sound-meaning integration double model

ActiveCN121560160BSpoken languageStereotaxis
The application provides a Chinese spoken language intention recognition brain-computer interface system based on a sound-meaning integration double model, belongs to the technical field of biomedical engineering, and relates to language brain-computer interface technology. Taking sound-meaning integration as the core, the stereotactic intracranial electroencephalogram (sEEG) technology is adopted to collect neural signals of the brain articulatory motor coding area and the semantic concept organization coding area. The system comprises a voice initiation decoder, a speech decoder, a semantic decoder and a Chinese word speech-semantic fusion synthesizer, the target decoder is constructed by extracting high gamma band features of key brain areas of the frontal lobe (left inferior frontal gyrus, premotor cortex, etc.), the temporal lobe (anterior temporal lobe, dorsolateral temporal lobe, etc.). At the same time, a visual and auditory induction training paradigm is matched, three tasks of listening to sound to group words, looking at words to group words and word association are set, and the subjects are supported to generate words independently. The system effectively solves the homonym and near homonym word ambiguity problem in Chinese spoken language recognition, and provides a precise interactive tool for ALS and other speech disorder patients.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models

The invention provides a Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models, belongs to the technical field of biomedical engineering, and relates to a language brain-computer interface technology. Sound-sense integration is taken as a core, and a stereotactic intracranial electroencephalogram (sEEG) technology is adopted to acquire neural signals of a brain phonetic motion coding region and a semantic concept organization coding region. The system comprises a sound production starting decoder, a voice decoder, a semantic decoder and a Chinese word voice-semantic fusion synthesizer, and a target decoder is constructed by extracting high gamma wave band characteristics of key brain regions of frontal lobe (left subfrontal gyrus, anterior cortex of motion and the like) and temporal lobe (anterior temporal lobe, dorsal lateral temporal lobe and the like). Meanwhile, an audio-visual induction training normal form is matched, three tasks of listening word combination, character reading word combination and vocabulary association are set, and subjects are supported to autonomously generate vocabularies. The system effectively solves the problem of ambiguity of homophonous and near-phonetic words in spoken Chinese recognition, and provides an accurate interaction tool for speech disorder patients such as ALS and the like.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Speech disorder detection method, device, equipment and readable storage medium

ActiveCN120318639BCharacter and pattern recognitionSensorsSpeech disorderSpeech sounds
The present disclosure relates to a speech disorder detection method, device, equipment and readable storage medium. By acquiring standard audio-visual materials, in response to the pronunciation operation of the to-be-detected object to the standard audio-visual materials, multi-modal pronunciation data is collected, audio acoustic features are extracted based on the pronunciation audio, video visual features are extracted based on the video of the face and oral cavity activity, the audio acoustic features, the video visual features and the demographic information coding data are fused to obtain a fusion feature vector, and based on the fusion feature vector and a pre-trained prediction model, a speech disorder detection result of the to-be-detected object is obtained. Compared with the prior art, the embodiment of the present disclosure can improve the accuracy and comprehensiveness of speech disorder detection, improve the diagnosis efficiency, reduce the dependence on professionals, reduce the burden of medical resources, and clearly determine the specific type of pronunciation problem, thereby providing a scientific basis for subsequent individualized intervention treatment.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Voice interaction method and device, equipment and storage medium

The invention discloses a voice interaction method and device, equipment and a storage medium, and relates to the technical field of voice processing, and the method comprises the steps: determining an optimal voice collection mode of a voice signal under the condition that the voice signal of a user is detected; acquiring a real-time voice signal of the user through the optimal voice acquisition mode; detecting whether abnormal voice signals exist in the real-time voice signals or not, wherein the abnormal voice signals are voice signals with pronunciation problems; if yes, repairing the real-time voice signal; and generating and executing a voice control instruction according to the restored voice signal. According to the invention, when the real-time voice signal of the user has the pronunciation problem, the real-time voice signal can be repaired, and the voice control instruction is generated and executed according to the repaired voice signal, so that the problem that the existing far-field voice recognition system cannot accurately recognize the voice instruction with the voice problem is solved, and the voice recognition efficiency is improved. And the voice interaction is limited.
Owner:SHENZHEN SKYWORTH DISPLAY TECH CO LTD

Captioned telephone service system for user with speech disorder

ActiveUS12592993B2Special service for subscribersSpeech recognitionSpeech ProcessorSpeech disturbances
A captioned telephone service (CTS) system provides a transcription service and a speech-to-text and text-to-speech converting service for the deaf or hard-of-hearing user with a speech disorder during a phone call between the user and the peer. The CTS system transcribes the peer's voice into text to be displayed on the user's device. The CTS system transcribes the user's voice into text based on the database storing user's spoken audios and corresponding texts, and converts the text into a clear and articulate speech to be sent to the peer's device instead of the user's voice in order to help the peer better understand what the user said. The CTS system includes a speech-to-text handler and a text-to-speech handler. The speech-to-text handler transcribes user's voice into text using the database, and the text-to-speech handler converts the text into a clear and articulate speech.
Owner:MEZMO CORP

Speech data processing method and system based on large language model

The application discloses a speech data processing method and system based on a large language model, relates to the technical field of speech data processing, and comprises the following steps: performing multi-channel feature decomposition on a received original speech signal, and constructing an acoustic state representation tensor; constructing a semantic candidate distribution space, generating multiple sets of semantic hypothesis vectors, and constructing a semantic evolution path graph; generating a semantic uncertainty function representing semantic ambiguity and speech disturbance sensitivity; dynamically constructing a reasoning depth control parameter and inputting the same to a multi-layer reasoning path scheduling unit of the large language model, constructing an intention structure vector, and mapping the intention structure vector into a structured semantic output. The technical problems that in the prior art, under a complex acoustic environment, it is difficult to accurately and effectively separate acoustic features, leading to low speech understanding accuracy of high ambiguity, and lacking dynamic reasoning ability to cope with semantic uncertainty risks are solved, and the technical effects of improving semantic understanding precision, ambiguity resolution ability of speech interaction, and reducing business misjudgment rate and risk are achieved.
Owner:GUANGDONG JINWAN INFORMATION TECH CO LTD

Computer programs for terminal devices, computer programs for speech recognition servers, and communication systems

To provide a technology that can display appropriate text strings in situations where both the voice of a user with a speech impediment and the voice of a user without a speech impediment may be input. [Solution] The computer program causes the computer of the terminal device to function as follows: a supply unit that supplies voice data to a speech recognition engine; an acquisition unit that obtains a string set including first string data and second string data from the speech recognition engine; and a display control unit that causes the display unit to display a string screen containing a string corresponding to at least one of the first string data and the second string data. The first string data is data generated based on a first speech recognition model for a user with a speech disorder, and the second string data is data generated based on a second speech recognition model for a user without a speech disorder.
Owner:BROTHER KOGYO KK

Speech disorder assessment method and system based on multi-modal large model

The application discloses a speech disorder evaluation method and system based on a multi-modal large model. The system comprises: acquiring corresponding multi-modal data for a target, the multi-modal data including audio signals, lip videos and tongue ultrasonic images; performing cross-modal feature extraction and semantic alignment on the multi-modal data to map the extracted modal features to a unified semantic space aligned with a large model text embedding space, obtaining a semantic vector sequence; using the semantic vector sequence as input, simulating clinical multi-level reasoning logic using the large model to obtain a preliminary evaluation result of articulation disorder, the preliminary evaluation result including severity level and disorder type; using key information in the preliminary evaluation result to set a retrieval query strategy to guide the large model to generate an evaluation report and rehabilitation suggestions. The application improves the accuracy, real-time performance and robustness of speech disorder evaluation.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Captioned telephone service system for user with speech disorder

A captioned telephone service system is provided for assisting users with speech disorders in communicating with peers. The system comprises a CTS server, a user device, and a user application configured with a sentence refinement unit and a text-to-speech handler. The user inputs one or more words, which are processed by the sentence refinement unit to correct errors, insert missing grammar, and generate grammatically complete sentences. A contextual adaptation module analyzes conversation and peer information to select the most relevant sentence. The user confirms, regenerates, or edits the sentence, which is then converted to natural speech by the text-to-speech handler and transmitted to the peer's device. The system further employs user and peer databases to learn from prior interactions, thereby improving accuracy and facilitating effective communication for speech-impaired individuals.
Owner:MEZMO CORP

Speech recognition based sentence correction method and apparatus, device, and storage medium

The present application relates to the technical field of sentence correction, and particularly relates to a sentence correction method and device based on speech recognition, equipment and a storage medium, wherein a comprehensive bundle search technology is adopted to recognize speech data, and a first candidate sentence with the highest bundle search score and a plurality of candidate sentences are extracted in terms of semantic features and pinyin features, and feature fusion is performed; the first candidate sentence is corrected by using the fused features after feature fusion, thereby reducing the negative influence caused by pronunciation problems of users and multi-pronunciation character problems of texts, and improving the accuracy and efficiency of sentence correction.
Owner:GUANGZHOU YANLI NETWORK TECH CO LTD +3

Artificial intelligence-based vocal music training method and system

The present application relates to the technical field of voice analysis, and particularly discloses a vocal music training method and system based on artificial intelligence, which collects and analyzes body posture parameters of a target vocal music training person, detects posture abnormalities after preprocessing, forms a first abnormal training set, simultaneously captures and analyzes pronunciation feature parameters, identifies pronunciation abnormalities, forms a second abnormal training set, comprehensively generates an abnormal parameter training total set, and matches a correction set, and the abnormal degree and causes of body posture and the abnormal degree and causes of pronunciation are visually displayed, so that the target vocal music training person can directly see his / her own posture and pronunciation problems, and is facilitated to make self-adjustment and correction.
Owner:SHANGLUO UNIV

Personalized synthesis and recognition enhancement of dysarthric speech

ActiveCN120412540BSpeech synthesisDysarthric speechSpeech disorder
This invention discloses a personalized synthesis and recognition enhancement method for speech disorders. The speech disorder synthesis model includes a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module. The input of the speech disorder synthesis model includes samples, and the output includes synthesized speech disorder speech. The samples are speech disorder text sequences. The input of the long-range dependent feature encoding module includes samples, and the output is an alignment vector z. The input of the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The input of the decoding module is the final embedding representation, and the output is the synthesized speech disorder speech. The speech disorder synthesis model of this invention improves the ability to extract personalized features of speech disorder speech, enhances speech synthesis performance, and improves the fine-grained expression of speech disorder speech features.
Owner:TIANJIN UNIV

Session track division method and device, electronic equipment and storage medium

ActiveCN121545506ASpeech recognitionMedical equipmentSpeech disorderNoise
The invention provides a session track division method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining audio and video data recorded when a plurality of objects speak, the plurality of objects including a target object with dysarthria; processing audio data in the audio and video data to obtain text data and voice embedded data, wherein the text data carries voice composition state information; processing video data in the audio and video data to obtain video feature data including a time sequence speaking probability corresponding to each object; fusing the text data, the voice embedded data and the video feature data to obtain multi-modal fusion feature data; and carrying out speaker track division processing based on the multi-modal fusion feature data, and determining a track division result comprising the speaking information corresponding to each object. According to the method and the device, the accuracy and the anti-interference capability of multi-person session track division can be improved, and the judgment defect of a single mode in noise, overlapped voice and other scenes can be avoided.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +1

Speech training system for persons with speech disorders, speech training method for persons with speech disorders, communication support system for persons with speech disorders, communication support method for persons with speech disorders, analysis system for persons with speech disorders, analysis method for persons with speech disorders, program and recording medium

PendingJP2026084736AHealthcare managementReadingSpeech trainingSpeech rate
This system provides speech therapy for individuals with speech disorders, including children with disabilities, enabling them to easily and independently continue their training at home or elsewhere, without time or location constraints, even after discharge from the hospital or during outpatient visits. [Solution] The speech training support system for persons with speech disorders is configured to present a model voice converted to a slower speed using speech rate conversion technology to the person with a speech disorder who is the target of the speech training support, and then use speech rate conversion technology to convert the voice spoken by the person with a speech disorder, who imitates the model voice converted to a slower speed, back to the same speed as the model voice before conversion, and then present the high-speed converted voice to the person with a speech disorder or to the listener. By presenting the high-speed converted voice to the person with a speech disorder, training is conducted that focuses attention on the accuracy of articulation movements. The speed of the model voice converted to a slower speed is the fastest speed at which the person with a speech disorder can speak with accurate articulation without difficulty, and is between 1 / 2 and 1 times the speed of the model voice before conversion.
Owner:THE UNIV OF TOKYO +1

A method and device for evaluating multi-modal speech ability based on generative artificial intelligence

This invention provides a multimodal speech ability assessment method and apparatus based on generative artificial intelligence. It efficiently processes multimodal data through an asynchronous parallel mechanism and employs specialized techniques for in-depth analysis of different modal characteristics: for speech manuscripts, it utilizes structured cue words to guide a large language model, achieving multi-dimensional and standardized semantic assessment of text quality; for presentation slides, it applies vector retrieval-enhanced generation technology to accurately locate core content and assess its structure and design; for audio data, it calls professional interfaces for streaming analysis to obtain overall dimensional scores and word-level diagnostics that pinpoint specific word pronunciation problems; for video data, it identifies the speaker's emotional state through a multimodal model that integrates spatiotemporal, audio, and text features. The standardized integration and unified display of the assessment results from each modality achieves a comprehensive and in-depth assessment of speech ability, significantly improving assessment efficiency and practical teaching value.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Methods and systems for customized multimedia sessions and treatments of speech disorders using customized multimedia sessions

PendingUS20260080801A1SensorsDiagnostic recording/measuringSpeech disorderUser device
A system and method for providing customized interactive multimedia sessions using machine learning. A method includes obtaining a transcript for multimedia content. The transcript is analyzed using a machine learning architecture in order to generate questions and corresponding expected answers for the multimedia content. The questions are provided to a user device. Responses to the questions may be received and analyzed in order to analyze user performance. Some techniques described include methods for treating speech disorders using customized interactive multimedia sessions.
Owner:SPEECHBUDDY LLC