Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Speech disorder" patented technology

A communication disorder in which normal speech is impaired.

Apparatus for the self-administration of therapies for the treatment of speech disorders

PCT designated stageWO2026028021A1Psychotechnic devicesSensorsAuditory stimuliSound sources
The present invention relates to an apparatus (10) for the self-administration of therapies for the treatment of speech disorders comprising a stimulus delivery device (11,12) with a horizontal development substantially curved along an angular portion so as to outline a portion of a circle, comprising at least one panel (11) for the delivery of visual stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); at least one acoustic source (12) for the delivery of acoustic stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); and a local electronic processing unit configured to control the generation of visual and acoustic stimuli by means of the at least one panel (11) and the at least one acoustic source (12), and characterized in that it comprises at least one first acoustic detector (15) for detecting a reproduction of a verbal element produced by the patient, wherein the at least one first acoustic detector (15) is carried by the at least one panel (11) and positioned so as to detect acoustic signals generated in a space inside the portion of circle outlined by the stimulus delivery device (11,12).
Owner:LINARI MEDICAL SRL

A brain-computer interface system for recognizing the intention of Chinese oral language based on a sound-meaning integration double model

ActiveCN121560160BSpoken languageStereotaxis
The application provides a Chinese spoken language intention recognition brain-computer interface system based on a sound-meaning integration double model, belongs to the technical field of biomedical engineering, and relates to language brain-computer interface technology. Taking sound-meaning integration as the core, the stereotactic intracranial electroencephalogram (sEEG) technology is adopted to collect neural signals of the brain articulatory motor coding area and the semantic concept organization coding area. The system comprises a voice initiation decoder, a speech decoder, a semantic decoder and a Chinese word speech-semantic fusion synthesizer, the target decoder is constructed by extracting high gamma band features of key brain areas of the frontal lobe (left inferior frontal gyrus, premotor cortex, etc.), the temporal lobe (anterior temporal lobe, dorsolateral temporal lobe, etc.). At the same time, a visual and auditory induction training paradigm is matched, three tasks of listening to sound to group words, looking at words to group words and word association are set, and the subjects are supported to generate words independently. The system effectively solves the homonym and near homonym word ambiguity problem in Chinese spoken language recognition, and provides a precise interactive tool for ALS and other speech disorder patients.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models

The invention provides a Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models, belongs to the technical field of biomedical engineering, and relates to a language brain-computer interface technology. Sound-sense integration is taken as a core, and a stereotactic intracranial electroencephalogram (sEEG) technology is adopted to acquire neural signals of a brain phonetic motion coding region and a semantic concept organization coding region. The system comprises a sound production starting decoder, a voice decoder, a semantic decoder and a Chinese word voice-semantic fusion synthesizer, and a target decoder is constructed by extracting high gamma wave band characteristics of key brain regions of frontal lobe (left subfrontal gyrus, anterior cortex of motion and the like) and temporal lobe (anterior temporal lobe, dorsal lateral temporal lobe and the like). Meanwhile, an audio-visual induction training normal form is matched, three tasks of listening word combination, character reading word combination and vocabulary association are set, and subjects are supported to autonomously generate vocabularies. The system effectively solves the problem of ambiguity of homophonous and near-phonetic words in spoken Chinese recognition, and provides an accurate interaction tool for speech disorder patients such as ALS and the like.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Speech disorder detection method, device, equipment and readable storage medium

ActiveCN120318639BCharacter and pattern recognitionSensorsSpeech disorderSpeech sounds
The present disclosure relates to a speech disorder detection method, device, equipment and readable storage medium. By acquiring standard audio-visual materials, in response to the pronunciation operation of the to-be-detected object to the standard audio-visual materials, multi-modal pronunciation data is collected, audio acoustic features are extracted based on the pronunciation audio, video visual features are extracted based on the video of the face and oral cavity activity, the audio acoustic features, the video visual features and the demographic information coding data are fused to obtain a fusion feature vector, and based on the fusion feature vector and a pre-trained prediction model, a speech disorder detection result of the to-be-detected object is obtained. Compared with the prior art, the embodiment of the present disclosure can improve the accuracy and comprehensiveness of speech disorder detection, improve the diagnosis efficiency, reduce the dependence on professionals, reduce the burden of medical resources, and clearly determine the specific type of pronunciation problem, thereby providing a scientific basis for subsequent individualized intervention treatment.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Information processing device, information processing method, and program

[Problem] To realize a technique capable of improving quality of life for a subject such as a patient who is impaired in vocalization. [Solution] An information processing device 1 includes a moving image data acquisition unit 11a, a preprocessing unit 11b, a feature amount acquisition unit 11c, a learning model generation unit 11d, an utterance content estimation unit 11e, and a substitute voice output unit 11f. The moving image data acquisition unit 11a acquires a moving image including lips of a subject. The preprocessing unit 11b extracts a lip portion in the moving image. The feature amount acquisition unit 11c acquires a feature amount of the extracted lip portion. The utterance content estimation unit 11e estimates an utterance content of the subject corresponding to the lip portion extracted by the preprocessing unit 11b, on the basis of a learning model that has been machine-learned in advance by associating a feature amount of a lip portion in a moving image for machine learning including the movement of lips during utterance with the utterance content uttered by the lip portion.
Owner:KEIO UNIV

Captioned telephone service system for user with speech disorder

ActiveUS12592993B2Special service for subscribersSpeech recognitionSpeech ProcessorSpeech disturbances
A captioned telephone service (CTS) system provides a transcription service and a speech-to-text and text-to-speech converting service for the deaf or hard-of-hearing user with a speech disorder during a phone call between the user and the peer. The CTS system transcribes the peer's voice into text to be displayed on the user's device. The CTS system transcribes the user's voice into text based on the database storing user's spoken audios and corresponding texts, and converts the text into a clear and articulate speech to be sent to the peer's device instead of the user's voice in order to help the peer better understand what the user said. The CTS system includes a speech-to-text handler and a text-to-speech handler. The speech-to-text handler transcribes user's voice into text using the database, and the text-to-speech handler converts the text into a clear and articulate speech.
Owner:MEZMO CORP

Empathy-Based Wearable Systems for Perceptual and Communicative Assistance

PendingUS20260179644A1Speech recognitionHeadphonesEmpathy
The invention provides a unified assistive technology system, the ADA Empathy Wearable Bundle, enhancing perception, communication, and self-expression for individuals with visual, auditory, or speech impairments.It integrates Empathy Glasses, Empathy Headphones, Word Articulation Guidance (WAGS), Word Articulation Response Monitoring (W.A.R.M.), and Conversational MIDI (C-MIDI), using AI-driven narrative interpretation, augmented reality, and blockchain-based security.The system delivers vivid narrative audio for blind users, multimodal AR overlays for hearing-impaired users, and real-time visual and auditory articulation guidance for speech-impaired or non-verbal users. A tone-based trust protocol ensures operational integrity, while blockchain storage preserves privacy and consent.Grounded in ethical AI, empathy, and human dignity, the system enables users to perceive, understand, and engage with the world fully, offering a holistic, compassionate, and transformative solution in assistive technology.
Owner:BOWES DEAN

Speech disorder assessment method and system based on multi-modal large model

The application discloses a speech disorder evaluation method and system based on a multi-modal large model. The system comprises: acquiring corresponding multi-modal data for a target, the multi-modal data including audio signals, lip videos and tongue ultrasonic images; performing cross-modal feature extraction and semantic alignment on the multi-modal data to map the extracted modal features to a unified semantic space aligned with a large model text embedding space, obtaining a semantic vector sequence; using the semantic vector sequence as input, simulating clinical multi-level reasoning logic using the large model to obtain a preliminary evaluation result of articulation disorder, the preliminary evaluation result including severity level and disorder type; using key information in the preliminary evaluation result to set a retrieval query strategy to guide the large model to generate an evaluation report and rehabilitation suggestions. The application improves the accuracy, real-time performance and robustness of speech disorder evaluation.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Captioned telephone service system for user with speech disorder

A captioned telephone service system is provided for assisting users with speech disorders in communicating with peers. The system comprises a CTS server, a user device, and a user application configured with a sentence refinement unit and a text-to-speech handler. The user inputs one or more words, which are processed by the sentence refinement unit to correct errors, insert missing grammar, and generate grammatically complete sentences. A contextual adaptation module analyzes conversation and peer information to select the most relevant sentence. The user confirms, regenerates, or edits the sentence, which is then converted to natural speech by the text-to-speech handler and transmitted to the peer's device. The system further employs user and peer databases to learn from prior interactions, thereby improving accuracy and facilitating effective communication for speech-impaired individuals.
Owner:MEZMO CORP

Adaptive audio and audiovisual recursive self-feedback for speech therapy

Systems and methods are provided for generating, managing, adapting, and delivering speech therapy to a user via a mobile device in an at-home or out-of-clinic setting. Users may be persons experiencing aphasia or other speech conditions. For example, prompts may be communicated to a user, and their spoken responses recorded and analyzed. Based on analysis of the responses, these systems and methods may determine a mode of response including but not limited to playback of the user's spoken response to allow the user to recursively self-assess and / or self-correct. User performance and improvement trends may be assessed, and utilized to adapt future therapy sessions.
Owner:UNIV OF SOUTH FLORIDA

Personalized synthesis and recognition enhancement of dysarthric speech

ActiveCN120412540BSpeech synthesisDysarthric speechSpeech disorder
This invention discloses a personalized synthesis and recognition enhancement method for speech disorders. The speech disorder synthesis model includes a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module. The input of the speech disorder synthesis model includes samples, and the output includes synthesized speech disorder speech. The samples are speech disorder text sequences. The input of the long-range dependent feature encoding module includes samples, and the output is an alignment vector z. The input of the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The input of the decoding module is the final embedding representation, and the output is the synthesized speech disorder speech. The speech disorder synthesis model of this invention improves the ability to extract personalized features of speech disorder speech, enhances speech synthesis performance, and improves the fine-grained expression of speech disorder speech features.
Owner:TIANJIN UNIV

Compounds and Methods for Modulating UBE3A-ATS

PendingUS20260132402A9Sugar derivativesScreening processSpeech disorderUbiquitin-Protein Ligases
Provided are compounds, methods, and pharmaceutical compositions for reducing the amount or activity of UBE3A-ATS, the endogenous antisense transcript of ubiquitin protein ligase E3A (UBE3A) in a cell or subject, and in certain instances increasing the expression of paternal UBE3A and the amount of UBE3A protein in a cell or subject. Such compounds, methods, and pharmaceutical compositions are useful to ameliorate at least one symptom or hallmark of a neurogenetic disorder. Such symptoms and hallmarks include developmental delays, ataxia, speech impairment, sleep problems, seizures, and EEG abnormalities. Such neurogenetic disorders include Angelman Syndrome.
Owner:IONIS PHARMACEUTICALS INC

Adaptive audio and audiovisual recursive self-feedback for speech therapy

PendingUS20260188132A1Spoken languageSpeech disorder
Systems and methods are provided for generating, managing, adapting, and delivering speech therapy to a user via a mobile device in an at-home or out-of-clinic setting. Users may be persons experiencing aphasia or other speech conditions. For example, prompts may be communicated to a user, and their spoken responses recorded and analyzed. Based on analysis of the responses, these systems and methods may determine a mode of response including but not limited to playback of the user's spoken response to allow the user to recursively self-assess and / or self-correct. User performance and improvement trends may be assessed, and utilized to adapt future therapy sessions.
Owner:UNIV OF SOUTH FLORIDA

Session track division method and device, electronic equipment and storage medium

ActiveCN121545506ASpeech recognitionMedical equipmentSpeech disorderNoise
The invention provides a session track division method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining audio and video data recorded when a plurality of objects speak, the plurality of objects including a target object with dysarthria; processing audio data in the audio and video data to obtain text data and voice embedded data, wherein the text data carries voice composition state information; processing video data in the audio and video data to obtain video feature data including a time sequence speaking probability corresponding to each object; fusing the text data, the voice embedded data and the video feature data to obtain multi-modal fusion feature data; and carrying out speaker track division processing based on the multi-modal fusion feature data, and determining a track division result comprising the speaking information corresponding to each object. According to the method and the device, the accuracy and the anti-interference capability of multi-person session track division can be improved, and the judgment defect of a single mode in noise, overlapped voice and other scenes can be avoided.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +1

Speech training system for persons with speech disorders, speech training method for persons with speech disorders, communication support system for persons with speech disorders, communication support method for persons with speech disorders, analysis system for persons with speech disorders, analysis method for persons with speech disorders, program and recording medium

PendingJP2026084736AHealthcare managementReadingSpeech trainingSpeech rate
This system provides speech therapy for individuals with speech disorders, including children with disabilities, enabling them to easily and independently continue their training at home or elsewhere, without time or location constraints, even after discharge from the hospital or during outpatient visits. [Solution] The speech training support system for persons with speech disorders is configured to present a model voice converted to a slower speed using speech rate conversion technology to the person with a speech disorder who is the target of the speech training support, and then use speech rate conversion technology to convert the voice spoken by the person with a speech disorder, who imitates the model voice converted to a slower speed, back to the same speed as the model voice before conversion, and then present the high-speed converted voice to the person with a speech disorder or to the listener. By presenting the high-speed converted voice to the person with a speech disorder, training is conducted that focuses attention on the accuracy of articulation movements. The speed of the model voice converted to a slower speed is the fastest speed at which the person with a speech disorder can speak with accurate articulation without difficulty, and is between 1 / 2 and 1 times the speed of the model voice before conversion.
Owner:THE UNIV OF TOKYO +1

Methods and systems for customized multimedia sessions and treatments of speech disorders using customized multimedia sessions

PendingUS20260080801A1SensorsDiagnostic recording/measuringSpeech disorderUser device
A system and method for providing customized interactive multimedia sessions using machine learning. A method includes obtaining a transcript for multimedia content. The transcript is analyzed using a machine learning architecture in order to generate questions and corresponding expected answers for the multimedia content. The questions are provided to a user device. Responses to the questions may be received and analyzed in order to analyze user performance. Some techniques described include methods for treating speech disorders using customized interactive multimedia sessions.
Owner:SPEECHBUDDY LLC

Gesture Vox

GestureVox is an innovative AI-powered software system designed to convert sign language into spoken words in real-time. Utilizing advanced machine learning techniques, including frameworks such as TensorFlow, PyTorch, Keras, and Scikit-learn, GestureVox offers a seamless and accurate gesture recognition and speech synthesis process. The system's architecture includes modules for data collection, pre-processing, model training, testing, hyperparameter tuning, and deployment. Key features include the ability to process live video feeds, a user-friendly interface, and scalability to handle a large number of concurrent users, potentially utilizing cloud services such as AWS, Azure, and Google Cloud. GestureVox significantly enhances communication for individuals with speech impairments, providing an inclusive and accessible solution.
Owner:SELVAM HARIVATSAN