Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27 results about "Dysarthria" patented technology

Improper speech where words run into one another and cannot be distinctly recognized.

Speech recognition for assisting patients with speech difficulties

ActiveUS12488786B1Speech recognitionDysarthriaAcoustics
Disclosed herein are novel speech recognition methods and systems for assisting users or patients that have speech difficulties (e.g., as a symptom of one or more disorders or conditions). Specifically disclosed are speech recognition methods and systems that enable dysarthria patients to communicate more clearly and effectively.
Owner:ARTIK LLC

Automated recommendation tool to improve intelligiblity in speech dysarthria

A computing device to improve intelligibility in an individual with speech dysarthria. The computing device accesses recorded speech of an individual with a speech dysarthria condition. The computing device accesses a pre-trained language improvement model. The pre-trained language improvement model identifies one or more exercises to be performed by the individual with speech dysarthria to improve intelligibility of the individual, the one or more exercises identified by the pre-trained language improvement model based upon an analysis of the recorded speech by the pre-trained language improvement model. A digital interface presents the one or more exercises to be performed by the individual.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Method and device for evaluating dysarthria

The present disclosure relates to a method which a server receives voice data of a user from a user terminal and evaluates dysarthria of the user, the server including a processor and a database and communicates with the terminal, and the terminal including a processor, a memory, a display, and a microphone and communicating with the server. The method comprises steps in which the server: causing the terminal to describe a dysarthria evaluation process to the user; causing the user terminal to check whether the surrounding environment of the user is suitable to evaluate the voice of the user; causing the user terminal to induce the user to perform at least one from among sustained phonation, articulatory diadochokinesis, word reading, and sentence reading for sustained phonation evaluation, and record the voice of the user; and evaluating degrees of dysarthria of the user on the basis of the recorded voice data.
Owner:HAII CORP

Phoneme recognition model training method, phoneme recognition method and phoneme recognition device

PendingCN121506109ASpeech recognitionPhoneme recognitionDysarthria
The invention relates to the technical field of speech recognition, and discloses a phoneme recognition model training method, a phoneme recognition method and a phoneme recognition device.A first sample phoneme sequence corresponding to normal speech data is recognized through a first phoneme recognition model, and a second sample phoneme sequence corresponding to dysarthria speech data is recognized through a second phoneme recognition model; and taking the second sample phoneme sequence as a reference phoneme sequence in a scene without a fixed text. According to the method, fine tuning training is performed on the basis of the first phoneme recognition model according to the individual dysarthria data of the dysarthria patient to obtain the individualized second phoneme recognition model, and the situation of free dialogue or poor text compliance can be covered even in a scene without a fixed text, so that the accuracy of recognition of the dysarthria is improved. The actual pronunciation ability of the patient can be reflected in a natural communication context, dependence on fixed texts and artificial phoneme level alignment labeling is avoided, higher flexibility and adaptability are achieved, and a reliable basis can be provided for rehabilitation training and curative effect tracking.
Owner:SHENZHEN UNIV +1

Automatic dysarthria assessment method and system based on visual speech motor features

The application discloses an automatic dysarthria evaluation method and system based on visual speech movement characteristics. Video stream data when expressing multiple unit sounds is acquired; after normalizing the human face image in each frame image, human face feature point calibration is performed; according to the time sequence of the video stream data and the calibrated human face feature points, the movement trajectory sequence of the lips and the lower jaw in the video stream data is extracted; according to the movement trajectory sequence, the distance from each detection point to the fixed point is calculated according to the preset detection point and the fixed point, and a distance sequence is formed; the formed distance sequence is input into the trained automatic dysarthria evaluation model, speech construction probability prediction is performed, and a dysarthria evaluation result is obtained. The application adopts more objective, comprehensive and detailed automatic evaluation indexes to assist the diagnosis and evaluation work of doctors, and provides a reliable and convenient method for subsequent speech therapists to formulate further speech rehabilitation treatment schemes, disease monitoring and the like for dysarthria patients.
Owner:TIANJIN UNIV

Ataxia dysarthria voice airflow synchronous analysis and training system and method

The invention discloses an ataxia dysarthria voice and airflow synchronous analysis and training system and method, and belongs to the technical field of medical rehabilitation, and the system comprises a data collection module, a signal preprocessing module, a voice and airflow synchronization analysis module, a personalized training generation module and a training feedback and evaluation module. Based on a differential geometry principle, a Riemannian manifold theory is innovatively introduced, voice and air flow signals are respectively mapped to different Riemannian manifold spaces, accurate time sequence alignment of the voice-air flow signals is realized by constructing a geodesic flow field, a double-flow network structure is designed to extract voice-air flow synchronization features, and the accuracy of voice-air flow synchronization is improved. According to the method, an accurate dysarthria evaluation result is generated, the system generates a personalized training scheme according to the evaluation result, a training closed loop is formed through real-time feedback and dynamic adjustment, accurate evaluation and effective training of ataxia dysarthria are achieved, and compared with a traditional method, the accuracy of dysarthria analysis is improved by about 35%, and the training efficiency is improved by 30%.
Owner:ZHEJIANG PROVINCIAL PEOPLES HOSPITAL

Speech synthesis feedback method and system for dysarthria correction

PendingCN121905141ASpeech synthesisDysarthriaAcoustics
The invention provides a speech synthesis feedback method and a speech synthesis feedback system for dysarthria correction, which are applied to the technical field of speech synthesis, and are used for determining a simplified feedback strategy currently applied to a target patient and constructing a corresponding dynamic reference basis by acquiring a speech signal and speech motion data of the target patient. And performing comparative analysis on the basis of the voice signal, the speech motion data and the dynamic reference basis to determine the pronunciation adaptation state of the target object, and generating synthetic voice feedback according to the pronunciation adaptation state, especially when the pronunciation adaptation state indicates that the target object is over-adapted to the simplified feedback strategy. The generation parameters of the synthetic voice feedback are adjusted to introduce acoustic variation features tending to natural pronunciation, so that over-adaptation of a voice composition motion mode can be effectively avoided, and it is ensured that acoustic simplification does not interfere or distort accurate perception of a patient on key acoustic features of target phonemes. The beneficial effect of stable and efficient transition from simplification to natural pronunciation is realized.
Owner:THE FIRST AFFILIATED HOSPITAL OF SHANTOU UNIV MEDICAL COLLEGE

A stroke dysarthria speech feature selection method based on dynamic weight compensation

This invention discloses a method for selecting speech features in stroke-related dysarthria based on dynamic weight compensation. First, the patient's acoustic features are acquired and initial feature weights are calculated, dividing the data into high-weight and low-weight feature candidate sets. Second, outliers in high-weight features are identified using a dual-threshold outlier detection algorithm, and a dynamic weight compensation strategy is employed to correct weight biases while preserving potential pathological information, resulting in an optimized compensated feature dataset. Then, a training feature dataset is constructed using the low-weight candidate set, input into a pre-trained temporal dependency validation model for analysis, and weights are updated based on the feature's contribution to the classification results to obtain the final feature weights. Finally, features are classified according to these weights, and an adaptive truncation strategy is used to select the final feature subset. This invention constructs a two-layer selection mechanism from coarse to fine screening, effectively extracting the core acoustic features most relevant to pathology, improving feature discriminative power, interpretability, and model generalization ability, and providing high-quality feature input for assisted diagnosis and assessment.
Owner:DALIAN NEUSOFT UNIV OF INFORMATION

Speech disorder assessment method and system based on multi-modal large model

The application discloses a speech disorder evaluation method and system based on a multi-modal large model. The system comprises: acquiring corresponding multi-modal data for a target, the multi-modal data including audio signals, lip videos and tongue ultrasonic images; performing cross-modal feature extraction and semantic alignment on the multi-modal data to map the extracted modal features to a unified semantic space aligned with a large model text embedding space, obtaining a semantic vector sequence; using the semantic vector sequence as input, simulating clinical multi-level reasoning logic using the large model to obtain a preliminary evaluation result of articulation disorder, the preliminary evaluation result including severity level and disorder type; using key information in the preliminary evaluation result to set a retrieval query strategy to guide the large model to generate an evaluation report and rehabilitation suggestions. The application improves the accuracy, real-time performance and robustness of speech disorder evaluation.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Dysarthria detection method, dysarthria detection device, and recording medium

A dysarthria detection method includes an obtaining step and a detecting step. In the obtaining step, voice information regarding voice uttered by a subject is obtained. In the detecting step, it is detected whether the subject has dysarthria, based on an output result obtained by inputting the voice information obtained in the obtaining step into a detection model. The detection model has been trained by machine learning to output information regarding whether the subject has dysarthria by using voice information inputted.
Owner:PANASONIC HOLDINGS CORP

Parkinson's disease dysarthria early recognition method based on voiceprint features

ActiveCN121545557ASpeech analysisDysarthriaMedical diagnosis
The invention discloses a Parkinson's disease dysarthria early recognition method based on voiceprint features, and belongs to the technical field of medical diagnosis. The method comprises the following steps: performing voice analysis and voice recognition on original voice data of a to-be-recognized object, extracting pronunciation deviation and an initial feature set related to voiceprint in a fragment, screening out target features with remarkable discrimination degree for recognition of dysarthria of Parkinson's disease, and obtaining a dysarthria recognition result of the to-be-recognized object through a model. Through accurate vowel and consonant fragment extraction and pronunciation deviation recognition, in combination with association of a character sequence and voice data, accurate interception of phoneme fragments is realized, quantitative judgment of pronunciation deviation is realized, key dimensions of dysarthria recognition are reserved through feature extraction, early signals of dysarthria of Parkinson's disease are captured, and the recognition accuracy of dysarthria of Parkinson's disease is improved. High-risk groups are marked, and an early warning result is calculated and output in combination with a joint risk value, so that the problem of unobvious symptoms in a prosperous period is solved, a review interval is dynamically adjusted, and whole-process management from identification to monitoring is realized.
Owner:SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL

Method for early recognition of parkinsonian dysarthria based on voiceprint features

ActiveCN121545557BDysarthriaMedical diagnosis
The application discloses a Parkinson disease dysarthria early identification method based on voiceprint features and belongs to the technical field of medical diagnosis. The original speech data of a to-be-identified object is subjected to speech analysis and speech recognition, and an initial feature set related to voiceprints in a pronunciation deviation and a segment is extracted, target features with significant discriminability for Parkinson disease dysarthria identification are screened out, a dysarthria identification result of the to-be-identified object is obtained through a model, accurate vowel and consonant segment extraction and pronunciation deviation identification are realized, phoneme segment accurate cutting is realized by combining the association of a text sequence and speech data, the quantization determination of the pronunciation deviation is realized, the key dimension of the dysarthria identification is reserved through feature extraction, the early signals of Parkinson disease dysarthria are captured, the high-risk groups are marked, the early warning result is output in combination with the joint risk value calculation, the problem that the prodromal symptoms are not obvious is solved, the review interval is dynamically adjusted, and the whole-process management from identification to monitoring is realized.
Owner:SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL

A method and device for speech recognition of stroke dysarthria

ActiveCN119049517BSpeech recognitionSyllableDysarthria
The present application discloses a method and device for speech recognition of dysarthria caused by stroke. The technical solution of the present application classifies the acquired speech sample data according to the syllable category represented by the audio and the normal and patient categories, and obtains the audio spectrogram through transformation. Then, in the stage of building the network model, the front-end network uses the core processing module designed by plant morphology and physiology, and forms the rhizome with continuous Downsample after the STEM module to quickly transport the computing nodes to a higher receptive field area. The back-end network is based on the alternating configuration of the deep separable convolution and attention mechanism based on the Xception module to form a vine cross structure. The attention mechanism is selectively placed in the alternating convolution modules to improve the recognition ability and accuracy of key speech features, thereby being able to capture the global receptive field at multiple scales, accurately learn and discriminate the significant characteristic information of stroke pathology, and solve the technical problem of low accuracy in the existing articulation analysis of stroke patients.
Owner:GUANGDONG UNIV OF TECH

Non-invasive screening method and device for Alzheimer's disease proactive period based on expression features

The invention provides a non-invasive screening method and a non-invasive screening device for an Alzheimer's disease proactive period based on expression features. The non-invasive screening method comprises the following steps: S1, collecting facial video data of a user; s2, preprocessing the face video data; s3, subject extraction is carried out on the preprocessed video clips based on subject modeling, so that positive emotion clips are eliminated, and negative emotion clips and neutral emotion clips are reserved; s4, extracting a key frame from the reserved video clip; s5, extracting micro-expression features of the user from the key frame; s6, based on the key frame, analyzing lip motion features of the user to perform dysarthria assessment; s7, carrying out multi-modal feature fusion on the micro expression features and the dysarthria features; and S8, inputting the fused features into a pre-trained pAD risk assessment model to obtain a pAD risk assessment result. Noninvasive, low-cost and high-efficiency Alzheimer's disease prophase screening is realized for the first time, and a feasible technical path is provided for early screening of large-scale crowds.
Owner:张涛

Speech recognition method and system for dysarthria group

ActiveCN120656444BSpeech recognitionDysarthriaVoice transformation
The application relates to a speech recognition method and system for a dysarthria group, the method comprising: collecting dysarthria speech data, preprocessing the dysarthria speech data, and obtaining effective speech segments; inputting the effective speech segments into a dysarthria speech recognition model to obtain phoneme-level or Chinese character-level recognition results; the dysarthria speech recognition model is obtained by training a Conformer model using a first training set, the first training set comprising: fake dysarthria audio data; the fake dysarthria audio data is obtained by speech conversion based on a CycleGAN-VC speech conversion model; in the process of training the model in the first training set, the model parameters of the Conformer model are adjusted, and the Conformer model is optimized through a whale optimization algorithm. The application can improve the accuracy and robustness of dysarthria speech recognition.
Owner:BEIJING UNION UNIVERSITY

Large language model-based dysarthria voice real-time conversion system

The invention discloses a dysarthria voice real-time conversion system based on a large language model, and the system comprises a voice recognition module based on ASR, which employs a Whisper ASR model to convert the input voice of a dysarthria patient into an initial text; based on an LLM semantic correction module, integrating a Qwen2.5-7B-Instruct large language model, and performing semantic error correction and emotion enhancement on an initial text through a double-stage prompt engineering technology; the TTS-based speech synthesis module is used for converting the corrected text into natural speech by adopting a CosyVoice TTS model and outputting the natural speech; the real-time optimization module is used for controlling end-to-end delay to meet real-time factors through a dynamic voice buffering mechanism, an edge-cloud collaborative architecture and a model quantification technology; the personalized federated learning module is used for carrying out user self-adaptive fine tuning on the ASR model and the LLM model by adopting a LightFed-Cluster framework in combination with differential privacy protection; the semantic accuracy, the voice definition, the voice naturalness and the conversion time delay are greatly improved, and the method is more suitable for auxiliary and alternative communication of dysarthria patients.
Owner:SOUTHEAST UNIV

Device for assisting pronunciation of dysarthria and exercising temporomandibular joint

The utility model discloses an auxiliary pronunciation and temporomandibular joint exercise device for dysarthria, which comprises a first movable plate, a second movable plate, a third movable plate and a third movable plate, the second movable plate is provided with a second hinge piece, and the second hinge piece is hinged to the first hinge piece; the limiting inserting rod is provided with a fixing position and an adjusting position, and the limiting inserting rod at the fixing position is inserted into the clamping hole so that the limiting inserting rod and the first hinge piece can be fixed in the circumferential direction; and the elastic piece is used for applying elastic acting force to the limiting insertion rod. The included angle is adjusted through relative rotation of the second movable plate and the first movable plate, so that the height of the whole structure is adjusted to adapt to different calibers needing to be opened; through cooperation of the limiting insertion rod and the elastic piece, a user can easily switch between a fixing position and an adjusting position, and rapid adjustment is achieved; according to the device, the adaptability of control over different mandibular movements is improved, and the trouble caused by frequent replacement of a sounding rubber tube is avoided.
Owner:THE THIRD XIANGYA HOSPITAL OF CENT SOUTH UNIV

Method for evaluating malocclusion with concomitant articulatory disorders in children based on multi-modal analysis

PendingCN122658364ADysarthriaModal analysis
This invention discloses a method for assessing articulation disorders associated with malocclusion in children based on multimodal analysis, belonging to the field of multimodal analysis technology. The method includes: Step 1, identifying a subset of malocclusion-sensitive phonemes; simultaneously acquiring single-channel audio signals, frontal view color video of the mouth, and lateral view color video of the mouth while the child completes a guided articulation task sequence consisting of words corresponding to the subset of malocclusion-sensitive phonemes; obtaining audio-visual segments of each malocclusion-sensitive phoneme through phoneme-level forced alignment and synchronous slicing; Step 2, obtaining labial, lingual, and dental key points from the audio-visual segments of the articulation segments through oral key point detection; performing specific acoustic and visual feature extraction according to the key constraints of the articulatory organs for each malocclusion-sensitive phoneme; and aligning and splicing them along the time axis to form a bimodal malocclusion-sensitive feature pair for the corresponding malocclusion-sensitive phoneme; Step 3, configuring a phoneme-specific expert network for each malocclusion-sensitive phoneme; and outputting auxiliary assessment results after compensatory coupling channels and aggregation according to the type of oral and maxillofacial structural limitation. This invention improves the relevance, robustness, and interpretability of assessing articulation disorders associated with malocclusion in children.
Owner:BEIJING SHANMAI TECHNOLOGY CO LTD

Real-time dysarthria speech restoration method based on progressive model distillation

The invention discloses a real-time dysarthria speech restoration method based on progressive model distillation, and belongs to the technical field of speech signal processing. The invention provides a combined solution for the problem of double interference of voice degradation and environmental noise faced by disabled old people in a complex nursing environment. Firstly, a pseudo-parallel corpus is constructed through a self-supervised repair strategy, an ideal target speech is generated by using a high-precision speech conversion model, and the problem of truth value missing in a pathological speech enhancement task is solved. Secondly, constructing a complex convolutional loop network based on voiceprint embedding, and introducing a speaker embedding vector to suppress non-target human voice interference; meanwhile, a gradual receptive field perception distillation strategy is adopted, a non-causal wide convolution teacher model is compressed into a causal narrow convolution student model through micro sparsification, and real-time processing of an embedded terminal is realized. And finally, phoneme perception loss is introduced, and semantic level supervision is performed by using the pre-trained ASR model, so that the intelligibility of the repaired voice is remarkably improved.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Large model database AI phonetic symbol phonetic notation system based on speech recognition

InactiveCN121528250ASpeech analysisDecision networksDysarthria
The invention relates to the technical field of voice signal processing and rehabilitation medicine, and discloses a large model database AI phonetic symbol phonetic notation system based on voice recognition. The method comprises six steps of pathological feature recognition and database matching, multi-task collaborative modeling and shared feature construction, pathological voice rhythm reconstruction processing, collaborative decision optimization and conflict detection, intelligent analysis and collaborative feedback scoring, and personalized strategy making and database updating. Precise evaluation of pathological speech such as aphasia, dysarthria and speech development retardation is realized, and a pathological speech rhythm reconstruction network is adopted to identify and repair rhythm abnormity. Constructing a multi-task collaborative decision network to realize multi-dimensional collaborative evaluation of pronunciation accuracy, rhythm recovery, emotion expression, cognitive level and the like; comprehensively processing an evaluation result by applying a rehabilitation collaborative feedback algorithm and dynamically adjusting a strategy weight; according to the method, the accuracy and systematicness of pathological voice rehabilitation evaluation can be remarkably improved, personalized rehabilitation guidance is provided for patients, and the method has wide application prospects in the fields of voice rehabilitation treatment, language pathological diagnosis, special education and the like.
Owner:SHENZHEN BESTWAY CLOUD INTELLIGENCE TECHNOLOGY CO LTD

A dysarthria speech recognition system and method based on retrieval-augmented learning

PendingCN122637767ADysarthriaFeature data
The application discloses a dysarthria speech recognition system and method based on retrieval reinforcement learning, and the method comprises the following steps: firstly, performing front-end feature processing on the spectrogram of dysarthria speech; then inputting a retrieval mapping encoder, performing hierarchical distribution alignment by using a loss function, and utilizing hierarchical weights decreasing with depth to retain acoustic features in shallow weak alignment and extract semantic features in deep strong alignment, so as to map distorted features to a non-dysarthria speech manifold; then, based on a pre-constructed non-dysarthria speech feature database, searching for the most similar non-dysarthria speech features by using Euclidean distance; finally, inputting a retrieval reinforcement decoder, taking self-attention output as a query vector and retrieval embedding as a key-value vector, performing feature correction and decoding by using multi-head cross-attention, and outputting recognized text. The application effectively corrects acoustic distortion while retaining pathological features, and significantly improves the recognition accuracy of dysarthria speech.
Owner:GUANGDONG UNIV OF TECH

Chinese multi-dialect-oriented dysarthria detection and grading method

PendingCN121884868ASpeech recognitionDysarthriaSemantic feature
The invention discloses a Chinese multi-dialect-oriented dysarthria detection and grading method, and belongs to the field of speech processing. According to the method, detection and severity grading of dysarthria are synchronously realized through a multi-task deep learning framework. The collected count-off voice is preprocessed, the sampling rate is unified, and the voice is segmented into short segments; a Chinese optimized wav2vec2.0 model is adopted to extract deep semantic features, and key acoustic features such as pronunciation fatigue, tone stability and formant dynamics are extracted in combination with a specially designed predefined feature module; thirdly, a confusion perception attention mechanism is introduced, feature interference caused by dialect variation and dysarthria pronunciation similarity is inhibited, and the generalization ability of the model in different regional crowds is improved; and synchronously outputting health / illness judgment and 0-3-level severity rating through a multi-task classification network in combination with a dynamic weighted loss function. The objectivity, accuracy and practicability of dysarthria detection are improved.
Owner:BEIJING UNIV OF TECH

Snn-based post-stroke dysarthria pathological speech recognition method and device

ActiveCN119601043BSpeech analysisSyllableDysarthria
The application discloses a post-stroke dysarthria pathological speech recognition method and device based on SNN, and aims to solve the technical problem that the existing dysarthria pathological speech recognition method based on a deep learning algorithm causes excessive power consumption in the model operation process. The method comprises the following steps: obtaining a plurality of initial audio data of a patient to be detected, and respectively pre-processing each initial audio data to output a mel-frequency cepstral coefficient feature corresponding to each initial audio data; inputting each mel-frequency cepstral coefficient feature into a pulse neural network that fuses word attention and channel attention for recognition, and outputting a plurality of syllable category disease probability; and generating a post-stroke dysarthria pathological speech recognition result corresponding to the patient to be detected according to the plurality of syllable category disease probability.
Owner:GUANGDONG UNIV OF TECH

Practice device for assisting exercise dysarthria patient

ActiveCN223336710UGymnastic exercisingDysarthriaApparatus instruments
The utility model relates to the technical field of medical instruments, in particular to an exercise device for assisting a patient with athletic dysarthria, which comprises a round bucket-shaped body, an air inlet pipe and an air tap, a first opening end of the round bucket-shaped body is connected with the air inlet pipe, and the tail end of the air inlet pipe is connected with the air tap. An opening of the air nozzle gradually extends towards the side away from the air inlet pipe to be in a funnel shape, a rotatable tongue blocking plate is arranged at the end, close to the air inlet pipe, of the interior of the air nozzle, and an air channel between the air nozzle and the air inlet pipe is in a communicated state or a separated state through rotation of the tongue blocking plate. The anti-resistance respiratory training device for dysarthria rehabilitation can achieve the anti-resistance respiratory training function of an anti-resistance respiratory training device for dysarthria rehabilitation in the prior art, and meanwhile helps a patient to train the strength and flexibility of tongue muscles.
Owner:FIRST HOSPITAL AFFILIATED TO GENERAL HOSPITAL OF PLA

Method and apparatus for evaluating neurolinguistic disorder using explainable artificial intelligence

The present disclosure provides a method for evaluating the severity of a neurolinguistic disorder using explainable artificial intelligence, implemented by a server. The present method may comprise the steps of: receiving, by a server, a patient's voice data collected through a user terminal; performing audio preprocessing and voice segmentation on the received voice data; extracting dysarthria-specific features from the processed voice data; classifying the severity of dysarthria on the basis of the extracted features using a tree-based machine learning algorithm; generating visualization information by analyzing and visualizing the contribution of the features that influenced the prediction using explainable artificial intelligence techniques; and generating natural language-type large language model (LLM)-generated statements explaining the rationale for a prediction result using an LLM.
Owner:HAII CORP

Quality analysis method and device for synthetic speech, equipment and medium

PendingCN121811848ASpeech synthesisDysarthriaSpeech synthesis
The invention relates to the technical field of speech semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a quality analysis method, device, equipment and medium for synthetic speech, and the method comprises the steps: obtaining an initial audio of a target dysarthria person and a speech text to be synthesized, carrying out the audio coding of the initial audio, and obtaining a speech text to be synthesized; obtaining an initial audio feature; performing voice synthesis according to the initial audio feature and the to-be-synthesized voice text to obtain a synthesized voice; performing text analysis on the synthesized voice according to the to-be-synthesized voice text to obtain voice intelligibility; performing acoustic analysis on the synthesized speech according to the initial audio features to obtain acoustic similarity; performing rhythm analysis on the synthesized voice according to the initial audio to obtain rhythm similarity; and performing quality analysis on the synthetic speech according to the speech intelligibility, the acoustic similarity and the rhythm similarity to obtain the synthetic speech quality. The quality analysis accuracy of the synthetic speech can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition method and device, equipment and storage medium

PendingCN120998205ASpeech recognitionSpeech disorderDysarthria
The embodiment of the invention relates to a voice recognition method and device, equipment and a storage medium. The method provided by the invention comprises the following steps: adjusting the loudness of acquired voice data to a preset loudness range, wherein the voice data is from an object related to dysarthria; extracting at least one speech segment associated with the speech activity from the speech data; providing audio features of the at least one speech segment to a model to generate first text content, where the model is trained based on training speech associated with a preset loudness range; and based on the at least one text expression constraint, adjusting the text expression of the first text content to determine an identification result corresponding to the voice data. In this way, the embodiments of the present disclosure can improve the accuracy of the text content identified by the voice data (e.g., voice data including dysarthria). In addition, through the embodiment of the invention, the interaction efficiency of dysarthria people can be improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1