An AI algorithm-based voice disorder evaluation and training system

By integrating multi-source data and AI algorithms and dynamically adjusting training content, the system addresses the issues of limited diagnostic dimensions and rigid training programs in voice disorder assessment systems. This enables efficient and personalized voice disorder assessment and training, improving diagnostic accuracy and rehabilitation efficiency.

CN120753677BActive Publication Date: 2026-02-24THE SECOND AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510982269.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2026-02-24
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing voice disorder assessment systems lack multi-source data fusion, have a single diagnostic dimension, rigid training programs, poor adaptability, low recognition accuracy when faced with edge scenarios or small sample data, low training efficiency, and poor compliance.

Method used

A voice disorder assessment and training system based on AI algorithms was constructed. The system integrates vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data and dynamic images through a data acquisition module. It adopts a high-order formant detection algorithm and graph neural network, combined with transfer learning and reinforcement learning, to dynamically adjust the training content and frequency, thereby achieving personalized rehabilitation training.

Benefits of technology

It significantly improves the generalization ability of voice disorder assessment, enhances diagnostic accuracy and training efficiency, strengthens the scientific nature and compliance of the rehabilitation process, and supports real-time adjustment of personalized training programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120753677B_ABST
    Figure CN120753677B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice evaluation, in particular to a voice disorder evaluation and training system based on an AI algorithm. The system comprises a data acquisition module, a feature extraction module, a disorder evaluation module, a personalized training module and a collaborative regulation module. The acoustic information of a voice disorder patient is acquired through a sensor network, a high-order formant detection algorithm is used to extract disorder-related features, and a multi-modal joint feature vector is constructed. Then, an evaluation model is established by using transfer learning combined with a graph neural network, so that accurate identification and evaluation of the voice disorder are realized. The personalized training module dynamically generates a rehabilitation training scheme based on a reinforcement learning strategy network; the collaborative regulation module supports real-time two-way interaction between doctors and patients, and intelligently adjusts the rehabilitation plan according to the training feedback. The application can significantly improve the scientificity, accuracy and operability of voice disorder diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice assessment technology, and more particularly to a voice disorder assessment and training system based on AI algorithms. Background Technology

[0002] Voice disorders are a common speech and language disorder characterized by difficulty in vocalization, abnormal voice quality, and weakened voice intensity control. They are prevalent among individuals who use their voice intensively for extended periods, such as teachers, singers, and customer service personnel. They are also frequently seen in patients after laryngeal surgery and those with voice disorders caused by neurological diseases. In severe cases, they can significantly impact a patient's communication abilities and mental health. Currently, clinical assessment and intervention for voice disorders primarily rely on physician subjective judgment, acoustic analysis software, or laryngoscopy. However, these methods have the following limitations: existing systems largely depend on speech acoustic signals, lacking the fusion and analysis of multi-source data such as vocal cord vibration, physiological electromyography, and dynamic imaging, resulting in limited diagnostic dimensions and insufficient accuracy; current voice disorder assessment models are mostly statically trained shallow classifiers, lacking the ability to adapt to different populations and scenarios, leading to low accuracy when dealing with marginal scenarios or small sample data; traditional voice training methods use static, uniform training templates, ignoring differences in patients' pathological states and recovery progress, making it difficult to flexibly adjust training content and pace, resulting in low efficiency and poor adherence during the rehabilitation process. Summary of the Invention

[0003] To address the aforementioned issues, this invention provides an AI-based voice disorder assessment and training system. This invention constructs a system capable of integrating multi-source data such as vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, and dynamic images. It comprehensively applies transfer learning and graph neural networks to enhance the generalization ability of voice disorder assessment, and combines reinforcement learning to achieve dynamic optimization and adjustment of vocal training content, frequency, and intensity. This solves the problems of existing systems having limited diagnostic dimensions, poor adaptability, and rigid training schemes.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0005] A voice disorder assessment and training system based on AI algorithms includes a data acquisition module, a feature extraction module, a noise disorder assessment module, a personalized training module, and a collaborative control module that are connected in sequence via communication.

[0006] The data acquisition module is used to collect acoustic information of patients with voice disorders through a sensor network. The acoustic information includes vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, vocal cord vibration load data, and vocal cord dynamic images.

[0007] The feature extraction module is used to perform time-frequency domain analysis based on the acoustic information using a high-order formant detection algorithm, extract high-order acoustic features associated with obstacles, and construct a multimodal joint feature vector;

[0008] The noise impairment assessment module is used to construct a noise impairment assessment model based on the multimodal joint feature vector, using transfer learning and graph neural network algorithms, to classify and identify different types of voice impairments and generate voice impairment assessment results.

[0009] The personalized training module is used to generate a personalized rehabilitation training plan by using a reinforcement learning-based training strategy optimization algorithm based on the voice disorder assessment results and dynamically adjusting the vocal training content, frequency, and intensity through a strategy network. The training content includes breathing control, resonance adjustment, and glottal closure.

[0010] The collaborative control module is used to establish a real-time two-way data interaction channel between the doctor's terminal and the patient's terminal, and dynamically coordinate the arrangement and adjustment suggestions of rehabilitation tasks between doctors and patients in combination with the user status and the personalized rehabilitation training plan.

[0011] Furthermore, the operation of the feature extraction module includes the following steps:

[0012] The collected vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, and vocal cord dynamic images were preprocessed simultaneously.

[0013] Based on the vocal cord vibration signal and speech acoustic data, a high-order formant detection algorithm is used to extract obstacle-related features of high-order formant frequency, amplitude, frequency drift and harmonic distortion.

[0014] Based on the aforementioned obstacle-related features, combined with laryngeal electromyography signals and vocal cord dynamic images, abnormal patterns of vocal control and physiological movement are identified by calculating electromyography activation parameters and vocal cord kinematic indices. Furthermore, a feature fusion algorithm is used to construct a multimodal joint feature vector.

[0015] Furthermore, the multimodal joint feature vector includes fundamental frequency, frequency jitter, amplitude jitter, harmonic-to-noise ratio, formant frequency change rate, harmonic energy distribution, vocal cord vibration frequency, and glottal closure characteristics.

[0016] Furthermore, the operation of the noise barrier assessment module includes the following steps:

[0017] Based on the multimodal joint feature vector, a patient feature heterogeneous graph structure is constructed, with each patient as a graph unit, which includes vocal cord vibration feature nodes, laryngeal electromyography nodes, and dynamic image keyframe nodes.

[0018] A transfer learning algorithm based on joint maximum mean difference is adopted to map the graph embedding distribution of historical patient sample voice disorder data to the current patient feature heterogeneous graph structure, forming a unified graph feature space for transfer initialization;

[0019] Based on the unified graph feature space, a noise obstacle assessment model is constructed using a graph neural network algorithm. The graph convolution and graph attention mechanisms are integrated to extract cross-modal dependencies and high-order interaction features between key obstacle nodes in the obstacle feature map.

[0020] By combining the graph embedding output, a multi-label classifier is trained to determine the impairment label for each patient sample in the graph. The output includes prediction results and corresponding scores for multiple impairment types, including glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, and spectral instability.

[0021] The classification output is mapped to voice disorder assessment results, including disorder type, severity level, corresponding feature weight interpretation, and assessment confidence interval.

[0022] Furthermore, the construction process of the transfer learning algorithm includes the following steps:

[0023] Based on historical patient sample voice disorder data in the source domain and acoustic information of current patients in the target domain, a heterogeneous graph structure containing vocal cord vibration nodes, electromyography nodes, and dynamic image nodes is constructed as the original graph input for the source and target graphs.

[0024] Based on the heterogeneous graph structure, a joint maximum mean difference transfer learning algorithm is used to perform joint matching of the marginal and conditional distributions of the node embedding representations of the source and target graphs, mapping the two types of graph structures to a unified graph feature space.

[0025] By embedding the unified graph feature space as input and modeling obstacle patterns in the input graph neural network, a noise obstacle recognition model is trained by jointly optimizing the transfer alignment loss and the multi-label classification loss.

[0026] Furthermore, the formula for the noise barrier assessment model is as follows:

[0027]

[0028] Among them, S i denoted as the comprehensive score result for the i-th type of voice disorder; M represents the number of modalities in the multimodal features; K represents the number of key feature points in each modality; h represents the embedding vector; γ represents the embedding vector of the k-th node in the j-th modality at the L-th layer of the graph neural network; j and β j Represents the score intensity adjustment factor and bias term for the j-th modality; α ijThis indicates the importance of the j-th modality to the i-th type of voice disorder; λ represents the centroidal embedding of all embedded nodes in the j-th mode; j Represents the modal regularization factor; R represents the squared Euclidean distance between the embedding feature of the k-th node of the j-th modality and the modality center; jk δ represents the prediction residual recorded by the k-th node of the j-th modality during training; j This represents the adjustment factor for the residual term.

[0029] Furthermore, the construction process of the personalized training module includes the following steps:

[0030] Based on the voice disorder assessment results, the disorder type label, severity level and key acoustic features corresponding to the patient are extracted to construct an individual state space, which serves as the input state for the reinforcement learning strategy network.

[0031] A personalized training model is constructed using a reinforcement learning training strategy optimization algorithm, with the improvement in vocal performance as the reward function, and the vocal training action sequence is dynamically adjusted; the action sequence includes breathing control, resonance adjustment and glottal closure exercises.

[0032] Based on the action sequence, the optimal training action is selected for the current state through a policy network, and an instant reward signal is generated by combining the acoustic improvement after training with the patient's subjective feedback.

[0033] The policy network and value function network are optimized in real time based on the reward signal, and the training content, frequency and intensity are dynamically updated to form a multi-round interactive learning process, which ultimately generates a personalized training scheme instruction set.

[0034] Furthermore, the formula for the personalized training model is as follows:

[0035]

[0036] Wherein, J(π) θ ) represents the policy network π θ The objective function for optimization under parameter θ; a represents the current vocal training action; π θ This represents the training policy network; In terms of strategy π θ Expectation calculation of the sampled action distribution 'a'; s t Indicates the state at step t; a t π represents the training action taken in step t; θ (a t |s t ) represents the output of the policy network in state s t Choose action a t The probability of; Indicates action a t Immediate reward signal after execution; ΔA t Indicates the extent of improvement in key acoustic features before and after training; ω1, ω2, and ω3 represent the weighting coefficients of the improvement enhancement term, strategy optimization term, and load penalty term, respectively, and T represents the time step.

[0037] Furthermore, the collaborative control module establishes a remote synchronous interaction mechanism with the doctor's end, allowing the doctor to remotely adjust personalized training plans based on the patient's training history data, current status, and evaluation results.

[0038] The beneficial effects of this invention are as follows:

[0039] This invention significantly enhances the comprehensive perception of voice disorder characteristics by integrating multi-dimensional information such as vocal cord vibration, electromyography signals, acoustic data, and dynamic images through a data acquisition module, providing higher-quality input data for subsequent assessment. The introduction of a high-order formant detection algorithm, combined with a time-frequency domain analysis mechanism, enables in-depth mining of disorder-related acoustic features, constructing a highly discriminative multimodal joint feature vector, effectively enhancing the system's ability to express complex pathological vocal patterns. Based on a collaborative strategy of transfer learning and graph neural networks, the noise disorder assessment module not only adapts to the heterogeneity of various patient voice samples but also fully models the nonlinear relationships and structural dependencies between features, improving the model's diagnostic performance with few or no samples. By constructing a policy network through reinforcement learning algorithms, the content, frequency, and intensity of vocal training are dynamically adjusted, achieving adaptive optimization of the personalized training module. This makes the training process more aligned with the patient's recovery rhythm and physiological feedback, improving rehabilitation efficiency and compliance. The collaborative control module establishes a two-way data interaction mechanism between the doctor's end and the patient's end, which can synchronize training progress and physical status in real time, support doctors to dynamically intervene and adjust the plan, enhance the scientific nature and operability of remote rehabilitation, and promote the implementation of precision medicine. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of a voice disorder assessment and training system based on AI algorithms according to the present invention.

[0041] Figure 2 This is a flowchart illustrating the operation of a noise barrier assessment module provided in an embodiment of the present invention.

[0042] Figure 3 This is a flowchart illustrating the operation of a personalized training module provided in an embodiment of the present invention. Detailed Implementation

[0043] Please see Figure 1-3As shown, this invention relates to a voice disorder assessment and training system based on AI algorithms.

[0044] Example

[0045] A voice disorder assessment and training system based on AI algorithms includes a data acquisition module, a feature extraction module, a noise disorder assessment module, a personalized training module, and a collaborative control module that are connected in sequence via communication.

[0046] The data acquisition module is used to collect acoustic information of patients with voice disorders through a sensor network. The acoustic information includes vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, vocal cord vibration load data, and vocal cord dynamic images.

[0047] It should be noted that the data acquisition module integrates a wearable flexible electronic laryngeal patch, a laryngeal microphone, a surface electromyography (sEMG) sensor, an IMU sensor, and a medical image acquisition interface. All sensors are networked via Bluetooth Low Energy (BLE 5.0), enabling synchronous acquisition, data fusion, and automatic fault detection. The main control terminal is equipped with a time synchronization module (such as the IEEE 1588 protocol) to ensure millisecond-level alignment of multi-channel signals.

[0048] Data collection specifically includes the following:

[0049] Vocal cord vibration signal acquisition: A flexible electronic laryngeal patch is attached to the thyroid cartilage to acquire highly sensitive mechanical vibration signals. Its sensor chip has a sampling rate of up to 5kHz, and its dynamic range covers the range of normal phonation and abnormal vibration. The built-in analog front-end (AFE) features adaptive gain and low-noise amplification, and has strong resistance to power frequency and environmental interference.

[0050] Laryngeal electromyography (EMG) signal acquisition: Surface EMG electrodes are attached to both sides of the thyroid cartilage and cricoid cartilage to accurately record the electrophysiological activity of laryngeal muscle groups (such as the cricoid cartilage muscle and vocal cord adductor muscles) during phonation. The EMG signal acquisition frequency is 1-2kHz, and real-time bandpass filtering (20-500Hz) and a 50Hz notch filter are used to eliminate power frequency noise.

[0051] Speech acoustic data acquisition: A directional microphone placed close to the throat acquires the raw audio signal during the speech process in real time, with a sampling rate of 48kHz and 16-bit quantization, supplemented by an adaptive noise cancellation (ANC) algorithm to reduce interference from environmental speech and background noise.

[0052] Vocal cord vibration load data acquisition: An IMU sensor (including a triaxial accelerometer and a triaxial gyroscope) is integrated inside the laryngeal patch to capture the minute dynamics of the laryngeal muscles and trachea during the phonation process, and to calculate the vibration frequency, acceleration and instantaneous load changes.

[0053] Medical image acquisition: Supports interface with portable ultrasound probes or electronic laryngoscopes, importing dynamic image sequences of vocal cord vibration and anatomical structures via the standard DICOM protocol. The image acquisition end has a synchronous trigger function, enabling precise timing alignment with physiological signals.

[0054] All data acquisition devices are periodically synchronized using the main control terminal as the clock source, and all signal streams are timestamped. The system automatically detects packet loss and latency, and utilizes interpolation algorithms and sliding window technology to achieve missing data repair and seamless stitching of multi-channel data. Real-time signal integrity detection is included, capable of detecting anomalies such as electrode detachment, microphone distortion, and IMU drift, automatically prompting the user or adjusting acquisition parameters to achieve high reliability and adaptive operation. All acquired data is AES encrypted and uploaded in batches to the cloud server and local backup to ensure user privacy and data security. The system supports breakpoint resume and dynamic bandwidth management, adapting to remote rehabilitation scenarios.

[0055] The feature extraction module is used to perform time-frequency domain analysis based on the acoustic information using a high-order formant detection algorithm, extract high-order acoustic features associated with obstacles, and construct a multimodal joint feature vector;

[0056] The operation of the feature extraction module includes the following steps:

[0057] The collected vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, and vocal cord dynamic images were preprocessed simultaneously.

[0058] Based on the vocal cord vibration signal and speech acoustic data, a high-order formant detection algorithm is used to extract obstacle-related features of high-order formant frequency, amplitude, frequency drift and harmonic distortion.

[0059] Specifically, by combining vocal cord vibration signals and speech acoustic signals, an automatic high-order formant detection process is adopted. The signal is first divided into frames, and spectrum analysis is performed frame by frame to automatically identify the positions of multiple formants such as F1-F5.

[0060] It should be noted that F1 to F5 (the first to fifth resonance peaks):

[0061] F1 (first formant): mainly determined by the degree of vocal tract opening, reflecting the degree of oral cavity opening and chin height.

[0062] F2 (second formant): closely related to the front and back position of the tongue, affecting the brightness of vowels.

[0063] F3 (third formant): Involves the local configuration of the tongue tip, lip shape, and vocal tract, and is more sensitive to individual vowels and pathological changes.

[0064] F4 and F5 (fourth and fifth resonant peaks): These reflect higher-order resonant structures of the vocal tract and are highly sensitive to subtle anatomical abnormalities, vocal cord lesions, and complex disorders. They are often used to distinguish between pathological and healthy acoustic signals.

[0065] In voice disorder analysis, F1 to F3 are generally used to judge the basic acoustic state, while higher-order formants such as F4 and F5 often reveal deep voice disorders and abnormal fine tract movements, which are important bases for early detection and classification diagnosis.

[0066] The system performs frame-by-frame processing on vocal cord vibration signals and speech acoustic data, extracting frequency, amplitude, peak width, and other indicators for F1 to F5 in each frame. For F1 and F2, the system focuses on analyzing their changes under different vowels and training movements to determine the basic vocal state. For F3, F4, and F5, the system focuses on analyzing frequency drift, amplitude abrupt changes, and peak distortion of higher-order peaks to locate pathological features such as irregular vibration of the vocal cord edges, incomplete glottal closure, and upper respiratory tract abnormalities. Using a dynamic tracking algorithm, the system monitors the time-varying curves of F1 to F5 within the vocal cycle, automatically identifying abnormal fluctuations, momentary loss of sound, or noise coverage.

[0067] When abnormal changes are detected in higher-order resonance peaks (F3-F5), the system automatically marks the segment as a "higher-order resonance abnormality zone." If the drift range of F4 and F5 exceeds the normal threshold, or if the peak amplitude weakens or disappears, it suggests a potential risk of dystonia or vocal cord microstructural lesions. Abnormalities in basic peaks such as F1 and F2 (e.g., a significant decrease in frequency or a significant increase in amplitude) suggest possible abnormalities in airflow control or vocal muscle activation. The system automatically generates a resonance peak trajectory diagram, visualizing the dynamic fluctuations of F1-F5 to help doctors quickly pinpoint the time period and type of the disorder.

[0068] Due to individual differences in vocal tract structure, the baseline frequency and amplitude distribution of F1 to F5 vary among patients. The system possesses adaptive adjustment and individualized reference range functions. By combining harmonic energy distribution, the proportion and distribution pattern of harmonic energy near each formant peak are analyzed to further quantify the impact of vocal cord lesions on acoustic characteristics and assist AI in performing fine-grained disorder classification.

[0069] In the multimodal joint feature vector finally output by the system, the dynamic features of F1 to F5 (such as mean, variance, volatility, and anomaly labeling) serve as key input features and participate in the reasoning and decision-making of AI models in subsequent obstacle assessment, personalized training scheme generation, and other processes.

[0070] The system automatically detects minute, irregular changes in frequency and amplitude during vocalization, highlighting areas with severe vibration to aid in the diagnosis of pathological phenomena such as dystonia and asymmetrical vocal cord vibration. It assesses the energy proportion of each harmonic, identifying typical disorder characteristics such as abnormal harmonic enhancement or attenuation and sudden increases in noise components; it automatically outputs key indicators such as harmonic-to-noise ratio and noise energy proportion to quantify the clarity of vocalization. Based on different patients' historical data, the system can adaptively set individualized thresholds for various abnormal indicators, reducing false alarms and improving sensitivity to early or mild voice disorders.

[0071] Based on the aforementioned obstacle-related features, combined with laryngeal electromyography signals and vocal cord dynamic images, abnormal patterns of vocal control and physiological movement are identified by calculating electromyography activation parameters and vocal cord kinematic indices. A feature fusion algorithm is then used to construct a multimodal joint feature vector. The multimodal joint feature vector includes fundamental frequency, frequency jitter, amplitude jitter, harmonic-to-noise ratio, formant frequency change rate, harmonic energy distribution, vocal cord vibration frequency, and glottal closure characteristics.

[0072] Specifically, through multi-channel electromyography (EMG) acquisition, the system automatically analyzes the activation start and end points, activation duration, peak amplitude, and synergistic activation sequence of voice-related muscle groups. The system can intelligently identify abnormal patterns such as "premature activation," "delayed activation," and "asymmetric activation," and automatically issue graded warnings based on the rehabilitation stage. It monitors the changing trends of EMG waveforms during voice training; for example, a gradual decrease in signal amplitude or sudden, disordered high-frequency jitter is automatically flagged as a risk of muscle fatigue or spasm, suggesting adjustments to the training strategy. The module features real-time electrode contact detection, noise level monitoring, and artifact recognition, automatically prompting the user to readjust the equipment when abnormal acquisition occurs, improving data reliability. A deep learning segmentation model is used to automatically locate the vocal cord edges and extract the opening and closing trajectory and area change curve of the vocal cords throughout the entire voice cycle. It automatically identifies phenomena such as incomplete vocal cord closure, excessively short closure duration, or abnormal opening and closing rates. The system registers and compares the movement trajectories of the left and right vocal cords, quantifies symmetry, and assists in the diagnosis of lesions such as unilateral vocal cord paralysis. All detected motion abnormalities, such as delayed closure, local stagnation, and violent shaking, are automatically highlighted on the image sequence by the system, and a list of abnormal events is generated to facilitate doctors in quickly locating the lesion.

[0073] The system automatically compares the activation time of electromyography (EMG) with the actual start and end times of vocal cord movement to detect whether the coupling between the two is abnormal (e.g., EMG activation but vocal cord movement not responding in time), uncovering complex abnormal relationships among the neuromuscular-mechanical systems. Identified multimodal abnormal events are automatically labeled and attributed, such as "delayed EMG activation leading to delayed vocal cord closure," providing multi-dimensional evidence for AI assessment and physician diagnosis.

[0074] The noise impairment assessment module is used to construct a noise impairment assessment model based on the multimodal joint feature vector, using transfer learning and graph neural network algorithms, to classify and identify different types of voice impairments and generate voice impairment assessment results.

[0075] The operation of the noise barrier assessment module includes the following steps:

[0076] Based on the multimodal joint feature vector, a patient feature heterogeneous graph structure is constructed, with each patient as a graph unit, which includes vocal cord vibration feature nodes, laryngeal electromyography nodes, and dynamic image keyframe nodes.

[0077] Specifically, each patient is treated as a graph instance, and different types of features from the multimodal feature vector are assigned to different node types.

[0078] Vocal cord vibration nodes: These include features such as vocal cord vibration frequency, amplitude, and periodic jitter, with each dimension serving as a separate node or node attribute.

[0079] Laryngeal electromyographic nodes: including electromyographic activation level, muscle group synergy indicators, and temporal characteristics.

[0080] Dynamic image key frame nodes: composed of automatically selected key motion frames, with features including vocal cord closure area, symmetry, opening and closing speed, etc.

[0081] The specific relationships between edges and nodes are as follows:

[0082] Temporal edge: Connects similar nodes in adjacent time slices (e.g., vocal cord vibration nodes in consecutive frames).

[0083] Synchronization edge: Connects nodes of different modalities in the same time slice to strengthen the synchronous association between features.

[0084] Collaborative edges: If electromyography abnormalities are found to be accompanied by imaging abnormalities, “collaborative abnormality” edges are automatically generated between relevant nodes to improve the ability to express cross-modal abnormalities.

[0085] Spatial and anatomical structural edges: For example, two nodes with similar anatomical structures in an image can have spatial edges added to support the expression of physical proximity.

[0086] Each node and each edge is accompanied by metadata such as acquisition timestamp, device ID, acquisition quality, and feature score, providing a foundation for subsequent migration and traceability. This forms a heterogeneous graph with complex relationships between acoustic, physiological, and imaging nodes, which can reflect anomalies in a single modality as well as capture cross-modal dependencies.

[0087] A transfer learning algorithm based on joint maximum mean difference is adopted to map the graph embedding distribution of historical patient sample voice disorder data to the current patient feature heterogeneous graph structure, forming a unified graph feature space for transfer initialization;

[0088] Based on the unified graph feature space, a noise obstacle assessment model is constructed using a graph neural network algorithm. The graph convolution and graph attention mechanisms are integrated to extract cross-modal dependencies and high-order interaction features between key obstacle nodes in the obstacle feature map.

[0089] Specifically, the model first aggregates basic information of all nodes in the graph through a base graph convolutional layer (GCN), capturing feature interactions within the local neighborhood. Then, a multi-head graph attention layer (GAT) focuses on key nodes and edges that contribute most to obstacle determination, automatically adjusting the weights of different types of nodes / edges to achieve fine-grained cross-modal aggregation. The model architecture embeds a "node type embedding module" that automatically identifies and processes different types of nodes (such as acoustic, electromyographic, and imaging) according to adaptive strategies, allowing different types of feature information in the same graph to participate in aggregation and decision-making, improving the ability to distinguish complex obstacle patterns. For cross-modal nodes with synchronized timestamps or causal relationships (such as an electromyographic activation node and a corresponding vocal cord movement node), the model automatically constructs special "cooperative edges," assigning higher weights to these edges to ensure that cross-modal abnormal events (such as "delayed electromyographic activation leading to delayed vocal cord closure") can be effectively perceived and identified by the model.

[0090] The model supports multi-level information flow, meaning that each node not only receives information from its immediate neighbors but also captures deep dependencies with distant nodes through a multi-hop propagation mechanism (such as the spatiotemporal correlation between "early electromyography abnormalities and late acoustic disorders"). The model has a built-in anomaly detection submodule that automatically marks abnormal nodes and edges through attention weight analysis and node embedding feature distribution, providing a basis for subsequent disorder type determination and feature interpretation.

[0091] During the model training phase, multi-label manual annotation of historical samples and soft labels output from transfer learning are combined and jointly optimized using a hybrid loss function to improve the ability to identify new types of obstacles and difficult boundary samples. To address real-world challenges such as small sample sizes and imbalanced data, the model automatically adjusts the learning rate and employs an early stopping mechanism to prevent overfitting. After each inference, the model automatically generates a feature contribution weight table and a node / edge heatmap for obstacle determination, helping doctors understand the basis for their judgment and the core abnormalities.

[0092] By combining the graph embedding output, a multi-label classifier is trained to determine the impairment label for each patient sample in the graph. The output includes prediction results and corresponding scores for multiple impairment types, including glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, and spectral instability.

[0093] It should be noted that the unified feature embedding of each patient from the graph neural network output is fed into a multi-label classification network. This classification network typically employs a multilayer perceptron (MLP), residual fully connected network, or lightweight Transformer structure, automatically adapting the number of output nodes based on the number of disorder types. Each output node corresponds to a specific voice disorder type (such as glottal closure dysfunction, laryngeal neuromuscular abnormality, resonance disorder, spectral instability, etc.), outputting the predicted probability and score for each disorder type, rather than mutually exclusive single selections. This allows the system to identify concurrent disorders and the coexistence of multiple complex pathologies. Hierarchical labels for disorder types are supported; for example, "resonance disorder" can be further subdivided into subtypes such as "nasopharyngeal insufficiency." The system can automatically adapt the label structure based on training data to meet the needs of refined diagnosis.

[0094] The classification output is mapped to voice disorder assessment results, including disorder type, severity level, corresponding feature weight interpretation, and assessment confidence interval.

[0095] The construction process of the transfer learning algorithm includes the following steps:

[0096] Based on historical patient sample voice disorder data in the source domain and acoustic information of current patients in the target domain, a heterogeneous graph structure containing vocal cord vibration nodes, electromyography nodes, and dynamic image nodes is constructed as the original graph input for the source and target graphs.

[0097] Based on the heterogeneous graph structure, a joint maximum mean difference transfer learning algorithm is used to perform joint matching of the marginal and conditional distributions of the node embedding representations of the source and target graphs, mapping the two types of graph structures to a unified graph feature space.

[0098] By embedding the unified graph feature space as input and modeling obstacle patterns in the input graph neural network, a noise obstacle recognition model is trained by jointly optimizing the transfer alignment loss and the multi-label classification loss.

[0099] Specifically, the system incorporates a graph generation engine that performs unified node classification, attribute extraction, and edge relationship modeling on multimodal data (vocal cord vibration, electromyography, and imaging) collected from historical patient samples (source domain) and current patients (target domain). This ensures structural consistency between the source and target graphs, facilitating subsequent distribution alignment and feature transfer. Batch normalization and node labeling: The feature data of all graph nodes (such as amplitude, activation level, and dynamic trajectory of each node) are normalized to eliminate baseline bias caused by different data batches and devices. Simultaneously, each type of node is labeled (e.g., "EMG node - left circular muscle," "Imaging node - vocal cord closure," etc.) to ensure comparability of nodes across samples.

[0100] The system employs multi-round automated analysis to first assess the overall distribution of all nodes in the embedding space of both the source and target images. Then, under each obstacle label category, it performs conditional distribution analysis on nodes of the same type (e.g., nodes all belonging to the "EMG abnormality" category) to capture the transfer-sensitive features of "normal" and "obstacle" nodes. During the transfer process, the system can perform distribution alignment for each modality or type of node, gradually reducing the inter-domain distribution differences of features for a particular type of node, improving transfer robustness, and preventing a single feature from dominating the overall transfer process. The transfer learning engine can automatically adjust the transfer step size and alignment strength based on parameters such as the sample size, number of nodes, and label categories of the source and target domains, ensuring effective transfer even in small sample / weak label scenarios. The system can automatically identify samples with abnormal distributions and significant embedding shifts during the transfer process, minimizing the risk of transfer failure through mechanisms such as data augmentation, feature perturbation, or soft label correction.

[0101] Ultimately, after multiple rounds of distribution alignment and node embedding, the multimodal heterogeneous graph data from all patients are unified into a single high-dimensional feature space, ensuring the input standardization and cross-domain generalization ability of subsequent graph neural network models. The system is equipped with a transfer alignment visualization tool, allowing doctors and engineers to intuitively view the changes in graph structure distribution before and after transfer learning, and quickly assess the effectiveness of transfer learning and data quality.

[0102] Furthermore, the formula for the noise barrier assessment model is as follows:

[0103]

[0104] Among them, S i denoted as the comprehensive score result for the i-th type of voice disorder; M represents the number of modalities in the multimodal features; K represents the number of key feature points in each modality; h represents the embedding vector; γ represents the embedding vector of the k-th node in the j-th modality at the L-th layer of the graph neural network; j and β j The scoring intensity adjustment factor and bias term for the j-th mode are used to control the response amplitude and baseline shift of modes such as vocal cord vibration in the scoring, respectively; α ij This indicates the importance of the j-th modality to the i-th type of voice disorder; λ represents the centroidal embedding of all embedded nodes in the j-th mode; j Represents the modal regularization factor; R represents the squared Euclidean distance between the embedding feature of the k-th node of the j-th modality and the modality center; jk δ represents the prediction residual recorded by the k-th node of the j-th modality during training; j This represents the adjustment factor for the residual term.

[0105] The calculation formula is as follows:

[0106]

[0107] in, The output of node k in the acoustic mode may represent the rate of change of formant frequency, harmonic-to-noise ratio, etc.

[0108] R jk The calculation formula is as follows:

[0109]

[0110] in, This represents the true obstacle label of the corresponding node in the training sample; This represents the obstacle prediction result output by the graph neural network.

[0111] The personalized training module is used to generate a personalized rehabilitation training plan by using a reinforcement learning-based training strategy optimization algorithm based on the voice disorder assessment results and dynamically adjusting the vocal training content, frequency, and intensity through a strategy network. The training content includes breathing control, resonance adjustment, and glottal closure.

[0112] The construction process of the personalized training module includes the following steps:

[0113] Based on the voice disorder assessment results, the disorder type label, severity level and key acoustic features corresponding to the patient are extracted to construct an individual state space, which serves as the input state for the reinforcement learning strategy network.

[0114] Specifically, the module automatically reads the noise barrier assessment results, including each barrier type label (such as glottal closure barrier, resonance barrier, etc.), severity level (such as mild, moderate, severe), and key acoustic features (such as F0, HNR, formant frequency drift, closure area, etc.).

[0115] Combine patient history of training response, subjective self-evaluation (such as vocal difficulty score, fatigue feedback), and basic physiological parameters (such as age, past medical history).

[0116] The aforementioned multidimensional data is combined into an individual state space, which serves as the environmental state input to the reinforcement learning policy network. For example:

[0117] State vector = [obstacle type 1, severity, F0 abnormality, closure feature abnormality, breathing control ability, improvement in the last training, fatigue feedback, etc.].

[0118] The state space can be dynamically expanded, such as by adding real-time physiological parameters of the patient (heart rate, respiratory rate) and training compliance indicators, to achieve a comprehensive dynamic characterization.

[0119] A personalized training model is constructed using a reinforcement learning training strategy optimization algorithm, with the improvement in vocal performance as the reward function, and the vocal training action sequence is dynamically adjusted; the action sequence includes breathing control, resonance adjustment and glottal closure exercises.

[0120] Specifically, each patient's training environment is adaptively constructed based on their disorder type, severity, previous training response, and physiological characteristics. The system automatically sets the type, range, and limitations of available training movements based on the assessment report. For example, for patients with severe glottic closure disorder, only basic breathing and mild closure training are initially available to prevent overtraining. The personalized training model supports dynamically adding or removing movement types based on the patient's recovery progress and real-time feedback; for example, advanced resonance regulation training is gradually introduced as performance improves, while high-intensity training movements are temporarily disabled when performance declines.

[0121] The reward system not only relies on improvements in acoustic parameters detected by AI (such as increased F0, improved HNR, and abnormally reduced formant peaks), but also integrates multiple indicators such as electromyographic signals (such as improved activation coordination), imaging parameters (such as optimized vocal cord closure symmetry), and patient subjective self-evaluations (such as training comfort and confidence scores), achieving a "soft and hard" combination of reward signals. If the system detects physiological abnormalities (such as over-activation of electromyographic signals, abnormal fatigue after training, sore throat, etc.), it immediately applies a negative reward and forcibly adjusts the subsequent training plan to ensure training safety and scientific rigor.

[0122] Reinforcement learning models can automatically plan action sequences based on the data output of each training round. For example, they might start with 5 minutes of breathing training, followed by a brief resonance adjustment, and finally arrange moderate-intensity glottal closure exercises, achieving progressive training from easy to difficult and from basic to advanced levels. If acoustic parameters, electromyography, and subjective evaluations all meet the standards for several consecutive rounds, the model automatically unlocks higher-level training content. For instance, resonance training actions will be expanded from single phonations to emotional intonation and complex syllable combinations.

[0123] Based on the action sequence, the optimal training action is selected for the current state through a policy network, and an instant reward signal is generated by combining the acoustic improvement after training with the patient's subjective feedback.

[0124] Specifically, the system collects the patient's current state vector in real time (including the latest acoustic data, electromyographic feedback, training fatigue level, historical completion status, etc.) and inputs it into the policy network. The network outputs the optimal next training action (e.g., "perform 3 minutes of diaphragmatic breathing training") through forward inference and dynamically adjusts training parameters (e.g., training time, rest interval). If the patient's subjective feedback is fatigue, the system prioritizes recommending low-intensity training or arranging rest; if fluctuations in the patient's state are detected, the system temporarily adjusts the training content to avoid monotonous or high-intensity operations that could lead to decreased training compliance.

[0125] After the training exercises are performed, the module synchronously collects multimodal data (such as the fundamental frequency and harmonic-to-noise ratio of acoustic signals, the activation level of electromyography signals, and the vocal cord closure status from image analysis), combined with the patient's subjective evaluation for that session (such as "Did this training feel easier / is the pronunciation more stable?"). The system pre-sets a multi-dimensional reward weighting scheme: for example, acoustic improvement is assigned a weight of 40%, electromyography and image improvement are each assigned 20%, and subjective feedback is assigned 20%. If any dimension shows a negative indicator (such as self-reported excessive fatigue or a decline in acoustic parameters), the overall reward is reduced, forcing the model to adjust its next strategy.

[0126] Based on the reward signals received, the policy network immediately updates its parameters (e.g., through Q-values ​​or policy gradients in reinforcement learning) to ensure the model continuously adapts towards the optimal rehabilitation goal. All interactions, reward feedback, and training actions are automatically archived into the patient's personal training archive. The model is periodically reviewed and retrained to identify historically efficient action combinations and repeatedly ineffective actions, further optimizing future action selections.

[0127] When a patient experiences unexpected reactions or abnormal progress, the system automatically sends a notification to the doctor. Doctors can modify the training exercise library online, adjust reward parameters, and even temporarily switch to manual intervention mode, enhancing system security and professionalism. Based on the patient's progress, the system regularly sends incentive messages (such as rehabilitation badges and training achievement reminders) to increase patient engagement and create an efficient positive feedback mechanism.

[0128] The policy network and value function network are optimized in real time based on the reward signal, and the training content, frequency and intensity are dynamically updated to form a multi-round interactive learning process, which ultimately generates a personalized training scheme instruction set.

[0129] Specifically, based on accumulated immediate rewards, the parameters of the policy network and value function network are updated in real time through backpropagation and policy gradient methods, improving the adaptability to diverse patient states. An experience replay and exploration mechanism is employed to ensure that the model not only utilizes historically optimal actions but also explores new actions to discover potentially efficient training methods.

[0130] The strategy network can adjust parameters such as training movement type, training intensity (e.g., gradually increasing the difficulty of breathing control training), and frequency (automatic push of daily / weekly training plans) in real time. If adverse reactions occur or progress is slow, the system automatically reduces the intensity or guides the patient to a more suitable type of exercise.

[0131] Finally, the system summarizes the optimal action sequences and parameters from multiple rounds of interactive learning to generate a personalized training program instruction set for patients, including: daily training action arrangements (the order and frequency of breathing-resonance-closing); detailed parameters for each training exercise (such as number of repetitions, number of sets, and target acoustic improvement indicators); and suggestions for rest and assessment intervals.

[0132] Furthermore, the formula for the personalized training model is as follows:

[0133]

[0134] Wherein, J(π) θ ) represents the policy network π θ The objective function for optimization under parameter θ; a represents the current vocal training action; π θ This represents the training policy network; In terms of strategy π θ Expectation calculation of the sampled action distribution 'a'; s t Indicates the state at step t; a t π represents the training action taken in step t; θ (a t |s t ) represents the output of the policy network in state s t Choose action a t The probability of; Indicates action a t Immediate reward signal after execution; ΔA t Indicates the extent of improvement in key acoustic features before and after training; ω1 represents the vocal load index at step t, reflecting the degree of vocal cord fatigue or stress caused by the current training; ω1, ω2, and ω3 represent the weighting coefficients of the improvement enhancement term, strategy optimization term, and load penalty term, respectively; and T represents the time step, the total number of steps in each round of vocal training.

[0135] ΔA t The calculation formula is as follows:

[0136]

[0137] in, and κ represents the numerical values ​​of the j-th acoustic feature before and after the t-th training step, such as frequency jitter, amplitude jitter, and harmonic-to-noise ratio; j The value represents the contribution weight of the j-th feature index; N represents the number of higher-order acoustic features, typically 7–10.

[0138] The collaborative control module is used to establish a real-time two-way data interaction channel between the doctor's terminal and the patient's terminal. It dynamically coordinates the arrangement and adjustment suggestions of rehabilitation tasks between doctors and patients by combining the user's status and the personalized rehabilitation training plan. The collaborative control module establishes a remote synchronous interaction mechanism with the doctor's terminal, which can remotely adjust the personalized training plan based on the patient's training history data, current status and evaluation results.

[0139] Specifically, the collaborative control module consists of a patient terminal (such as an app) and a doctor-side management platform, communicating via an encrypted internet channel. The core of the module is a real-time task scheduling engine with built-in hierarchical access control and message push functionality. All data, task scheduling, and communication processes rely on a cloud service platform for real-time synchronization and backup, and are equipped with a local caching strategy to ensure fault tolerance in weak network or network outage situations.

[0140] The system supports encrypted synchronization and targeted push of training data, rehabilitation progress, real-time acquired signals, AI assessment reports, and other information between patient and doctor terminals. It employs WebSocket or MQTT protocols to achieve low-latency, highly reliable data transmission. Events such as training completion, anomaly detection, and significant data changes are automatically triggered and pushed to both doctor and patient terminals, enabling proactive reminders and anomaly alerts.

[0141] Based on the impairment assessment results, historical training performance, and rehabilitation goals set by the doctor, the system automatically generates personalized training plans and daily rehabilitation tasks, which are then assigned to the patient. The plan includes parameters such as training content, time, frequency, and intensity. Doctors can adjust the rehabilitation plan online, modify training goals, or issue new training exercises based on real-time assessment data, patient feedback, and AI suggestions, with all changes instantly synchronized to the patient. The module automatically records training completion rate, patient compliance, and training effectiveness scores, providing visual reports to the doctor for follow-up and decision-making reference.

[0142] The module integrates an encrypted video consultation channel, allowing doctors to remotely guide patients in vocal training and correct movement in real time, supporting screen sharing and training demonstrations. It supports diverse communication between doctors and patients, including instant messaging, voice, images, and documents. Doctors can regularly send training reminders, rehabilitation suggestions, incentives, or warnings to improve patient adherence. For scenarios where doctors manage multiple patients simultaneously, the module supports task prioritization, batch adjustments, and abnormal grouping alerts, improving doctor efficiency.

[0143] All data transmissions within the module are encrypted using SSL / TLS, while core sensitive data such as medical records and training logs are stored using AES256 encryption. Different users (doctors / patients / administrators) have tiered access permissions, and all remote operations, adjustment suggestions, and critical task changes automatically generate audit logs for easy traceability and compliance oversight. The module supports data anonymization, privacy mode switching, and one-click privacy authorization for patients, ensuring the secure and compliant transfer of sensitive data in doctor-patient collaboration.

[0144] In summary, this invention integrates a flexible electronic laryngeal patch, sEMG, a condenser microphone, an IMU, and an imaging interface to ensure millisecond-level alignment of multiple signal channels, including acoustic, electromyographic, and dynamic imaging signals. This significantly improves the temporal consistency and data quality of multi-source signals, providing a solid foundation for subsequent AI analysis. By employing a high-order formant detection algorithm to deeply mine key acoustic features such as F1 to F5, combined with dynamic tracking, spectral anomaly identification, and harmonic energy analysis, it can precisely capture early, occult lesions (such as resonance disorders and vocal cord incomplete closure), improving the accuracy and sensitivity of early screening for voice disorders.

[0145] This invention constructs a heterogeneous feature map of "acoustics-electromyography-imaging" and introduces a graph neural network model. The system can effectively integrate cross-modal and cross-temporal vocal abnormality features, automatically discover potential synergistic abnormal relationships (such as delayed electromyography leading to closure abnormalities), and visualize the judgment criteria through an attention mechanism, enhancing the clinical interpretability of AI diagnosis. The Joint Maximum Mean Difference (JMMD) algorithm in transfer learning is used to align the distribution of historical data with the current sample, alleviating the modeling challenges caused by the scarcity of clinical samples and individual differences, ensuring that the model has good adaptability and generalizability across different patient groups.

[0146] This invention constructs a training strategy optimization network based on reinforcement learning. It adjusts the training content, frequency, and intensity in real time based on evaluation results, self-assessment feedback, and physiological state, and generates multi-dimensional rewards through multimodal feedback signals, effectively improving the scientific rigor, safety, and recovery efficiency of training. The collaborative control module establishes a low-latency, encrypted, two-way remote interaction mechanism, supporting doctors to adjust training plans, push feedback, and provide video guidance in real time. The system automatically records task completion and training response, improving doctor management efficiency and patient motivation, and is suitable for remote medical scenarios such as home rehabilitation.

[0147] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A voice disorder assessment and training system based on AI algorithms, characterized in that, It includes a data acquisition module, a feature extraction module, a noise barrier assessment module, a personalized training module, and a collaborative control module that are connected in sequence via communication. The data acquisition module is used to collect acoustic information of patients with voice disorders through a sensor network. The acoustic information includes vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, vocal cord vibration load data, and vocal cord dynamic images. The feature extraction module is used to perform time-frequency domain analysis based on the acoustic information using a high-order formant detection algorithm, extract high-order acoustic features associated with obstacles, and construct a multimodal joint feature vector; The noise impairment assessment module is used to construct a noise impairment assessment model based on the multimodal joint feature vector, employing transfer learning and graph neural network algorithms. This model classifies and identifies different types of voice impairments and generates voice impairment assessment results. The construction process of the transfer learning and graph neural network algorithms includes the following steps: Based on historical patient sample voice disorder data in the source domain and acoustic information of current patients in the target domain, a patient feature heterogeneous graph structure containing vocal cord vibration nodes, electromyography nodes, and dynamic image nodes is constructed as the original graph input for the source and target graphs. Based on the heterogeneous graph structure of the patient features, a transfer learning algorithm with joint maximum mean difference is used to perform joint matching of the marginal and conditional distributions of the node embedding representations of the source and target graphs, mapping the two types of graph structures to a unified graph feature space. The unified graph feature space is used as the input embedding, and obstacle pattern modeling is performed in the input graph neural network algorithm. The noise obstacle assessment model is trained by jointly optimizing the transfer alignment loss and the multi-label classification loss. The personalized training module is used to generate a personalized rehabilitation training plan by using a reinforcement learning-based training strategy optimization algorithm based on the voice disorder assessment results and dynamically adjusting the vocal training content, frequency, and intensity through a strategy network. The training content includes breathing control, resonance adjustment, and glottal closure. The collaborative control module is used to establish a real-time two-way data interaction channel between the doctor's terminal and the patient's terminal, and dynamically coordinate the arrangement and adjustment suggestions of rehabilitation tasks between doctors and patients in combination with the user status and the personalized rehabilitation training plan.

2. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that, The operation of the feature extraction module includes the following steps: The collected vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data, and vocal cord dynamic images were preprocessed simultaneously. Based on the vocal cord vibration signal and speech acoustic data, a high-order formant detection algorithm is used to extract obstacle-related features of high-order formant frequency, amplitude, frequency drift and harmonic distortion. Based on the aforementioned obstacle-related features, combined with laryngeal electromyography signals and vocal cord dynamic images, abnormal patterns of vocal control and physiological movement are identified by calculating electromyography activation parameters and vocal cord kinematic indices. Furthermore, a feature fusion algorithm is used to construct a multimodal joint feature vector.

3. The voice disorder assessment and training system based on AI algorithm according to claim 2, characterized in that, The multimodal joint feature vector includes fundamental frequency, frequency jitter, amplitude jitter, harmonic-to-noise ratio, formant frequency change rate, harmonic energy distribution, vocal cord vibration frequency, and glottal closure characteristics.

4. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that, The operation of the noise barrier assessment module includes the following steps: Based on the multimodal joint feature vector, a patient feature heterogeneous graph structure is constructed, with each patient as a graph unit, which includes vocal cord vibration feature nodes, laryngeal electromyography nodes, and dynamic image keyframe nodes. A transfer learning algorithm based on joint maximum mean difference is adopted to map the graph embedding distribution of historical patient sample voice disorder data to the current patient feature heterogeneous graph structure, forming a unified graph feature space for transfer initialization; Based on the unified graph feature space, a noise obstacle assessment model is constructed using a graph neural network algorithm. The graph convolution and graph attention mechanisms are integrated to extract cross-modal dependencies and high-order interaction features between key obstacle nodes in the obstacle feature map. By combining the graph embedding output, a multi-label classifier is trained to determine the impairment label for each patient sample in the graph. The output includes prediction results and corresponding scores for multiple impairment types, including glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, and spectral instability. The classification output is mapped to voice disorder assessment results, including disorder type, severity level, corresponding feature weight interpretation, and assessment confidence interval.

5. The voice disorder assessment and training system based on AI algorithm according to claim 4, characterized in that, The formula for the noise barrier assessment model is as follows: in, denoted as the comprehensive score result for the i-th type of voice disorder; M represents the number of modalities in the multimodal features; K represents the number of key feature points in each modality; h represents the embedding vector; This represents the embedding vector output by the k-th node in the j-th modality at the L-th layer of the graph neural network; and This represents the score intensity adjustment factor and bias term for the j-th mode; This indicates the importance of the j-th modality to the i-th type of voice disorder; This represents the center-mean embedding of all embedded nodes in the j-th modality; Represents the modal regularization factor; Let Euclidean distance be the squared Euclidean distance between the embedding feature of the k-th node of the j-th modality and the modality center. This represents the prediction residual recorded by the k-th node of the j-th modality during training; This represents the adjustment factor for the residual term.

6. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that, The construction process of the personalized training module includes the following steps: Based on the voice disorder assessment results, the disorder type label, severity level and key acoustic features corresponding to the patient are extracted to construct an individual state space, which serves as the input state for the reinforcement learning strategy network. A personalized training model is constructed using a reinforcement learning training strategy optimization algorithm, with the improvement in vocal performance as the reward function, and the vocal training action sequence is dynamically adjusted; the action sequence includes breathing control, resonance adjustment and glottal closure exercises. Based on the action sequence, the optimal training action is selected for the current state through a policy network, and an instant reward signal is generated by combining the acoustic improvement after training with the patient's subjective feedback. The policy network and value function network are optimized in real time based on the reward signal, and the training content, frequency and intensity are dynamically updated to form a multi-round interactive learning process, which ultimately generates a personalized training scheme instruction set.

7. The voice disorder assessment and training system based on AI algorithm according to claim 6, characterized in that, The formula for the personalized training model is as follows: in, Representational Policy Network In parameters The optimization objective function is as follows; 'a' represents the current vocal training action; This represents the training policy network; Indicating in strategy Expectation calculation under the sampled action distribution a; Indicates the state at step t; This indicates the training action taken in step t; This indicates the output of the policy network in the state. Select action The probability of; Indicates action Immediate reward signal after execution; Indicates the extent of improvement in key acoustic features before and after training; This represents the vocal load index at step t; , and represents the weighting coefficients of the improvement enhancement term, the strategy optimization term, and the load penalty term, respectively, and T represents the time step.

8. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that, The collaborative control module establishes a remote synchronous interaction mechanism with the doctor's end, allowing the doctor to remotely adjust personalized training plans based on the patient's training history data, current status, and evaluation results.

Citation Information

Patent Citations

  • System and application for evaluation of voice and speech disorders and speech-language therapy customized for parkinson patients

    KR102668964B1

  • System and method for pathological voice recognition and computer-readable storage medium

    US20230386504A1