Throat disorder assessment and training system based on AI algorithm

By integrating multi-source data and AI algorithms, a voice disorder assessment and training system is constructed, which solves the problems of the existing system's single diagnostic dimension and rigid training program, and realizes efficient and personalized voice disorder assessment and training.

CN120753677AActive Publication Date: 2025-10-10THE SECOND AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510982269.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-10
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

The existing voice disorder assessment system lacks multi-source data fusion, has a single diagnostic dimension, poor adaptability, and rigid training programs, resulting in insufficient diagnostic accuracy and low rehabilitation efficiency.

Method used

Build a voice disorder assessment and training system based on AI algorithms, integrate vocal cord vibration signals, laryngeal electromyography signals, speech acoustic data and dynamic images through the data acquisition module, use transfer learning and graph neural networks to improve the assessment generalization ability, and combine reinforcement learning to dynamically adjust the training content and intensity.

Benefits of technology

Significantly improve the comprehensive perception ability of voice disorder characteristics, improve diagnostic accuracy and rehabilitation efficiency, realize personalized training, and enhance the adaptability and operability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120753677A_ABST
    Figure CN120753677A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice assessment, in particular to a voice disorder assessment and training system based on an AI algorithm. Comprising a data acquisition module, a feature extraction module, an obstacle assessment module, a personalized training module and a cooperative regulation and control module. Acoustic information of a patient suffering from voice disorder is acquired through a sensor network, disorder related features are extracted by adopting a high-order formant detection algorithm, and a multi-modal joint feature vector is constructed. And then, establishing an evaluation model by using transfer learning in combination with a graph neural network, so as to realize accurate recognition and evaluation of the voice impairment. The personalized training module dynamically generates a rehabilitation training scheme based on the reinforcement learning strategy network; the cooperative regulation and control module supports real-time bidirectional interaction between doctors and patients and intelligently adjusts a rehabilitation plan according to training feedback. According to the invention, scientificity, accuracy and operability of diagnosis and treatment of the voice disorder can be obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice evaluation, in particular to a voice disorder evaluation and training system based on an AI algorithm. BACKGROUND

[0002] Voice disorder is a common speech and language disorder, which is manifested as difficulty in vocalization, abnormal voice quality, weakened sound intensity control ability, etc. It is widely present in long-term high-intensity vocalization groups such as teachers, singers and customer service personnel, and is also common in postoperative patients with larynx and patients with vocalization disorder caused by nervous system diseases. In severe cases, it can significantly affect the communication ability and mental health of patients. At present, the voice disorder is mainly evaluated and intervened by means of subjective judgment of doctors, acoustic analysis software or laryngoscope examination. However, these methods have the following disadvantages: the existing system mostly relies on speech acoustic signals, lacks the fusion collection and analysis of multi-source data such as vocal cord vibration, physiological electromyography and dynamic image, resulting in limited diagnostic dimension and insufficient accuracy; the current voice disorder evaluation model is mostly a shallow classifier trained statically, lacks migration adaptation ability for different groups and different scenes, and has low recognition accuracy when facing edge scenes or small sample disorder data; the traditional vocalization training method uses a static and unified training template, ignores the differences in pathological state and recovery progress of patients, and the training content and rhythm are difficult to adjust flexibly, resulting in low efficiency and poor compliance in the rehabilitation process. SUMMARY

[0003] To solve the above problems, the present application provides a voice disorder evaluation and training system based on an AI algorithm, which can fuse multi-source data such as vocal cord vibration signal, laryngeal electromyography signal, speech acoustic data and dynamic image, comprehensively apply transfer learning and graph neural network to improve the generalization ability of voice disorder evaluation, and realize dynamic optimization and adjustment of vocalization training content, frequency and intensity by combining reinforcement learning, thereby solving the problems of single diagnostic dimension, poor adaptability and rigid training scheme of the existing system.

[0004] To achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0005] A voice disorder evaluation and training system based on an AI algorithm, comprising a data collection module, a feature extraction module, a voice disorder evaluation module, a personalized training module and a collaborative regulation module connected in sequence;

[0006] The data collection module is used to collect acoustic information of voice disorder patients through a sensor network, and the acoustic information includes vocal cord vibration signal, laryngeal electromyography signal, speech acoustic data, vocal cord vibration load data and vocal cord dynamic image;

[0007] The feature extraction module is used to perform time-frequency domain analysis based on the acoustic information using a high-order formant detection algorithm to extract high-order acoustic features associated with the obstacle and construct a multimodal joint feature vector;

[0008] The noise disorder assessment module is used to construct a noise disorder assessment model based on the multimodal joint feature vector using transfer learning and graph neural network algorithms to classify and identify different types of voice disorders and generate voice disorder assessment results;

[0009] The personalized training module is used to generate a personalized rehabilitation training program based on the voice disorder assessment results, using a training strategy optimization algorithm based on reinforcement learning, and dynamically adjusting the vocal training content, frequency, and intensity through a strategy network; the training content includes breathing control, resonance regulation, and glottal closure;

[0010] The collaborative control module is used to establish a real-time two-way data interaction channel between the doctor's terminal and the patient's terminal, and dynamically coordinate the rehabilitation task arrangement and adjustment suggestions between the doctor and the patient in combination with the user status and the personalized rehabilitation training plan.

[0011] Furthermore, the operation process of the feature extraction module includes the following steps:

[0012] Synchronously pre-process the collected vocal cord vibration signals, laryngeal electromyographic signals, speech acoustic data and vocal cord dynamic images;

[0013] Based on the vocal cord vibration signal and speech acoustic data, a high-order formant detection algorithm is used to extract obstacle-related features such as high-order formant frequency, amplitude, frequency drift, and harmonic distortion;

[0014] Based on the obstacle-related features, combined with laryngeal electromyographic signals and vocal cord dynamic images, the abnormal patterns of vocal control and physiological movements are explored by calculating electromyographic activation parameters and vocal cord kinematic indicators, and a feature fusion algorithm is used to construct a multimodal joint feature vector.

[0015] Furthermore, the multimodal joint feature vector includes fundamental frequency, frequency jitter, amplitude jitter, harmonic-to-noise ratio, formant frequency change rate, harmonic energy distribution, vocal cord vibration frequency and glottal closure characteristics.

[0016] Furthermore, the operation process of the noise nuisance assessment module includes the following steps:

[0017] Based on the multimodal joint feature vector, each patient is used as a graph unit to construct a patient feature heterogeneous graph structure including vocal cord vibration feature nodes, laryngeal electromyography nodes, and dynamic image key frame nodes;

[0018] A transfer learning algorithm based on joint maximum mean difference is used to map the graph embedding distribution of historical patient sample voice disorder data to the current patient feature heterogeneous graph structure, forming a unified graph feature space for transfer initialization.

[0019] Based on the unified graph feature space, a graph neural network algorithm is used to build a noise obstacle assessment model. This model integrates graph convolution and graph attention mechanisms to extract cross-modal dependencies in the obstacle feature graph and high-order interaction features between key obstacle nodes.

[0020] Combined with the graph embedding output, a multi-label classifier is trained to determine the disorder label for each patient sample in the graph, outputting prediction results and corresponding scores for multiple disorder types, including glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, and spectral instability.

[0021] The classification output results are mapped to voice disorder assessment results, including disorder type, severity level, corresponding feature weight explanation and assessment confidence interval.

[0022] Furthermore, the construction process of the transfer learning algorithm includes the following steps:

[0023] Based on historical patient voice disorder data in the source domain and the acoustic information of current patients in the target domain, a heterogeneous graph structure containing vocal cord vibration nodes, myoelectric nodes, and dynamic image nodes is constructed as the original graph input of the source graph and target graph;

[0024] Based on the heterogeneous graph structure, a joint maximum mean difference transfer learning algorithm is used to perform joint matching of marginal and conditional distributions on the node embedding representations of the source graph and the target graph, mapping the two types of graph structures into a unified graph feature space;

[0025] The unified graph feature space is embedded as input and input into the graph neural network for obstacle pattern modeling. By jointly optimizing the transfer alignment loss and the multi-label classification loss, a noise obstacle recognition model is trained.

[0026] Furthermore, the noise barrier assessment model is formulated as follows:

[0027]

[0028] Among them, S i represents the comprehensive score result of the i-th type of voice disorder; M represents the number of modes of the multimodal feature; K represents the number of key feature points of each mode; h represents the embedding vector; represents the embedding vector output by the kth node in the jth modality in the Lth layer of the graph neural network; γ j and β j represents the score strength adjustment factor and bias term of the j-th modality; α ijrepresents the importance of the jth mode to the i-th type of voice disorder; represents the central mean embedding of all embedded nodes under the jth mode; λ j represents the modal regularization factor; represents the square of the Euclidean distance between the embedding feature of the kth node of the jth mode and the modal center; R jk represents the prediction residual recorded by the kth node of the jth modality during training; δ j Represents the adjustment factor of the residual term.

[0029] Furthermore, the construction process of the personalized training module includes the following steps:

[0030] Based on the voice disorder assessment results, extract the patient's corresponding disorder type label, severity level, and key acoustic features, and construct an individual state space as the input state of the reinforcement learning strategy network;

[0031] A personalized training model is constructed using a reinforcement learning training strategy optimization algorithm. The improvement in vocal performance is used as a reward function to dynamically adjust the vocal training action sequence, which includes breathing control, resonance regulation, and glottal closure exercises.

[0032] Based on the action sequence, the policy network selects the optimal training action for the current state, and generates an immediate reward signal by combining the acoustic improvement after training and the patient's subjective feedback;

[0033] The strategy network and value function network are optimized in real time according to the reward signal, and the training content, frequency and intensity are dynamically updated to form a multi-round interactive learning process, and finally generate a personalized training plan instruction set.

[0034] Furthermore, the formula of the personalized training model is as follows:

[0035]

[0036] Among them, J(π θ ) represents the policy network π θ The optimization objective function under the parameter θ; a represents the current vocal training action; π θ represents the training policy network; In the strategy π θ The expected operation under the distribution of the sampled action a; s t Indicates the state of step t; a t represents the training action taken in step t; π θ (a t |s t ) represents the output of the policy network in state s t Next select action a t probability; Indicates action a t The immediate reward signal after execution; ΔA t Indicates the improvement in key acoustic features before and after training; represents the phonation load index at step t; ω1, ω2, and ω3 represent the weighted coefficients of the improvement enhancement term, strategy optimization term, and load penalty term, respectively; and T represents the time step.

[0037] Furthermore, the collaborative control module establishes a remote synchronous interaction mechanism with the doctor side, and the doctor side can remotely adjust the personalized training plan based on the patient's training history data, current status and evaluation results.

[0038] The beneficial effects of the present invention are:

[0039] The present invention integrates multi-dimensional information such as vocal cord vibration, electromyographic signals, acoustic data, and dynamic images through the data acquisition module, which significantly improves the comprehensive perception of voice disorder characteristics and provides higher quality input data for subsequent evaluation. The introduction of high-order resonance peak detection algorithm, combined with the time-frequency domain analysis mechanism, can deeply mine the acoustic characteristics associated with the disorder, construct a multimodal joint feature vector with high discriminative power, and effectively enhance the system's ability to express complex pathological vocal patterns. Based on the collaborative strategy of transfer learning and graph neural network, the noise disorder assessment module can not only adapt to the heterogeneity of various patient voice samples, but also fully model the nonlinear relationship and structural dependency between features, and improve the diagnostic performance of the model with few samples or no samples. By constructing a strategy network through the reinforcement learning algorithm, the vocal training content, frequency and intensity are dynamically adjusted to achieve adaptive optimization of the personalized training module, so that the training process is more in line with the patient's recovery rhythm and physiological feedback, and improve rehabilitation efficiency and compliance. The collaborative control module establishes a two-way data interaction mechanism between the doctor and the patient, which can synchronize training progress and physical status in real time, support doctors' dynamic intervention and adjustment plans, enhance the scientific nature and operability of remote rehabilitation, and promote the implementation of precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a module diagram of a voice disorder assessment and training system based on AI algorithm of the present invention.

[0041] Figure 2 4 is a flowchart of the operation process of the noise nuisance assessment module provided by one embodiment of the present invention.

[0042] Figure 3 It is a flowchart of the operation process of the personalized training module provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0043] See also Figure 1-3As shown, the present invention relates to a voice disorder assessment and training system based on AI algorithm.

[0044] Example

[0045] A voice disorder assessment and training system based on an AI algorithm, comprising a data acquisition module, a feature extraction module, a noise disorder assessment module, a personalized training module, and a collaborative control module, which are sequentially connected in communication;

[0046] The data acquisition module is used to collect acoustic information of patients with voice disorders through a sensor network, wherein the acoustic information includes vocal cord vibration signals, laryngeal electromyographic signals, speech acoustic data, vocal cord vibration load data, and vocal cord dynamic images;

[0047] The data acquisition module integrates a wearable flexible electronic laryngeal patch, a laryngeal microphone, a surface electromyography (sEMG) sensor, an IMU sensor, and a medical image acquisition interface. These sensors are networked via Bluetooth Low Energy (BLE 5.0), enabling synchronized data acquisition, data fusion, and automatic fault detection. The main acquisition control terminal incorporates a time synchronization module (e.g., IEEE 1588 protocol) to ensure millisecond-level alignment of multi-channel signals.

[0048] Data collection specifically includes the following:

[0049] Vocal Cord Vibration Signal Acquisition: A flexible electronic laryngeal patch adheres to the thyroid cartilage to collect highly sensitive mechanical vibration signals. The sensor chip has a sampling rate of up to 5kHz, with a dynamic range covering both normal vocalization and abnormal vibration. The built-in analog front-end (AFE) features adaptive gain and low-noise amplification, providing strong immunity to power frequency and environmental interference.

[0050] Laryngeal EMG signal acquisition: Surface EMG electrodes are attached to the thyroid cartilage and cricoid cartilage to accurately record the electrophysiological activity of laryngeal muscles (such as the cricoid muscle and vocal cord adductor muscles) during phonation. The EMG signal acquisition frequency is 1-2kHz, using real-time bandpass filtering (20-500Hz) and a 50Hz notch filter to eliminate power-frequency noise.

[0051] Speech acoustic data collection: A directional microphone close to the throat collects the original audio signal during the vocalization process in real time, with a sampling rate of 48kHz and 16-bit quantization, supplemented by an adaptive noise cancellation (ANC) algorithm to reduce interference from ambient speech and background noise.

[0052] Vocal cord vibration load data collection: The IMU sensor (including a three-axis accelerometer and a three-axis gyroscope) is integrated into the laryngeal patch to capture the subtle dynamics of the laryngeal muscles and trachea during phonation, and calculate the vibration frequency, acceleration, and instantaneous load changes.

[0053] Medical Image Capture: This device supports interfacing with portable ultrasound probes or electronic laryngoscopes, importing dynamic image sequences of vocal cord vibration and anatomical structures via standard DICOM protocols. The image acquisition terminal features a synchronous trigger function, enabling precise timing alignment with physiological signals.

[0054] All types of acquisition equipment are regularly synchronized using the master control terminal as the clock source, and all signal streams are timestamped. The system automatically detects packet loss and delay, and uses interpolation algorithms and sliding window technology to repair missing data and seamlessly stitch multi-channel data. It is equipped with real-time signal integrity detection, which can detect anomalies such as electrode detachment, microphone distortion, and IMU drift, and automatically prompt the user or adjust the acquisition parameters to achieve high reliability and adaptive operation. All collected data is encrypted with AES and uploaded to the cloud server and local backup in batches to ensure user privacy and data security. The system supports breakpoint resumption and dynamic bandwidth management, which is suitable for remote rehabilitation scenarios.

[0055] The feature extraction module is used to perform time-frequency domain analysis based on the acoustic information using a high-order formant detection algorithm to extract high-order acoustic features associated with the obstacle and construct a multimodal joint feature vector;

[0056] The operation process of the feature extraction module includes the following steps:

[0057] Synchronously pre-process the collected vocal cord vibration signals, laryngeal electromyographic signals, speech acoustic data and vocal cord dynamic images;

[0058] Based on the vocal cord vibration signal and speech acoustic data, a high-order formant detection algorithm is used to extract obstacle-related features such as high-order formant frequency, amplitude, frequency drift, and harmonic distortion;

[0059] Specifically, the vocal cord vibration signal and speech acoustic signal are combined, and a high-order formant automatic detection process is adopted. The signal is first divided into frames, and spectrum analysis is performed frame by frame to automatically identify the positions of multiple formant peaks such as F1-F5.

[0060] It should be noted that F1 to F5 (the first to fifth formant):

[0061] F1 (first resonance peak): Mainly determined by the degree of vocal tract opening, reflecting the degree of oral opening and chin height.

[0062] F2 (second resonance peak): closely related to the front and back position of the tongue, affecting the lightness and darkness of the vowel.

[0063] F3 (third resonance peak): involves the tip of the tongue, lip shape and local configuration of the vocal tract, and is more sensitive to individual vowels and pathological changes.

[0064] F4 and F5 (fourth and fifth resonance peaks): reflect the higher-order resonant structure of the vocal tract, are significantly sensitive to subtle anatomical abnormalities, vocal cord lesions and complex disorders, and are often used to distinguish pathological from healthy acoustic signals.

[0065] In the analysis of voice disorders, F1 to F3 are generally used to judge the basic acoustic state. High-order resonance peaks such as F4 and F5 often reveal deep voice disorders and abnormalities in fine vocal tract movements, and are an important basis for early detection and classification diagnosis.

[0066] The system processes the vocal cord vibration signal and speech acoustic data in frames, extracting the frequency, amplitude, peak width and other indicators of F1 to F5 in each frame. For F1 and F2, the focus is on analyzing their changes under different vowels and training movements to determine the basic vocalization state. For F3, F4, and F5, the focus is on analyzing the frequency drift, amplitude mutation, peak shape distortion of high-order peaks, and locating pathological features such as irregular vibration of the vocal cord edge, incomplete glottal closure, and upper airway abnormalities. Using a dynamic tracking algorithm, the curves of F1 to F5 changing over time during the vocalization cycle are monitored to automatically identify abnormal fluctuations, instantaneous loss, or noise coverage.

[0067] When abnormal changes in high-order resonance peaks (F3 to F5) are detected, the system automatically marks the segment as a "high-order resonance abnormality zone". If the drift range of F4 and F5 exceeds the normal threshold, or the peak amplitude weakens or disappears, it indicates that there may be risks such as muscle tone disorders and vocal cord microstructure lesions. For abnormalities in basic peaks such as F1 and F2 (such as a sharp drop in frequency and a sharp increase in amplitude), it indicates that there may be abnormalities in airflow control or vocal muscle activation. The system automatically generates a resonance peak trajectory diagram and visualizes the dynamic fluctuations of F1 to F5 to assist doctors in quickly locating the time period and type of disorder.

[0068] Due to differences in vocal tract structure, the baseline frequency and amplitude distribution of F1-F5 varies among patients. The system features adaptive adjustment and personalized reference intervals. Combined with harmonic energy distribution, it analyzes the proportion and distribution pattern of harmonic energy near each resonance peak, further quantifying the impact of vocal cord pathology on acoustic characteristics and assisting AI in fine-grained disorder classification.

[0069] In the multimodal joint feature vector finally output by the system, the dynamic features of F1 to F5 (such as mean, variance, volatility, and anomaly labeling) serve as key input features and participate in the reasoning and decision-making of AI models such as subsequent obstacle assessment and personalized training plan generation.

[0070] It automatically detects minor irregularities in frequency and amplitude during vocalization, highlighting areas of severe jitter to assist in identifying pathological conditions such as dystonia and asymmetric vocal cord vibration. It also assesses the energy contribution of each harmonic order to identify typical impairments, such as abnormal harmonic enhancement or weakening and sudden increases in noise components. It automatically outputs key metrics such as the harmonic-to-noise ratio and noise energy contribution to quantify the degree of vocal clarity and turbidity. Based on the historical data of different patients, the system can adaptively set individual thresholds for various abnormal indicators, reducing false positives and increasing sensitivity to early or mild voice disorders.

[0071] Based on the disorder-related characteristics, combined with laryngeal electromyographic signals and vocal cord dynamic images, by calculating electromyographic activation parameters and vocal cord kinematic indicators, abnormal patterns of vocal control and physiological movements are explored, and a feature fusion algorithm is used to construct a multimodal joint feature vector; the multimodal joint feature vector includes fundamental frequency, frequency jitter, amplitude jitter, harmonic-to-noise ratio, resonance peak frequency change rate, harmonic energy distribution, vocal cord vibration frequency and glottal closure characteristics.

[0072] Specifically, through multi-channel electromyographic acquisition, the system automatically analyzes the activation start and end points, activation duration, peak amplitude, and coordinated activation sequence of voice-related muscle groups. The system intelligently identifies abnormal patterns such as premature activation, delayed activation, and asymmetric activation, and automatically issues graded warnings based on the rehabilitation stage. Trends in electromyographic waveforms during voice training, such as a gradual decrease in signal amplitude or sudden, disordered high-frequency jitter, are automatically flagged as signs of muscle fatigue or spasm risk, advising physicians to adjust training strategies. The module features real-time electrode contact detection, noise level monitoring, and artifact detection. It automatically prompts users to readjust the equipment when encountering acquisition anomalies, improving data reliability. A deep learning segmentation model automatically locates the vocal cord edges, extracting the opening and closing trajectories and area change curves of the vocal cords throughout the phonation cycle. It automatically identifies phenomena such as incomplete vocal cord closure, short closure duration, or abnormal opening and closing rates. The motion trajectories of the left and right vocal cords are aligned and compared to quantify symmetry, assisting in the diagnosis of conditions such as unilateral vocal cord paralysis. All detected movement abnormalities such as delayed closure, local freezes, violent shaking, etc., are automatically highlighted on the image sequence and a list of abnormal events is generated, making it easier for doctors to quickly locate the lesion.

[0073] The system automatically compares EMG activation time with the actual start and end points of vocal cord movement, detecting any abnormal coupling between the two (e.g., EMG activation without a timely response from vocal cord movement), and exploring the complex abnormal relationships between nerves, muscles, and mechanics. Identified multimodal abnormalities are automatically labeled and attributed, such as "delayed EMG activation leading to delayed vocal cord closure," providing multi-dimensional evidence for AI assessment and physician diagnosis.

[0074] The noise disorder assessment module is used to construct a noise disorder assessment model based on the multimodal joint feature vector using transfer learning and graph neural network algorithms to classify and identify different types of voice disorders and generate voice disorder assessment results;

[0075] The operation process of the noise nuisance assessment module includes the following steps:

[0076] Based on the multimodal joint feature vector, each patient is used as a graph unit to construct a patient feature heterogeneous graph structure including vocal cord vibration feature nodes, laryngeal electromyography nodes, and dynamic image key frame nodes;

[0077] Specifically, each patient is regarded as a graph instance, and different types of features in the multimodal feature vector are assigned to different node types.

[0078] Vocal cord vibration node: includes features such as vocal cord vibration frequency, amplitude, and periodic jitter, with each dimension as a separate node or node attribute.

[0079] Laryngeal electromyographic nodes: including electromyographic activation level, muscle group coordination index, timing characteristics, etc.

[0080] Dynamic image keyframe node: It consists of automatically selected key motion frames, with characteristics such as vocal cord closure area, symmetry, opening and closing speed, etc.

[0081] The relationship between edges and nodes is as follows:

[0082] Temporal edge: connects similar nodes in adjacent time slices (for example, vocal cord vibration nodes in consecutive frames).

[0083] Synchronous edge: connects different modal nodes in the same time slice to strengthen the synchronous association between features.

[0084] Collaborative edges: If EMG abnormalities are found to be accompanied by imaging abnormalities, "collaborative abnormality" edges are automatically generated between related nodes to improve the ability to express cross-modal abnormalities.

[0085] Spatial and anatomical structure edges: For example, spatial edges can be added to nodes with similar anatomical structures in an image to support the expression of physical proximity.

[0086] Each node and edge is accompanied by metadata such as acquisition timestamp, device ID, acquisition quality, and feature score, providing a foundation for subsequent migration and traceability. This creates a heterogeneous graph with complex connections between acoustic, physiological, and imaging nodes, which can not only reflect anomalies in a single modality but also capture cross-modal dependencies.

[0087] A transfer learning algorithm based on joint maximum mean difference is used to map the graph embedding distribution of historical patient sample voice disorder data to the current patient feature heterogeneous graph structure, forming a unified graph feature space for transfer initialization.

[0088] Based on the unified graph feature space, a graph neural network algorithm is used to build a noise obstacle assessment model. This model integrates graph convolution and graph attention mechanisms to extract cross-modal dependencies in the obstacle feature graph and high-order interaction features between key obstacle nodes.

[0089] Specifically, the model first aggregates basic information from all nodes in the graph through a basic graph convolutional layer (GCN), capturing feature interactions within the local neighborhood. A multi-head graph attention layer (GAT) then focuses on key nodes and edges that contribute most to obstacle detection, automatically adjusting the weights of different node / edge types to achieve fine-grained cross-modal aggregation. A "node type embedding module" embedded in the model architecture automatically identifies and processes different node types (e.g., acoustic, electromyographic, and imaging), allowing different types of feature information within the same graph to participate in aggregation and decision-making according to an adaptive strategy, improving the ability to discern complex obstacle patterns. For cross-modal nodes with synchronized timestamps or causal connections (e.g., a specific electromyographic activation node and a corresponding vocal cord movement node), the model automatically constructs special "cooperative edges" and assigns higher weights to these edges, ensuring that cross-modal anomalies (e.g., delayed electromyographic activation leading to delayed vocal cord closure) are effectively perceived and discriminated by the model.

[0090] The model supports multi-level information flow, meaning that each node not only receives information from directly adjacent nodes but also captures deep dependencies with remote nodes through a multi-hop propagation mechanism (such as the spatial and temporal correlation between early EMG abnormalities and late acoustic impairments). The model also includes a built-in anomaly detection submodule that automatically labels anomalous nodes and edges through attention weight analysis and node embedding feature distribution, providing a basis for subsequent obstacle type determination and feature interpretation.

[0091] During the model training phase, the model combines multi-label manual annotation of historical samples with soft labels generated through transfer learning, using a hybrid loss function for joint optimization to improve the recognition of new types of obstacles and difficult boundary samples. To address practical situations such as small sample sizes and data imbalance, the model automatically adjusts the learning rate and employs an early stopping mechanism to prevent overfitting. After each inference, the model automatically generates a feature contribution weight table and node / edge heat map for obstacle determination, helping doctors understand the basis for the determination and the core abnormalities.

[0092] Combined with the graph embedding output, a multi-label classifier is trained to determine the disorder label for each patient sample in the graph, outputting prediction results and corresponding scores for multiple disorder types, including glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, and spectral instability.

[0093] It should be noted that the unified feature embedding of each patient output by the graph neural network is used to feed it into a multi-label classification network. This classification network usually uses a multi-layer perceptron (MLP), a residual fully connected network, or a lightweight Transformer structure, and automatically adapts the number of output nodes according to the number of disorder types. Each output node corresponds to a type of voice disorder (such as glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, spectral instability, etc.), and the output is the predicted probability and score of each disorder type, rather than mutually exclusive single selections, so that the system can identify concurrent disorders and the coexistence of multiple complex pathologies. It supports hierarchical labels for disorder types, such as "resonance disorder" which is further subdivided into subtypes such as "nasopharyngeal insufficiency". The system can automatically adapt the label structure according to the training data to meet the needs of refined diagnosis.

[0094] The classification output results are mapped to voice disorder assessment results, including disorder type, severity level, corresponding feature weight explanation and assessment confidence interval.

[0095] The construction process of the transfer learning algorithm includes the following steps:

[0096] Based on historical patient voice disorder data in the source domain and the acoustic information of current patients in the target domain, a heterogeneous graph structure containing vocal cord vibration nodes, myoelectric nodes, and dynamic image nodes is constructed as the original graph input of the source graph and target graph;

[0097] Based on the heterogeneous graph structure, a joint maximum mean difference transfer learning algorithm is used to perform joint matching of marginal and conditional distributions on the node embedding representations of the source graph and the target graph, mapping the two types of graph structures into a unified graph feature space;

[0098] The unified graph feature space is embedded as input and input into the graph neural network for obstacle pattern modeling. By jointly optimizing the transfer alignment loss and the multi-label classification loss, a noise obstacle recognition model is trained.

[0099] Specifically, the system has a built-in graph generation engine that performs unified node classification, attribute extraction, and edge relationship modeling on multimodal data (vocal cord vibration, electromyography, and imaging) collected from historical patient samples (source domain) and current patients (target domain), ensuring that the source and target graphs are structurally consistent, facilitating subsequent distribution alignment and feature migration. Batch normalization and node labeling: Normalize the feature data of all graph nodes (such as the amplitude, activation level, and dynamic trajectory of each node) to eliminate baseline deviations caused by different data batches and devices. At the same time, label each type of node (such as "electromyography node - left circular muscle", "imaging node - vocal cord closure", etc.) to ensure node comparability across samples.

[0100] Through multiple rounds of automated analysis, the system first evaluates the overall distribution of all nodes in the source and target graphs in the embedding space, and then performs conditional distribution analysis on nodes of the same type (such as nodes that all belong to the "electromyography abnormality" category) under each obstacle label category to capture the migration-sensitive features of "normal" and "obstacle" nodes. During the migration process, the system can perform distribution alignment for each modality or each type of node separately, gradually reducing the inter-domain distribution differences of a certain type of node features, improving the robustness of migration, and avoiding a single feature dominating the overall migration process. The transfer learning engine can automatically adjust the migration step size and alignment strength based on parameters such as the sample size, number of nodes, and label category of the source and target domains to ensure that effective migration effects can still be achieved in small sample / weak label scenarios. The system can automatically identify samples with abnormal distribution and obvious embedding offset during the migration process, and minimize the risk of migration failure through mechanisms such as data enhancement, feature perturbation, or soft label correction.

[0101] Ultimately, after multiple rounds of distribution alignment and node embedding between the source and target graphs, the multimodal, heterogeneous graph data for all patients is unified into a single high-dimensional feature space, ensuring input standardization and cross-domain generalization for subsequent graph neural network models. The system is equipped with a transfer alignment visualization tool, allowing doctors and engineers to visually view changes in graph structure distribution before and after transfer, allowing them to quickly assess transfer learning effectiveness and data quality.

[0102] Furthermore, the noise barrier assessment model is formulated as follows:

[0103]

[0104] Among them, S i represents the comprehensive score result of the i-th type of voice disorder; M represents the number of modes of the multimodal feature; K represents the number of key feature points of each mode; h represents the embedding vector; represents the embedding vector output by the kth node in the jth modality in the Lth layer of the graph neural network; γ j and β j represents the scoring intensity adjustment factor and bias term of the jth mode, which are used to control the response amplitude and baseline offset of the vocal cord vibration mode in the scoring; α ij represents the importance of the jth mode to the i-th type of voice disorder; represents the central mean embedding of all embedded nodes under the jth mode; λ j represents the modal regularization factor; represents the square of the Euclidean distance between the embedding feature of the kth node of the jth mode and the modal center; R jk represents the prediction residual recorded by the kth node of the jth modality during training; δ j Represents the adjustment factor of the residual term.

[0105] The calculation formula is as follows:

[0106]

[0107] in, Represents the output of node k under the acoustic mode, which may represent the change rate of the resonance peak frequency, harmonic-to-noise ratio, etc.

[0108] R jk The calculation formula is as follows:

[0109]

[0110] in, represents the true obstacle label of the corresponding node in the training sample; Represents the obstacle prediction result output by the graph neural network.

[0111] The personalized training module is used to generate a personalized rehabilitation training program based on the voice disorder assessment results, using a training strategy optimization algorithm based on reinforcement learning, and dynamically adjusting the vocal training content, frequency, and intensity through a strategy network; the training content includes breathing control, resonance regulation, and glottal closure;

[0112] The process of constructing the personalized training module includes the following steps:

[0113] Based on the voice disorder assessment results, extract the patient's corresponding disorder type label, severity level, and key acoustic features, and construct an individual state space as the input state of the reinforcement learning strategy network;

[0114] Specifically, the module automatically reads the noise disorder assessment results, including each disorder type label (such as glottal closure disorder, resonance disorder, etc.), severity level (such as mild, moderate, severe), and key acoustic features (such as F0, HNR, resonance peak frequency drift, closure area, etc.).

[0115] Combined with the patient's historical training response, subjective self-assessment (such as voice difficulty score, fatigue feedback), and physiological basic parameters (such as age, medical history).

[0116] The above multi-dimensional data are combined into an individual state space as the environment state input of the reinforcement learning policy network. For example:

[0117] State vector = [obstacle type 1, severity, F0 abnormality, closure characteristic abnormality, breathing control ability, improvement degree from last training, fatigue feedback…].

[0118] The state space can be dynamically expanded, such as adding patients' real-time physiological parameters (heart rate, respiratory rate) and training compliance indicators, to achieve all-round dynamic characterization.

[0119] A personalized training model is constructed using a reinforcement learning training strategy optimization algorithm. The improvement in vocal performance is used as a reward function to dynamically adjust the vocal training action sequence, which includes breathing control, resonance regulation, and glottal closure exercises.

[0120] Specifically, the training environment for each patient is adaptively constructed based on the type of disorder, severity, previous training response, and physiological characteristics. The system automatically sets the type, range, and restrictions of optional training movements based on the assessment report. For example, for patients with severe glottal closure disorders, only basic breathing and mild closure training are initially open to prevent overtraining. The personalized training model supports dynamic increase and decrease of movement types based on the patient's rehabilitation progress and real-time feedback, such as gradually introducing high-difficulty resonance adjustment training when performance improves, and temporarily blocking high-intensity training movements when performance declines.

[0121] Rewards are not only based on improvements in acoustic parameters detected by AI (such as increased F0, improved HNR, and reduced formant anomalies), but also incorporate multiple indicators such as electromyographic signals (such as improved activation coordination), imaging parameters (such as optimized vocal cord closure symmetry), and patient self-assessment (such as training comfort and confidence scores), achieving a "soft and hard combination" of reward signals. If the system detects physiological abnormalities (such as excessive electromyographic signal activation, abnormal fatigue after training, sore throat, etc.), a negative reward is immediately applied and the subsequent training plan is forced to adjust to ensure training safety and scientific research.

[0122] The reinforcement learning model automatically plans an action sequence based on the data output of each training round. For example, it begins with a five-minute breathing exercise, followed by a brief resonance adjustment, and finally a moderately intense glottal closure exercise. This allows for progressive training from easy to difficult, from basic to advanced. If acoustic parameters, electromyography, and subjective evaluations meet the standards for multiple consecutive rounds, the model automatically unlocks more advanced training content. For example, resonance training exercises will be expanded from single pronunciations to emotional intonation and complex syllable combinations.

[0123] Based on the action sequence, the policy network selects the optimal training action for the current state, and generates an immediate reward signal by combining the acoustic improvement after training and the patient's subjective feedback;

[0124] Specifically, the system collects the patient's current state vector (including the latest acoustic data, electromyographic feedback, training fatigue, historical completion status, etc.) in real time and inputs it into the strategy network. The network outputs the optimal next training action (such as "perform 3 minutes of abdominal breathing training") through forward reasoning and dynamically adjusts training parameters (such as training time and rest intervals). If the patient's subjective feedback is fatigue, the system prioritizes low-intensity training or arranges rest; if it detects fluctuations in the patient's state, the system temporarily adjusts the training content to avoid monotonous or high-intensity operations that may lead to a decrease in training compliance.

[0125] After the training movement is completed, the module synchronously collects multimodal data (such as the fundamental frequency and harmonic-to-noise ratio of the acoustic signal, the activation level of the electromyographic signal, and the vocal cord closure status of the imaging analysis), combined with the patient's subjective evaluation of the session (such as "Does this training feel easier or more stable in pronunciation?"). The system has a preset multi-dimensional reward weighting scheme: for example, acoustic improvement is assigned a weight of 40%, electromyographic and imaging improvements are each assigned a weight of 20%, and subjective feedback is assigned a weight of 20%. If any dimension shows a negative indicator (such as self-rated excessive fatigue or deterioration of acoustic parameters), the overall reward is reduced, forcing the model to adjust its next strategy.

[0126] Based on the reward signal, the policy network immediately updates its parameters (e.g., through Q-values ​​or policy gradients in reinforcement learning), ensuring that the model continuously adapts toward the optimal rehabilitation goal. All interactions, reward feedback, and training actions are automatically archived in the patient's personal training archive. The model periodically reviews and retrains to identify historically effective action combinations and recurring ineffective actions, further optimizing future action selection.

[0127] When a patient exhibits unexpected reactions or abnormal progress, the system automatically sends notifications to the doctor. Doctors can modify the training action library online, adjust reward parameters, and even temporarily switch to manual intervention mode, enhancing system safety and professionalism. Based on patient progress, the system regularly sends incentive messages (such as recovery medals and training target achievement reminders) to enhance patient participation and form an efficient positive feedback mechanism.

[0128] The strategy network and value function network are optimized in real time according to the reward signal, and the training content, frequency and intensity are dynamically updated to form a multi-round interactive learning process, and finally generate a personalized training plan instruction set.

[0129] Specifically, based on accumulated immediate rewards, backpropagation and policy gradient methods are used to update the policy network and value function network parameters in real time, improving the ability to adapt to diverse patient conditions. Experience replay and exploration mechanisms are used to ensure that the model not only utilizes historically optimal actions but also explores new actions to discover potentially efficient training methods.

[0130] The policy network can adjust parameters such as training movement type, training intensity (e.g., gradually increasing the difficulty of breathing control training), and frequency (automatically pushing daily / weekly training plans) in real time. If adverse reactions or slow progress occur, the system automatically reduces intensity or guides the patient to a more appropriate type of exercise.

[0131] Ultimately, the system summarizes the optimal action sequences and parameters from multiple rounds of interactive learning and generates a personalized training program instruction set for patients, including: daily training action arrangements (the order and frequency of breathing-resonance-closure); detailed parameters for each training (such as the number of times, number of sets, target acoustic improvement indicators); rest and evaluation interval recommendations, etc.

[0132] Furthermore, the formula of the personalized training model is as follows:

[0133]

[0134] Among them, J(π θ ) represents the policy network π θ The optimization objective function under the parameter θ; a represents the current vocal training action; π θ represents the training policy network; In the strategy π θ The expected operation under the distribution of the sampled action a; s t Indicates the state of step t; a t represents the training action taken in step t; π θ (a t |s t ) represents the output of the policy network in state s t Next select action a t probability; Indicates action a t The immediate reward signal after execution; ΔA t Indicates the improvement in key acoustic features before and after training; represents the vocal load index of the tth step, reflecting the degree of vocal cord fatigue or stress caused by the current training; ω1, ω2 and ω3 represent the weighted coefficients of the improvement enhancement term, strategy optimization term and load penalty term, respectively; T represents the time step, the total number of steps in each round of vocal training.

[0135] ΔA t The calculation formula is as follows:

[0136]

[0137] in, and They represent the values ​​of the jth acoustic feature before and after the tth step of training, such as frequency jitter, amplitude jitter, and harmonic-to-noise ratio; κ j represents the contribution weight of the jth characteristic index; N represents the number of high-order acoustic features, usually 7–10 items.

[0138] The collaborative control module is used to establish a real-time two-way data interaction channel between the doctor's side and the patient's terminal, and dynamically coordinate the rehabilitation task arrangements and adjustment suggestions between doctors and patients based on the user status and the personalized rehabilitation training plan; the collaborative control module establishes a remote synchronous interaction mechanism with the doctor's side, and the doctor's side can remotely adjust the personalized training plan based on the patient's training history data, current status and evaluation results.

[0139] Specifically, the collaborative control module consists of a patient terminal (e.g., an app) and a doctor-side management platform, communicating via an encrypted internet channel. The core of the module is a real-time task scheduling engine with built-in permission tiers and push notifications. All data, task scheduling, and communication processes are synchronized and backed up in real time via a cloud service platform, with a local caching strategy ensuring fault tolerance in the event of weak or disconnected networks.

[0140] The system supports encrypted synchronization and targeted push of training data, rehabilitation progress, real-time collected signals, AI assessment reports, and other information between patients and doctors. It uses WebSocket or MQTT protocols for low-latency, highly reliable data transmission. Events such as training completion, anomaly detection, and important data changes automatically trigger push notifications to doctors and patients, providing proactive reminders and anomaly warnings.

[0141] Based on the results of the obstacle assessment, historical training performance, and the rehabilitation goals set by the doctor, the system automatically generates a personalized training plan and daily rehabilitation tasks and assigns them to the patient. The plan includes parameters such as training content, time, frequency, and intensity. Doctors can adjust the content of the rehabilitation plan online, modify training goals, or issue new training movements based on real-time assessment data, patient subjective feedback, and system AI suggestions, and can synchronize them to the patient in real time. The module automatically records the training completion rate, patient compliance, and training effect scores, and feeds back to the doctor through a visual report for follow-up and decision-making reference.

[0142] Integrated with an encrypted video consultation channel, doctors can provide real-time remote guidance on voice training and corrective movements, supporting screen sharing and training demonstrations. This module supports diverse communication between doctors and patients, including instant messaging, voice, images, and documents. Doctors can regularly send training reminders, rehabilitation advice, incentives, or warnings to improve patient compliance. For scenarios where doctors are managing multiple patients simultaneously, the module supports task prioritization, batch adjustments, and abnormal grouping alerts, improving doctor efficiency.

[0143] All module data transmission utilizes SSL / TLS encryption, and core sensitive data such as medical records and training records are stored using AES256 encryption. Different users (doctors, patients, and administrators) have tiered access permissions. Audit logs are automatically generated for all remote operations, adjustment suggestions, and key task changes, facilitating subsequent traceability and compliance monitoring. Support for data anonymization, privacy mode switching, and one-click privacy authorization from the patient side ensures the secure and compliant flow of sensitive data in doctor-patient collaboration.

[0144] In summary, the present invention integrates a flexible electronic laryngeal patch, sEMG, a condenser microphone, an IMU, and an imaging interface to ensure that multiple signal channels such as acoustics, electromyography, and dynamic imaging are aligned at the millisecond level, significantly improving the timing consistency and data quality of multi-source signals and providing a solid foundation for subsequent AI analysis. Through the high-order resonance peak detection algorithm, key acoustic features such as F1\~F5 are deeply mined. Combined with dynamic tracking, spectrum anomaly recognition, and harmonic energy analysis, it can accurately capture early, latent lesions (such as resonance disorders and incomplete vocal cord closure), improving the early screening accuracy and typing sensitivity of voice disorders.

[0145] By constructing a heterogeneous "acoustic-electromyographic-imaging" feature map and introducing a graph neural network model, the system can effectively integrate cross-modal and cross-temporal vocal abnormality features, automatically discover potential collaborative abnormal relationships (such as abnormal closure caused by electromyographic delay), and visualize the judgment basis through the attention mechanism, thereby enhancing the clinical interpretability of AI diagnosis. The joint maximum mean difference (JMMD) algorithm in transfer learning is used to perform distribution alignment on historical data and current samples, alleviating the modeling difficulties caused by the scarcity of clinical samples and individual differences, and ensuring that the model has good adaptability and generalizability among different patient groups.

[0146] This invention builds a training strategy optimization network based on reinforcement learning, adjusts training content, frequency, and intensity in real time based on assessment results, self-assessment feedback, and physiological status, and generates multi-dimensional rewards through multimodal feedback signals, effectively improving the scientific nature, safety, and recovery efficiency of training. The collaborative control module establishes a low-latency, encrypted, two-way remote interaction mechanism that supports doctors to adjust training plans, push feedback, and conduct video guidance in real time; the system automatically records task completion and training responses, improving doctor management efficiency and patient execution motivation, and adapting to telemedicine scenarios such as home rehabilitation.

[0147] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A voice disorder assessment and training system based on AI algorithm, characterized by: It includes a data acquisition module, a feature extraction module, a noise barrier assessment module, a personalized training module and a collaborative control module that are sequentially connected in communication; The data acquisition module is used to collect acoustic information of patients with voice disorders through a sensor network, wherein the acoustic information includes vocal cord vibration signals, laryngeal electromyographic signals, speech acoustic data, vocal cord vibration load data, and vocal cord dynamic images; The feature extraction module is used to perform time-frequency domain analysis based on the acoustic information using a high-order formant detection algorithm to extract high-order acoustic features associated with the obstacle and construct a multimodal joint feature vector; The noise disorder assessment module is used to construct a noise disorder assessment model based on the multimodal joint feature vector using transfer learning and graph neural network algorithms to classify and identify different types of voice disorders and generate voice disorder assessment results; The personalized training module is used to generate a personalized rehabilitation training program based on the voice disorder assessment results, using a training strategy optimization algorithm based on reinforcement learning, and dynamically adjusting the vocal training content, frequency, and intensity through a strategy network; the training content includes breathing control, resonance regulation, and glottal closure; The collaborative control module is used to establish a real-time two-way data interaction channel between the doctor's terminal and the patient's terminal, and dynamically coordinate the rehabilitation task arrangement and adjustment suggestions between the doctor and the patient in combination with the user status and the personalized rehabilitation training plan.

2. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that: The operation process of the feature extraction module includes the following steps: Synchronously pre-process the collected vocal cord vibration signals, laryngeal electromyographic signals, speech acoustic data and vocal cord dynamic images; Based on the vocal cord vibration signal and speech acoustic data, a high-order formant detection algorithm is used to extract obstacle-related features such as high-order formant frequency, amplitude, frequency drift, and harmonic distortion; Based on the obstacle-related features, combined with laryngeal electromyographic signals and vocal cord dynamic images, the abnormal patterns of vocal control and physiological movements are explored by calculating electromyographic activation parameters and vocal cord kinematic indicators, and a feature fusion algorithm is used to construct a multimodal joint feature vector.

3. The voice disorder assessment and training system based on AI algorithm according to claim 2, characterized in that: The multimodal joint feature vector includes fundamental frequency, frequency jitter, amplitude jitter, harmonic-to-noise ratio, formant frequency change rate, harmonic energy distribution, vocal cord vibration frequency and glottal closure characteristics.

4. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that: The operation process of the noise nuisance assessment module includes the following steps: Based on the multimodal joint feature vector, each patient is used as a graph unit to construct a patient feature heterogeneous graph structure including vocal cord vibration feature nodes, laryngeal electromyography nodes, and dynamic image key frame nodes; A transfer learning algorithm based on joint maximum mean difference is used to map the graph embedding distribution of historical patient sample voice disorder data to the current patient feature heterogeneous graph structure, forming a unified graph feature space for transfer initialization. Based on the unified graph feature space, a graph neural network algorithm is used to build a noise obstacle assessment model. This model integrates graph convolution and graph attention mechanisms to extract cross-modal dependencies in the obstacle feature graph and high-order interaction features between key obstacle nodes. Combined with the graph embedding output, a multi-label classifier is trained to determine the disorder label for each patient sample in the graph, outputting prediction results and corresponding scores for multiple disorder types, including glottal closure dysfunction, laryngeal neuromuscular abnormalities, resonance disorders, and spectral instability. The classification output results are mapped to voice disorder assessment results, including disorder type, severity level, corresponding feature weight explanation and assessment confidence interval.

5. The voice disorder assessment and training system based on AI algorithm according to claim 4, characterized in that: The construction process of the transfer learning algorithm includes the following steps: Based on historical patient voice disorder data in the source domain and the acoustic information of current patients in the target domain, a heterogeneous graph structure containing vocal cord vibration nodes, myoelectric nodes, and dynamic image nodes is constructed as the original graph input of the source graph and target graph; Based on the heterogeneous graph structure, a joint maximum mean difference transfer learning algorithm is used to perform joint matching of marginal and conditional distributions on the node embedding representations of the source graph and the target graph, mapping the two types of graph structures into a unified graph feature space; The unified graph feature space is embedded as input and input into the graph neural network for obstacle pattern modeling. By jointly optimizing the transfer alignment loss and the multi-label classification loss, a noise obstacle recognition model is trained.

6. The voice disorder assessment and training system based on AI algorithm according to claim 4, characterized in that: The noise barrier assessment model is formulated as follows: Among them, S i represents the comprehensive score result of the i-th type of voice disorder; M represents the number of modes of the multimodal feature; K represents the number of key feature points of each mode; h represents the embedding vector; represents the embedding vector output by the kth node in the jth modality in the Lth layer of the graph neural network; γ j and β j represents the score strength adjustment factor and bias term of the j-th modality; α ij represents the importance of the jth mode to the i-th type of voice disorder; represents the central mean embedding of all embedded nodes under the jth mode; λ j represents the modal regularization factor; represents the square of the Euclidean distance between the embedding feature of the kth node of the jth mode and the modal center; R jk represents the prediction residual recorded by the kth node of the jth modality during training; δ j Represents the adjustment factor of the residual term.

7. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that: The process of constructing the personalized training module includes the following steps: Based on the voice disorder assessment results, extract the patient's corresponding disorder type label, severity level, and key acoustic features, and construct an individual state space as the input state of the reinforcement learning strategy network; A personalized training model is constructed using a reinforcement learning training strategy optimization algorithm. The improvement in vocal performance is used as a reward function to dynamically adjust the vocal training action sequence, which includes breathing control, resonance regulation, and glottal closure exercises. Based on the action sequence, the policy network selects the optimal training action for the current state, and generates an immediate reward signal by combining the acoustic improvement after training and the patient's subjective feedback; The strategy network and value function network are optimized in real time according to the reward signal, and the training content, frequency and intensity are dynamically updated to form a multi-round interactive learning process, and finally generate a personalized training plan instruction set.

8. The voice disorder assessment and training system based on AI algorithm according to claim 7, characterized in that: The formula of the personalized training model is as follows: Among them, J(π θ ) represents the policy network π θ The optimization objective function under the parameter θ; a represents the current vocal training action; π θ Represents the training policy network; In the strategy π θ The expected operation under the distribution of the sampled action a; s t Indicates the state of step t; a t represents the training action taken in step t; π θ (a t |s t ) represents the output of the policy network in state s t Next select action a t probability; Indicates action a t The immediate reward signal after execution; ΔA t Indicates the improvement in key acoustic features before and after training; represents the phonation load index at step t; ω1, ω2, and ω3 represent the weighted coefficients of the improvement enhancement term, strategy optimization term, and load penalty term, respectively; and T represents the time step.

9. The voice disorder assessment and training system based on AI algorithm according to claim 1, characterized in that: The collaborative control module establishes a remote synchronous interaction mechanism with the doctor side, and the doctor side can remotely adjust the personalized training plan based on the patient's training history data, current status and evaluation results.

Citation Information

Patent Citations

  • Detection method and system for pathological voice

    CN103730130A

  • Voice timbre disorder intelligent rehabilitation system based on ICF-RFT framework

    CN116831533A

  • Noise disease recognition method based on MFCC and CNN

    CN119724250A

  • System and application for evaluation of voice and speech disorders and speech-language therapy customized for parkinson patients

    KR102668964B1

  • System and method for pathological voice recognition and computer-readable storage medium

    US20230386504A1