Multimodal learning-based laryngeal rehabilitation motion quantification evaluation system and method

By combining multimodal data collected by millimeter-wave radar and microphones, and using a deep learning neural network model to assess laryngeal rehabilitation movements, the subjectivity and incompleteness of traditional assessment methods are solved, enabling non-contact, comprehensive quantitative assessment of laryngeal rehabilitation movements and the provision of professional indicators.

CN120827371BActive Publication Date: 2025-12-05XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511323776.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-05
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Traditional laryngeal rehabilitation assessment methods lack multimodal data fusion, making it impossible to fully capture the complex dynamic changes in the laryngeal rehabilitation process. Furthermore, the lack of objective quantitative assessment indicators leads to significant subjective differences in the assessment of rehabilitation effectiveness.

Method used

A millimeter-wave radar module is used to capture three-dimensional point cloud data of the larynx and a standard microphone is used to collect acoustic signals. Combined with a deep learning neural network model, a point cloud processing module, an acoustic signal processing module, and a multimodal motion feature fusion module are used to achieve quantitative assessment of larynx rehabilitation movements.

Benefits of technology

It enables non-contact, comprehensive quantitative assessment of laryngeal rehabilitation movements, provides professional assessment indicators, improves the accuracy and objectivity of assessment, and adapts to individual differences and changes in the characteristics of different rehabilitation stages among patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120827371B_ABST
    Figure CN120827371B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of medical rehabilitation, in particular to a system and method for quantitatively evaluating laryngeal rehabilitation movements using multi-modal learning technology, the system comprising a millimeter wave radar module, a standard microphone and a deep learning model, for evaluating laryngeal rehabilitation movements, the millimeter wave radar collecting three-dimensional point cloud data of the larynx, the microphone acquiring audio signals, the deep learning model comprising a point cloud processing module, an acoustic processing module and a multi-modal fusion module, the point cloud module extracting cartilage movement parameters, the acoustic module extracting audio features, the fusion module realizing multi-modal feature fusion through topological perception mapping, self-attention mechanism and probability field reconstruction, and outputting quantitative evaluation results, the present application realizes objective quantitative evaluation of laryngeal rehabilitation movements, and improves the accuracy and comprehensiveness of the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical rehabilitation, in particular to a system and method for quantitatively evaluating laryngeal rehabilitation movements using multi-modal learning technology. BACKGROUND

[0002] Laryngeal rehabilitation therapy is of great significance in the recovery after laryngeal surgery, treatment of laryngeal dysfunction and voice disorders. Traditional evaluation of laryngeal rehabilitation movements mainly relies on the subjective judgment of doctors, lacking objective and quantitative evaluation criteria, resulting in large subjective differences in rehabilitation effect evaluation.

[0003] In the prior art, there are some evaluation methods for laryngeal rehabilitation, such as laryngeal observation based on endoscopy and voice function evaluation based on acoustic analysis. However, these methods can only evaluate laryngeal function from a single dimension and cannot fully capture the complex dynamic changes in the laryngeal rehabilitation process. The method based on endoscopy is invasive and causes discomfort to the patient, while the method based solely on acoustic analysis cannot directly observe the movement of the laryngeal structure.

[0004] Professional laryngeal rehabilitation training usually includes four basic methods: breathing training, relaxation training, resonance voice training and airflow voice training, which are aimed at different structures and functions of the larynx. Among them, breathing training is mainly aimed at the stability of the cricoid cartilage and airway, relaxation training is mainly aimed at the coordinated control of the thyroid cartilage and hyoid bone, resonance voice training mainly involves the adjustment of the epiglottic cartilage and laryngeal cavity, and airflow voice training mainly focuses on the movement of the arytenoid cartilage and the closing function of the vocal cords.

[0005] In addition, there is a lack of methods in the prior art to effectively fuse laryngeal structure movement data with acoustic data, which cannot fully utilize the complementarity of multi-modal data to improve the accuracy and comprehensiveness of evaluation. At the same time, most of the existing evaluation methods are qualitative evaluation, lacking precise quantitative indicators for different rehabilitation training methods, making it difficult to objectively measure the progress of rehabilitation.

[0006] Therefore, there is an urgent need for a laryngeal rehabilitation movement evaluation system and method that can realize multi-modal data fusion, fully cover key laryngeal structures, and provide objective quantitative evaluation indicators for different rehabilitation training methods. SUMMARY

[0007] The purpose of the present application is to provide a multi-modal learning laryngeal rehabilitation movement quantitative evaluation system and method, which realizes objective quantitative evaluation of laryngeal rehabilitation movements by fusing three-dimensional point cloud data collected by a millimeter wave radar and acoustic signals collected by a microphone.

[0008] The present application proposes a multi-modal learning laryngeal rehabilitation movement quantitative evaluation system, which comprises:

[0009] A millimeter wave radar module is used to capture a three-dimensional point cloud data stream of a human throat region, which includes thyroid cartilage, cricoid cartilage, arytenoid cartilage, epiglottis cartilage, and hyoid key structures.

[0010] A standard microphone is used to collect throat rehabilitation instruction and vocalization training audio signals.

[0011] A deep learning neural network model is used to receive the three-dimensional point cloud data stream and the audio signals, and obtain a quantitative evaluation result of a throat rehabilitation action. The deep learning neural network model includes three sub-networks: a point cloud processing module, an acoustic signal processing module, and a multi-modal action feature fusion module.

[0012] The point cloud processing module is used to process the three-dimensional point cloud data stream to obtain three-dimensional motion parameters of each key cartilage structure of the throat.

[0013] The acoustic signal processing module is used to extract features from the audio signals.

[0014] The multi-modal action feature fusion module is used to receive the three-dimensional motion parameters output by the point cloud processing module and the features output by the acoustic signal processing module, and use a multi-modal action feature fusion algorithm to fuse at the feature level to obtain the quantitative evaluation result of the throat rehabilitation action. The multi-modal action feature fusion module includes a topologically aware multi-dimensional feature space mapping structure, a self-attention dynamic weight mechanism on a matrix manifold, and an adaptive feature reconstruction system based on a probability field.

[0015] As a preferred embodiment, the millimeter wave radar module includes multiple millimeter wave radars for obtaining three-dimensional point cloud data streams from multiple perspectives covering each key cartilage structure of the throat.

[0016] As a preferred embodiment, the point cloud processing module takes three-dimensional point cloud data as input, first removes external background noise and samples to obtain corresponding throat region point cloud data, then extracts point clouds of the thyroid cartilage, cricoid cartilage, arytenoid cartilage, epiglottis cartilage, and hyoid from the region point cloud data, and finally segments these parts.

[0017] As a preferred embodiment, the point cloud processing module uses convolution blocks, pooling blocks, up-sampling blocks, and multi-layer perception blocks to extract three-dimensional motion parameter features, including: using a K-neighborhood screening algorithm for background noise removal and sampling; extracting point clouds of each key cartilage structure of the throat from the region point cloud data; using a point cloud key point detection algorithm to automatically label point markers of each cartilage key part; obtaining coordinate data of each cartilage key part; slicing the point cloud of the throat region to obtain each frame of point cloud sequence, labeling key points of each cartilage, and using key point displacement interpolation algorithm and key point coordinate conversion algorithm to obtain the position and motion trajectory of each cartilage structure.

[0018] As preferred, the acoustic signal processing module extracts motion acoustic features by using convolution blocks, pooling blocks, up-sampling blocks and multi-layer perception blocks, including: extracting high-frequency spectrum features and correlation features of the original audio signal; extracting breathing training related acoustic features, relaxation training related acoustic features, resonance voice training related acoustic features and airflow sound production training related acoustic features; and performing deep fusion on the extracted features by using an LSTM block and a time series strategy.

[0019] As preferred, the topologically perceived multi-dimensional feature space mapping structure includes:

[0020] A multi-dimensional feature space construction unit is configured to construct a point cloud feature space and an acoustic feature space.

[0021] A local structure preserving mapping unit is configured to design a feature mapping function to preserve the topological structure of the original feature space.

[0022] A manifold alignment unit is configured to define the distance between different feature spaces by using a Riemannian metric to realize the alignment of two manifolds.

[0023] A unified feature representation unit is configured to map the aligned features to a unified feature representation space.

[0024] As preferred, the self-attention dynamic weight mechanism on the matrix manifold includes:

[0025] A feature matrix construction unit is configured to organize the unified feature representation into a feature matrix.

[0026] A correlation matrix calculation unit is configured to calculate the inner product of the feature matrix to obtain a correlation matrix.

[0027] A multi-head self-attention unit is configured to decompose the correlation matrix into multiple sub-matrices to calculate attention weights respectively.

[0028] A matrix manifold optimization unit is configured to define a gradient on the matrix manifold to guide the optimization of the weight matrix.

[0029] As preferred, the adaptive feature reconstruction system based on the probability field includes:

[0030] A probability field modeling unit is configured to regard the feature distribution as a probability field to establish a probability density function of the feature.

[0031] An information entropy maximization unit is configured to calculate the information entropy of the feature distribution and maximize the information entropy under the constraint of preserving useful information.

[0032] A conditional random field construction unit is configured to establish a conditional dependency graph between features to describe the relationship between features.

[0033] An adaptive feature reconstruction unit is configured to dynamically adjust the distribution and representation of the features according to the characteristics of the probability field.

[0034] Preferably, the quantitative evaluation result of the laryngeal rehabilitation action includes: an action correctness discrimination result, an action start time, an action execution time, and an evaluation analysis result of an action correct execution time proportion; specialized evaluation indexes for four basic rehabilitation methods of breathing training, relaxation training, resonant voice training, and airflow sound production training; and a synergistic relationship analysis result between each laryngeal cartilage structure movement and the rehabilitation training method.

[0035] The multi-modal learning laryngeal rehabilitation action quantitative evaluation method comprises the following steps:

[0036] The three-dimensional point cloud data stream of multiple viewing angles collected by the millimeter wave radar module and the audio signal collected by the standard microphone are acquired, and then input into the trained deep learning neural network model, so as to obtain the laryngeal rehabilitation action quantitative evaluation result through the trained deep learning neural network model;

[0037] The deep learning neural network model is obtained by training a multi-modal data set, and the multi-modal data set includes a three-dimensional point cloud data set and an acoustic signal set;

[0038] The deep learning neural network model comprises three sub-networks: a point cloud processing module, an acoustic signal processing module, and a multi-modal action feature fusion module;

[0039] The point cloud processing module is configured to obtain three-dimensional motion parameters of key parts of the thyroid cartilage, the cricoid cartilage, the arytenoid cartilage, the epiglottis cartilage, and the hyoid bone;

[0040] The acoustic signal processing module extracts features from the audio signal, including acoustic features related to four basic rehabilitation methods of breathing training, relaxation training, resonant voice training, and airflow sound production training;

[0041] The multi-modal action feature fusion module fuses the three-dimensional motion parameters obtained by the point cloud processing module and the results of the acoustic signal processing module at the feature level, thereby outputting an evaluation analysis result including an action correctness discrimination result, an action start time, an action execution time, and an action correct execution time proportion, and specialized evaluation indexes for the four basic rehabilitation methods;

[0042] The multi-modal action feature fusion module includes a multi-dimensional feature space mapping structure through topological perception, a self-attention dynamic weight mechanism on a matrix manifold, and an adaptive feature reconstruction system based on a probability field, to realize deep fusion of point cloud features and acoustic features.

[0043] The present application has the following advantages:

[0044] 1. Through multi-modal data fusion, the complementary information of point cloud data and acoustic data is comprehensively utilized to capture the characteristics of laryngeal rehabilitation actions, improve the accuracy and comprehensiveness of evaluation.

[0045] 2. Non-invasive millimeter wave radar technology is used to obtain the three-dimensional motion parameters of each key cartilage structure of the larynx (including thyroid cartilage, ring cartilage, arytenoid cartilage, epiglottis cartilage and hyoid bone), avoiding the invasiveness of traditional endoscopy and improving patient comfort.

[0046] 3. The topological perception multi-dimensional feature space mapping structure, the self-attention dynamic weight mechanism on the matrix manifold and the adaptive feature reconstruction system based on the probability field are innovatively proposed to realize the deep fusion of point cloud features and acoustic features, and significantly improve the performance of the system.

[0047] 4. For four basic rehabilitation methods of breathing training, relaxation training, resonance voice training and airflow sound training, professional quantitative evaluation indicators are provided to provide objective reference for doctors and facilitate tracking of rehabilitation progress.

[0048] 5. The system can adapt to individual differences of different patients and characteristic changes of different rehabilitation stages, has good robustness and adaptability, and supports accurate adjustment of personalized rehabilitation programs. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 The overall architecture diagram of the multi-modal learning laryngeal rehabilitation action quantitative evaluation system of the application is shown in the figure;

[0050] Figure 2 The processing flowchart of the point cloud processing module of the application is shown in the figure;

[0051] Figure 3 The processing flowchart of the acoustic signal processing module of the application is shown in the figure;

[0052] Figure 4 The structure diagram of the multi-modal action feature fusion module of the application is shown in the figure;

[0053] Figure 5 The topological perception multi-dimensional feature space mapping structure of the application is shown in the figure;

[0054] Figure 6 The self-attention dynamic weight mechanism on the matrix manifold of the application is shown in the figure;

[0055] Figure 7 The adaptive feature reconstruction system based on the probability field of the application is shown in the figure;

[0056] Figure 8 The flowchart of the multi-modal learning laryngeal rehabilitation action quantitative evaluation method of the application is shown in the figure. Detailed Implementation

[0057] Please refer to Figure 1 - Figure 8 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0058] Reference Figure 1 The multimodal learning-based quantitative assessment system for laryngeal rehabilitation movements provided by this invention includes a millimeter-wave radar module 1, a standard microphone 2, and a deep learning neural network model 3. The deep learning neural network model 3 comprises three sub-networks: a point cloud processing module 31, an acoustic signal processing module 32, and a multimodal movement feature fusion module 33.

[0059] The millimeter-wave radar module 1 is used to capture three-dimensional point cloud data streams of the human laryngeal region, including key structures such as the thyroid cartilage, cricoid cartilage, arytenoid cartilage, epiglottis, and hyoid bone. In a preferred embodiment of the present invention, the millimeter-wave radar module 1 uses a 77GHz band millimeter-wave radar with a sampling frequency of 30Hz and a spatial resolution of 2mm, enabling precise capture of minute motion changes in the laryngeal region.

[0060] Standard microphone 2 is used to acquire audio signals of laryngeal rehabilitation instructions and vocal training sounds. Preferably, a standard microphone with a frequency response range of 10Hz-22kHz is used, with a sampling rate set to 24kHz and 24-bit quantization to ensure the quality of acoustic signal acquisition, especially to capture high-frequency harmonics related to resonance and low-frequency noise components related to airflow vocalization.

[0061] The deep learning neural network model 3 receives the three-dimensional point cloud data stream collected by the millimeter-wave radar module 1 and the audio signal collected by the standard microphone 2. Through the collaborative work of the point cloud processing module 31, the acoustic signal processing module 32 and the multimodal motion feature fusion module 33, the model finally outputs the quantitative evaluation results of the laryngeal rehabilitation movements.

[0062] In one embodiment of the present invention, the millimeter-wave radar module 1 includes multiple millimeter-wave radars, preferably 5 to 7, which capture three-dimensional point cloud data of the throat region from different angles to obtain a three-dimensional point cloud data stream covering multiple perspectives of key cartilage structures in the throat. Data acquisition from multiple perspectives helps eliminate data blind spots caused by a single perspective, improving the completeness and accuracy of the point cloud data.

[0063] Reference Figure 2, the point cloud processing module 31 takes the three-dimensional point cloud data as input, and mainly completes the following processing steps: first, remove the external background noise and sample to obtain the corresponding laryngeal region point cloud data, then extract the point cloud of the thyroid cartilage, the cricoid cartilage, the arytenoid cartilage, the epiglottis cartilage and the hyoid part from the region point cloud data, and finally segment these parts.

[0064] Specifically, the point cloud processing module 31 extracts three-dimensional motion parameter features using convolution blocks, pooling blocks, up-sampling blocks and multi-layer perception blocks. In a preferred embodiment of the present application, the processing flow of the point cloud processing module 31 includes:

[0065] 1) Background noise removal and sampling using K-neighborhood screening algorithm. The K-neighborhood screening algorithm is based on the K nearest neighbors of each point in the point cloud, and determines whether the point is a noise point by calculating the local density and distance threshold of the point. In this embodiment, the K value is preferably set to 15, and the distance threshold is set to 5mm. The algorithm can be expressed as:

[0066] ,

[0067] wherein, is the local density feature of point , and represents the average distance of point and its K nearest neighbors; is the number of neighborhood points, which is 15 in this embodiment; represents the set of K nearest neighbors of point ; is the point in the point cloud being processed; is the neighborhood point of ; represents the Euclidean distance between point and point , i.e. , wherein , , are the x, y, z coordinates of point , , are the x, y, z coordinates of point , , When is greater than a predetermined threshold (5mm in this embodiment), point is determined to be a noise point and is removed. The selection of 5mm as the threshold is based on experimental data analysis, and this value can effectively distinguish between laryngeal structures and background noise.

[0068] 2) Extract point clouds of each laryngeal cartilage structure from the region point cloud data. First, the thyroid cartilage is extracted by clustering segmentation, and then the cricoid cartilage, arytenoid cartilage, epiglottis cartilage and hyoid bone are segmented in turn based on anatomical positional relationship and morphological characteristics. In this embodiment, DBSCAN (density-based spatial clustering algorithm) is used for clustering segmentation, and the density threshold ε is set to 3 mm and the minimum point number MinPts is set to 10. The DBSCAN algorithm selects these parameters because experience shows that the point cloud density of each cartilage region in the larynx is high, and there is a clear spatial separation between them. At the same time, considering the anatomical positional relationship of different cartilage structures, R (k is the number of cartilages) to enhance the accuracy of segmentation.

[0069] 3) Use point cloud key point detection algorithm to automatically label the point marks of the key parts of the thyroid cartilage and hyoid bone. In this embodiment, the ISS (Intrinsic Shape Signatures) key point detection algorithm is used, which extracts key points based on the local geometric features of the point cloud. The algorithm calculates the eigenvalues of the covariance matrix of each point, and selects the key points according to the ratio relationship of the eigenvalues. The algorithm is represented as:

[0070] ,

[0071] ,

[0072] wherein 、 、 are the eigenvalues of the covariance matrix of the local region of the point cloud, sorted in descending order; represents the ratio of the second largest eigenvalue to the largest eigenvalue; represents the ratio of the smallest eigenvalue to the largest eigenvalue; represents the ratio of the smallest eigenvalue to the second largest eigenvalue. When and , (in this embodiment , ), the point is selected as a key point. and are threshold values set according to experience, used to control the selection criteria of key points. The values 0.8 and 0.7 are determined by experiments and can extract a sufficient number of key points with reasonable distribution in the thyroid cartilage and hyoid bone regions.

[0073] 4) Obtain the coordinate data of each key part of the cartilage. For the thyroid cartilage, extract the coordinates of the midpoint of the upper edge, the midpoint of the lower edge, and the edge points on both sides; for the cricoid cartilage, extract the characteristic points of the ring structure of the upper and lower edges; for the arytenoid cartilage, extract the coordinates of the apex and the base; for the epiglottis, extract the coordinates of the tip and the base; for the hyoid bone, extract the coordinates of the body and the large angle. These coordinate data are used to calculate the following key motion parameters:

[0074] The tilt angle and vertical displacement of the thyroid cartilage; the stability index and the cricothyroid joint activity of the cricoid cartilage; the adduction / abduction angle of the arytenoid cartilage; the lifting angle of the epiglottis; the vertical displacement and the forward / backward displacement of the hyoid bone; the vertical displacement calculation formula is:

[0075]

[0076] wherein, is the vertical displacement distance, in millimeters; is the z-coordinate when the structure is lifted to the highest point, in millimeters; is the z-coordinate when the structure is in the lowest position, in millimeters; represents the absolute value.

[0077] The angle change calculation formula is:

[0078]

[0079] wherein, is the angle change, in degrees; and are the direction vectors in two position states, respectively; represents the vector dot product; represents the module length of the vector ; represents the inverse cosine function.

[0080] 5) Slice the point cloud of the laryngeal region to obtain each frame of point cloud sequence, calibrate the key points of each cartilage, and then use the key point displacement interpolation algorithm and the key point coordinate conversion algorithm to obtain the position and motion trajectory of each cartilage structure. The displacement interpolation algorithm uses a cubic spline interpolation method to ensure the smoothness of the motion trajectory. The motion trajectory can be represented as a time sequence:

[0081]

[0082] wherein, is the time sequence of the motion trajectory; represents the three-dimensional coordinates of the key point of the cartilage at time t, in millimeters; is the time variable, in seconds; ​​Total time of action duration, unit: second. The acquisition of motion trajectory is of great significance for evaluating the smoothness and coordination of laryngeal rehabilitation action.

[0083] Through the above processing steps, the point cloud processing module 31 can extract the three-dimensional motion parameters of each laryngeal cartilage key position from the original three-dimensional point cloud data, providing basic data for subsequent multi-modal fusion.

[0084] Referring to Figure 3 , the acoustic signal processing module 32 extracts features from the audio signals collected by the standard microphone 2. In the preferred embodiment of the present application, the acoustic signal processing module 32 uses convolution blocks, pooling blocks, up-sampling blocks and multi-layer perception blocks to extract motion acoustic features.

[0085] Specifically, the processing flow of the acoustic signal processing module 32 includes:

[0086] 1) Extract high-frequency spectral features of the original audio signal, including Mel-Spectrogram, Chromgram and AudioSet features. Mel-Spectrogram converts sound signals to the Mel frequency scale, which is closer to the perceptual characteristics of the human ear. The calculation formula is:

[0087] ,

[0088] wherein, is the frequency corresponding to the Mel scale value, unit: mel; is the linear frequency, unit: Hz; is the natural logarithm function. This formula converts the linear frequency to the Mel frequency scale that is more consistent with human ear perception. The calculation of Mel-Spectrogram is realized by short-time Fourier transform (STFT) and Mel filter bank:

[0089] ,

[0090] wherein, is the output of the nth frame of the mth Mel filter, representing the energy value of Mel-Spectrogram; is the time frame index; is the Mel filter index; is the frequency index; is the FFT point number, which is set to 512 in this embodiment; is the STFT result of the nth frame, representing the complex value in time-frequency domain; represents the power spectrum; Let be the frequency response of the m-th Mel filter, which is a triangular filter. In this embodiment, the window length is set to 25ms, the frame shift to 10ms, and the number of Mel filters to 40. These parameters are commonly used settings in speech signal processing and can effectively capture the characteristics of human voices.

[0091] 2) Extract relevant features from the original audio signal, including Energetic, Shimmer, and Jitter features. These features reflect the stability and quality of the sound and are important for assessing the effectiveness of laryngeal rehabilitation. Shimmer represents the change in sound amplitude, calculated using the following formula:

[0092] ,

[0093] in, It is a dimensionless quantitative indicator of changes in sound amplitude. The total number of sound cycles; The peak amplitude of the i-th sound cycle; For the first Peak amplitude of each sound cycle; It is a logarithmic function with base 10; This indicates the absolute value. A smaller Shimmer value indicates a more stable sound amplitude. Jitter represents the change in sound frequency, and its calculation formula is:

[0094] ,

[0095] in, A quantitative indicator of sound frequency variation, measured in seconds; The total number of sound cycles; For the first The length of a sound cycle, in seconds; For the first The length of a sound cycle, in seconds; This indicates taking the absolute value. The smaller the value, the more stable the sound frequency. These parameters are important in laryngeal rehabilitation assessments, reflecting the stability and clarity of the sound.

[0096] Extracting specialized acoustic features for different rehabilitation training methods:

[0097] a. Acoustic characteristics related to breathing training: including inspiratory / expiratory ratio, airflow stability coefficient, and duration of breathing support. The formula for calculating the airflow stability coefficient is:

[0098] ,

[0099] in, is the airflow stability coefficient, whose value ranges from 0 to 1, and the larger the value, the more stable the airflow; is the standard deviation of sound energy; is the mean value of sound energy.

[0100] b. Relaxation training related acoustic features: including vocal cord tension index, voice onset time and vocal cord relaxation. The vocal cord tension index is calculated by the formula:

[0101] ,

[0102] wherein, is the vocal cord tension index, and the smaller the value, the more relaxed the vocal cord; is the weighting coefficient, which is set to 1.5 in this embodiment, for balancing the different contributions of Shimmer and Jitter to tension.

[0103] c. Resonant voice training related acoustic features: including resonance peak frequency distribution, high frequency energy ratio and sound brightness index. The high frequency energy ratio is calculated by the formula:

[0104] ,

[0105] wherein, is the high frequency energy ratio, indicating the proportion of high frequency component in total energy; is the energy at frequency is the energy at frequency is the high frequency threshold, which is set to 3000Hz in this embodiment; and are the lowest and highest frequencies of the spectrum, respectively.

[0106] d. Airflow phonation training related acoustic features: including vocal cord closure quotient, breathy voice ratio and voice onset articulation index. The vocal cord closure quotient is calculated by the formula:

[0107] ,

[0108] wherein, is the vocal cord closure quotient, and the smaller the value, the better the vocal cord closure effect; is the vocal cord open time, in seconds; is the total vocal cord vibration period time, in seconds.

[0109] 4) Deeply fuse the extracted features using LSTM blocks and time series strategy. LSTM (Long Short-Term Memory Network) can effectively capture the time sequence dependence of acoustic features and improve the quality of feature representation. The update formula of LSTM unit is:

[0110] ,

[0111] ,

[0112] ,

[0113] ,

[0114] ,

[0115] ,

[0116] wherein, is the forget gate, controlling the degree of discarding the cell state, taking the value range [0, 1]; is the input gate, controlling the degree of updating the cell state, taking the value range [0, 1]; is the output gate, controlling the degree of the cell state affecting the hidden state, taking the value range [0, 1]; is the cell state, storing long-term memory information; is the candidate cell state, representing new information of the current input; is the hidden state, representing the output of the LSTM unit; is the hidden state of the last moment; is the input feature of the current moment; , , are the weight matrices of the forget gate, the input gate, the candidate cell state and the output gate, respectively; , , are the corresponding bias vectors, respectively; is the sigmoid activation function, mapping the input to the interval [0, 1]; is the hyperbolic tangent activation function, mapping the input to the interval [-1, 1]; * represents the element product (Hadamard product), that is, the corresponding elements are multiplied. In this embodiment, the hidden layer dimension of the LSTM is set to 128, and the number of layers is 2. These parameter settings are obtained through experimental optimization, which can effectively capture the time sequence relationship of acoustic features.

[0117] 5) Construct a rehabilitation training type recognition module to automatically recognize the current rehabilitation training type (breathing training, relaxation training, resonance voice training or airflow sound training) based on acoustic features. On the basis of LSTM, a training type feature embedding vector

[0118] ,

[0119] wherein, For training type feature embedding vectors, the dimension is the same as the hidden state. For adaptive weight coefficients, the degree of influence of the training type features on the output is controlled.

[0120] Through the above processing steps, the acoustic signal processing module 32 can extract rich acoustic features from the original audio signal, especially professional features for different rehabilitation training methods, providing another dimension of data for subsequent multi-modal fusion.

[0121] Referring to Figure 4 , the multi-modal action feature fusion module 33 is the core innovative part of the present application, which is used to receive the three-dimensional motion parameters output by the point cloud processing module 31 and the features output by the acoustic signal processing module 32, and utilize a multi-modal action feature fusion algorithm to perform fusion at the feature level to obtain a quantitative evaluation result of the laryngeal rehabilitation action.

[0122] The multi-modal action feature fusion module 33 includes three key components: a topologically aware multi-dimensional feature space mapping structure 331, a self-attention dynamic weight mechanism 332 on a matrix manifold, and an adaptive feature reconstruction system 333 based on a probability field. These three components form a complete feature fusion solution that can effectively fuse point cloud features and acoustic features to improve the overall performance of the system.

[0123] Referring to Figure 5 , the topologically aware multi-dimensional feature space mapping structure 331 includes a multi-dimensional feature space construction unit 3311, a local structure preserving mapping unit 3312, a manifold alignment unit 3313, and a unified feature representation unit 3314.

[0124] The multi-dimensional feature space construction unit 3311 is used to construct a point cloud feature space and an acoustic feature space. The point cloud feature space is represented as wherein is a point cloud feature matrix; represents a real number set; is a feature dimension (set to 64 in this embodiment); is the number of point cloud feature samples. The acoustic feature space is represented as wherein is an acoustic feature matrix; is a feature dimension (set to 128 in this embodiment); is the number of acoustic feature samples. The selection of the feature dimension is based on experimental analysis, and these dimensions can adequately express the information of each modality.

[0125] In addition, for the four basic rehabilitation training methods, a dedicated feature subspace is constructed:

[0126] ,

[0127] ,

[0128] ,

[0129] ,

[0130] wherein, , , and represent the feature subspaces of breathing training, relaxation training, resonant voice training and airflow phonation training, respectively; , , and are the feature dimensions of each subspace, which are set to 64 in this embodiment; , , and are the sample numbers of each subspace.

[0131] The local structure preserving mapping unit 3312 is configured to design a feature mapping function to preserve the topology of the original feature space. For the point cloud feature space , the mapping function is designed, where is the mapping function of the point cloud feature; is the uniform feature space; for the acoustic feature space , the mapping function is designed, where is the mapping function of the acoustic feature. The design of the mapping function is based on the local linear embedding (LLE) algorithm, which first constructs a local neighborhood relationship for each data point, and then realizes dimension reduction by preserving these local relationships.

[0132] For each point in the point cloud feature space, first find its nearest neighbors , then calculate the reconstruction weight such that can be reconstructed by the linear combination of its neighbors:

[0133] ,

[0134] s.t. ,

[0135] wherein, is the point in the point cloud feature space; is the nearest neighbor of . To reconstruct the weights, the contribution of the nearest neighbor point To ; represents the square of the Euclidean distance; represents the minimization of the objective function; s.t. represents the constraint condition; represents the constraint that the sum of the weights is 1. Preferably, is set to 15, and the optimal value is determined by cross-validation. For each point in the acoustic feature space, the same method is used to calculate the reconstruction weight.

[0136] The manifold alignment unit 3313 is used to define the distance between different feature spaces by the Riemannian metric, and to realize the alignment of the two manifolds. The Riemannian metric defines the distance between points on the manifold, so that similar structures in different feature spaces can be kept similar in the unified feature space. In this embodiment, a variant of the Iterative Closest Point (ICP) algorithm is used to realize the manifold alignment, and the number of iterations is set to 50 and the convergence threshold is 0.001. These parameters are determined by experiments and can achieve good alignment effect under reasonable computational complexity.

[0137] The unified feature representation unit 3314 is used to map the aligned features to a unified feature representation space. The unified feature representation is , where is the unified feature matrix; is the unified feature dimension (set to 256 in this embodiment); is the number of feature samples. The unified feature representation is calculated by the following formula:

[0138] ,

[0139] where, is the mapping result of the point cloud feature; is the mapping result of the acoustic feature; [;] represents the feature splicing operation, that is, splicing in the feature dimension. In order to ensure the unified representation of the features, the spliced features are normalized:

[0140] ,

[0141] where, is the normalized unified feature; is the spliced feature; is the mean vector of the feature; is the standard deviation vector of the feature; represents the Z-score standardization of the feature, so that the mean of the feature is 0 and the standard deviation is 1. Normalization helps to eliminate the influence of different feature scales and improve the fusion effect.

[0142] Through the above processing steps, the topology-aware multi-dimensional feature space mapping structure 331 can map features of different modalities to a unified feature space, maintain the topology structure of the original features, and lay a foundation for subsequent feature fusion.

[0143] With reference to Figure 6 , the self-attention dynamic weight mechanism 332 on the matrix manifold includes a feature matrix construction unit 3321, a correlation matrix calculation unit 3322, a multi-head self-attention unit 3323, and a matrix manifold optimization unit 3324.

[0144] The feature matrix construction unit 3321 is configured to organize the unified feature representation into a feature matrix. The feature matrix is represented as wherein is the feature matrix; is the feature dimension (set to 256 in this embodiment); is the time step, which is dynamically adjusted according to the input data. To ensure that the condition number of the matrix is within a reasonable range, the feature matrix is initialized as follows:

[0145] ,

[0146] wherein is the initialized feature matrix; is the original feature matrix; is the feature dimension; represents dividing each element of the matrix by , which is a commonly used scaling factor in the Transformer model and helps stabilize the training process.

[0147] The correlation matrix calculation unit 3322 is configured to calculate the inner product of the feature matrix to obtain a correlation matrix. The correlation matrix is represented as wherein is the correlation matrix; is the transpose matrix of ; and represents matrix multiplication; is the time step. The correlation matrix reflects the correlation between features at different time steps. To avoid matrix ill-conditioning, the correlation matrix is regularized as follows:

[0148] ,

[0149] wherein is the regularized correlation matrix; is the original correlation matrix; is a regularization parameter (set to 0.01 in this embodiment); is an identity matrix with a dimension of The regularization parameter 0.01 is determined by experiment and can effectively prevent matrix ill-conditioning and improve computational stability.

[0150] The multi-head self-attention unit 3323 is used to decompose the correlation matrix into multiple sub-matrices and calculate the attention weight respectively. In this embodiment, 8 attention heads are used, and the attention weight is calculated independently for each attention head. The calculation formula of multi-head self-attention is as follows:

[0151] ,

[0152] ,

[0153] ,

[0154] where MultiHead is the output of multi-head attention; 、 、 are the query, key and value matrices, respectively, which are usually set to the same input matrix; Concat is a concatenation operation that concatenates the outputs of multiple heads in the feature dimension; head is the output of the th attention head; is the number of attention heads (set to 8 in this embodiment); 、 、 are the learnable weight matrices of the th head, which are used to transform the query, key and value, respectively; is the output transformation matrix; Attention is an attention calculation function; softmax is a softmax activation function that converts the input into a probability distribution; is the dimension of the key, which is usually equal to , where is the feature dimension. The advantage of using the multi-head attention mechanism is that it allows the model to learn information from different representation subspaces, enhancing the model's expressive power.

[0155] To enhance the relevance of different functional areas of the larynx, a laryngeal functional relevance attention mechanism is introduced:

[0156] ,

[0157] where, is a functional relevance attention function; is a functional relevance mask matrix, which is used to enhance the attention weight between

[0158] the features of the functional areas.

[0159] Dynamically adjust attention weights for different rehabilitation training methods:

[0160] ,

[0161] in, The training method's adaptive weight matrix; and These are weight matrices for breathing training, relaxation training, resonance voice training, and airflow vocalization training, respectively. , , and These are the corresponding weight coefficients, which are dynamically adjusted based on the current training type being identified, satisfying...

[0162] The matrix manifold optimization unit 3324 is used to define gradients on the matrix manifold to guide the optimization of the weight matrix. Optimization on the matrix manifold takes into account the geometric properties of the matrix, enabling more accurate optimization of the weight matrix. In this embodiment, the Riemann optimization algorithm is used, and the iterative optimization process is expressed as follows:

[0163] ,

[0164] in, For the first The weight matrix for the next iteration; For the first The weight matrix for the next iteration; for The contraction mapping at the point maps the directions in the tangent space back to the manifold; The learning rate is set to 0.001 in this embodiment. For the objective function exist The Riemann gradient at point 0.001. A learning rate of 0.001 is a commonly used initial learning rate in deep learning, achieving a good balance between convergence speed and stability.

[0165] Through the above processing steps, the self-attention dynamic weighting mechanism 332 on the matrix manifold can accurately capture the complex relationships between features, realize dynamic weight allocation, and improve the quality of feature fusion.

[0166] Reference Figure 7 The probability field-based adaptive feature reconstruction system 333 includes a probability field modeling unit 3331, an information entropy maximization unit 3332, a conditional random field construction unit 3333, and an adaptive feature reconstruction unit 3334.

[0167] The probability field modeling unit 3331 is used to treat the feature distribution as a probability field and establish the probability density function of the features. The feature distribution is represented as a probability density function. where is the probability density of the feature ; represents the feature vector. In this embodiment, the feature distribution is modeled using an exponential family distribution, denoted as

[0168] ,

[0169] where is the probability density of the feature ; is the natural exponential function; is the natural parameter, representing the parameter vector of the distribution; is the sufficient statistic, a function of the feature ; is the log-partition function, used to normalize the probability density, ensuring the probability integral is 1; represents the inner product of the parameter vector and the statistic. The prior distribution of the feature is determined based on domain knowledge and data analysis, and the energy function is used to describe the organization and structure of the feature.

[0170] The information entropy maximization unit 3332 is used to calculate the information entropy of the feature distribution, maximizing the information entropy under the constraint of maintaining useful information. The information entropy is denoted as

[0171] ,

[0172] where is the information entropy of the distribution , representing the uncertainty of the distribution; is the probability density of the feature ; is the natural logarithm; represents the integral; represents the integral of . Maximizing the information entropy under certain constraints can obtain the optimal feature distribution. The optimization objective is denoted as

[0173] ,

[0174] ,

[0175] where represents the maximization objective function; is the information entropy of the distribution ; represents the constraint condition; is the sufficient statistic of the feature; usually a certain function of the feature; is the corresponding expected value, usually estimated from the training data; The number of constraints. The theoretical basis of information entropy maximization is the maximum entropy principle, that is, choose the distribution with the maximum entropy as the optimal distribution under the condition of meeting the known constraints, so as to avoid introducing additional assumptions.

[0176] The conditional random field construction unit 3333 is used to establish the conditional dependence graph between features and describe the relationship between features. The conditional random field model is represented as:

[0177] ,

[0178] Wherein, is the conditional probability of the output label y given the input feature x; y is the output label, indicating the category of the laryngeal rehabilitation action; x is the input feature, indicating the fused feature vector; Z(x) is the normalization factor, so that the sum of the conditional probabilities is 1; exp is the natural exponential function; and are the weight parameters of the feature function and the edge feature function respectively; is the feature function, which is usually a function of the input feature and the output label; is the edge feature function, which describes the dependence between labels; and represent the sum of all feature functions and edge feature functions respectively. In this embodiment, the conditional random field adopts a full connection structure, and the number of nodes is equal to the feature dimension. The advantage of the conditional random field is that it can consider the dependence between output labels and improve the classification performance.

[0179] For different training targets, construct a training target conditional random field:

[0180] ,

[0181] Wherein, is the conditional probability of the output label given the input feature and the training method type ; is the corresponding normalization factor; and are the feature function and the edge feature function related to the training method respectively, which can adjust the behavior of the model according to the characteristics of different training methods.

[0182] The adaptive feature reconstruction unit 3334 is used to dynamically adjust the distribution and representation of features according to the characteristics of the probability field. The adaptive gating mechanism controls the contribution degree of different features, which is represented as:

[0183] ,

[0184] ,

[0185] wherein, is a gating value, ranging from [0, 1], representing the retention degree of the original feature; is a sigmoid activation function, mapping the input to the interval [0, 1]; and are learnable weight matrix and bias vector, respectively; represents matrix multiplication; is the reconstructed output feature; is the original feature; is the candidate feature, usually obtained through some transformation; represents element-wise product (Hadamard product), i.e., multiplying corresponding elements. In this embodiment, the initial value of the adaptive gating parameter is set to 0.5, which is dynamically adjusted through backpropagation. During the iterative optimization process, the number of iterations is set to 30, and the convergence threshold is 0.005. The advantage of the adaptive gating mechanism lies in its ability to dynamically adjust the importance of features according to different input features, thereby improving the adaptability of the model.

[0186] Add a patient individuality adaptation module:

[0187] ,

[0188] wherein, is an adaptive gating function considering patient individual differences; is the patient individuality feature vector, containing information such as physiological characteristics, medical history, and rehabilitation stage of the patient. This improvement enables the system to dynamically adjust the feature representation according to the individual differences of the patient, providing more personalized evaluation results.

[0189] Through the above processing steps, the adaptive feature reconstruction system 333 based on probability field can dynamically optimize the feature representation, improve the robustness and adaptability of the system, and adapt to the needs of different patients and different rehabilitation stages.

[0190] Through the above processing steps, the adaptive feature reconstruction system 333 based on probability field can dynamically optimize the feature representation, improve the robustness and adaptability of the system, and adapt to the needs of different patients and different rehabilitation stages.

[0191] The final output of the multi-modal action feature fusion module 33 includes: action correctness discrimination results, action start time, action execution time, and evaluation analysis results of action correct execution time proportion.

[0192] The action correctness discrimination result is represented as a binary result (correct / incorrect), which is obtained by comparing the similarity of the patient's executed action and the standard action. The similarity calculation formula is:

[0193] ,

[0194] wherein, is the similarity, which ranges from -1 to 1, and the greater the value, the more similar the two vectors are; represents the cosine similarity function; is the feature vector of the patient's performance of the action; is the feature vector of the standard action; represents the inner product of two vectors; represents the Euclidean norm of the vector , i.e. , wherein is each element of the vector . When the similarity is greater than a threshold value (set to 0.85 in this embodiment), the action is determined to be correct, otherwise it is determined to be incorrect. The threshold value of 0.85 is determined through experiments and can effectively distinguish between correct and incorrect actions.

[0195] The action start time is represented as the time interval from the issuance of the instruction to the patient starting to perform the action, in milliseconds. The action execution time is represented as the time taken by the patient to complete the entire action, in seconds. The action correct execution time ratio is represented as the percentage of the time taken by the patient to correctly perform the action out of the total execution time, and the calculation formula is:

[0196] ,

[0197] wherein, is the action correct execution time ratio, in percentage; is the time taken to correctly perform the action, in seconds; is the total execution time, in seconds; represents the conversion of the ratio to a percentage. This index can reflect the correctness of the patient during the entire action process, and is of great significance for evaluating the rehabilitation effect.

[0198] For the four basic rehabilitation methods, the system provides the following specialized evaluation indexes:

[0199] Respiratory training evaluation indexes:

[0200] Respiratory flow stability index, indicating the stability of respiratory airflow;

[0201] Chest and abdominal coordinated breathing score, evaluating the coordination of chest and abdominal breathing;

[0202] Cricoid cartilage stability quantitative value, reflecting the stability of the cricoid cartilage during breathing;

[0203] Relaxation training evaluation indexes:

[0204] thyroid tension index, a tension indicator calculated based on changes in thyroid cartilage position and acoustic characteristics;

[0205] larynx position stability score, assessing the stability of the overall larynx position based on changes in hyoid bone position;

[0206] vocal cord relaxation degree, assessing the relaxation degree of vocal cords based on cricothyroid cartilage abduction state and acoustic indicators;

[0207] resonance voice training evaluation indicators:

[0208] epiglottis lifting degree, quantifying the change in the lifting angle of the epiglottis cartilage;

[0209] laryngopharyngeal cavity resonance space index, evaluating the size and shape of the resonance space;

[0210] resonance peak energy distribution score, analyzing the energy distribution of resonance peaks in the sound spectrum;

[0211] airflow sound production training evaluation indicators:

[0212] glottal closure efficiency, assessing the closure effect of the glottis based on the adduction state of the cricothyroid cartilage;

[0213] breathy voice-vocal cord sound ratio, evaluating the ratio between breathy voice and sound produced by vocal cord vibration;

[0214] vocal cord vibration stability index, evaluating the stability and regularity of vocal cord vibration;

[0215] comprehensive training effect evaluation:

[0216] training method synergy index, evaluating the synergy effect between multiple training methods;

[0217] rehabilitation progress quantitative score, based on the longitudinal comparison results of multiple training data;

[0218] individualized improvement index, relative progress compared with baseline data of the patient;

[0219] In addition, the system also provides analysis of the synergistic relationship between the movement of various laryngeal cartilage structures and rehabilitation training methods, including:

[0220] thyroid cartilage and relaxation training synergy, analyzing the relationship between thyroid cartilage inclination changes and vocal cord tension

[0221] cricoid cartilage and breathing training synergy, evaluating the correlation between cricoid cartilage stability and breathing support effect;

[0222] cricothyroid cartilage and airflow sound production training synergy, quantifying the relationship between cricothyroid cartilage adduction / abduction and glottal closure efficiency;

[0223] The synergistic relationship between the epiglottis cartilage and resonance training, analyzing the relationship between the epiglottis lifting angle and the expansion of the resonance space

[0224] The synergistic relationship between the hyoid bone and multi-method training, evaluating the relationship between the hyoid bone position and the laryngeal position adjustment

[0225] These quantitative evaluation indicators and synergistic relationship analyses provide objective references for doctors, facilitating the tracking of patient rehabilitation progress, the adjustment of rehabilitation plans, and the realization of personalized treatment.

[0226] Reference Figure 8 The present application also provides a multi-modal learning laryngeal rehabilitation motion quantitative evaluation method, comprising the following steps:

[0227] Collecting multiple three-dimensional point cloud data streams from different angles collected by the millimeter wave radar module and audio signals collected by the standard microphone;

[0228] Inputting the collected data into the trained deep learning neural network model, and obtaining the laryngeal rehabilitation motion quantitative evaluation result through the trained deep learning neural network model;

[0229] The deep learning neural network model is trained through a multi-modal data set, which includes a three-dimensional point cloud data set and an acoustic signal set. In a preferred embodiment of the present application, the three-dimensional point cloud data set contains depth camera data, optical motion capture data, and expert score data based on multi-modal data, covering the motion data of key structures such as the thyroid cartilage, the ring-shaped cartilage, the arytenoid cartilage, the epiglottis cartilage, and the hyoid bone. The acoustic signal set contains standard audio samples and patient training audio samples of four basic rehabilitation methods: breathing training, relaxation training, resonance voice training, and airflow sound training.

[0230] The deep learning neural network model contains three sub-networks: a point cloud processing module, an acoustic signal processing module, and a multi-modal action feature fusion module. The point cloud processing module is used to obtain three-dimensional motion parameters of key parts such as the thyroid cartilage, the ring-shaped cartilage, the arytenoid cartilage, the epiglottis cartilage, and the hyoid bone; the acoustic signal processing module extracts features from the audio signal, including extracting acoustic features related to the four basic rehabilitation methods: breathing training, relaxation training, resonance voice training, and airflow sound training; the multi-modal action feature fusion module fuses the three-dimensional motion parameters obtained by the point cloud processing module and the results of the acoustic signal processing module at the feature level, thereby outputting evaluation analysis results including action correctness discrimination results, action start time, action execution time, and action correct execution time proportion, as well as specialized evaluation indicators for the four basic rehabilitation methods and synergistic relationship analysis results between the motion of each laryngeal cartilage structure and the rehabilitation training method.

[0231] Like the system embodiment, the multi-modal action feature fusion module includes a multi-dimensional feature space mapping structure through topology perception, a self-attention dynamic weight mechanism on a matrix manifold, and a self-adaptive feature reconstruction system based on a probability field, realizing deep fusion of point cloud features and acoustic features.

[0232] The multi-modal learning laryngeal rehabilitation action quantitative evaluation method of the application comprehensively captures the characteristics of laryngeal rehabilitation actions by fusing data of different modes, provides professional quantitative evaluation indexes, and provides strong support for laryngeal rehabilitation treatment.

[0233] The above is only a preferred embodiment of the application, and does not limit the scope of the application. Any equivalent structural transformation using the content of the application specification and drawings, or direct or indirect application in other related technical fields, is also included in the protection scope of the application.

Claims

1. A multi-modal learning laryngeal rehabilitation motion quantification evaluation system, characterized in that, The application relates to a laryngeal rehabilitation system based on multi-modal fusion of three-dimensional point cloud data and audio signals. The application comprises: a millimeter wave radar module for capturing a three-dimensional point cloud data stream of a human laryngeal region, wherein the laryngeal region comprises thyroid cartilage, cricoid cartilage, arytenoid cartilage, epiglottic cartilage and hyoid key structures; a standard microphone for collecting laryngeal rehabilitation instructions and vocalization training audio signals; a deep learning neural network model for receiving the three-dimensional point cloud data stream and the audio signals and obtaining a quantitative evaluation result of laryngeal rehabilitation actions; wherein the deep learning neural network model comprises three sub-networks: a point cloud processing module, an acoustic signal processing module and a multi-modal action feature fusion module; the point cloud processing module is used for processing the three-dimensional point cloud data stream to obtain three-dimensional motion parameters of each key cartilage structure of the larynx; the acoustic signal processing module is used for extracting features from the audio signals; 2. The multimodal learning-based laryngeal rehabilitation motion quantification evaluation system according to claim 1, wherein, the multi-modal action feature fusion module is used for receiving the three-dimensional motion parameters output by the point cloud processing module and the features output by the acoustic signal processing module, and fusing the features at a feature level by using a multi-modal action feature fusion algorithm to obtain the quantitative evaluation result of the laryngeal rehabilitation actions; wherein the multi-modal action feature fusion module comprises a topologically-aware multi-dimensional feature space mapping structure, a self-attention dynamic weight mechanism on a matrix manifold and an adaptive feature reconstruction system based on a probability field. 3.The laryngeal rehabilitation motion quantification evaluation system of multimodal learning according to claim 1, wherein, The millimeter wave radar module comprises a plurality of millimeter wave radars for obtaining three-dimensional point cloud data streams of multiple perspectives covering each key cartilage structure of the larynx.

4. The multimodal learning-based laryngeal rehabilitation motion quantification evaluation system according to claim 3, wherein, The point cloud processing module takes three-dimensional point cloud data as input, first removes external background noise and samples to obtain corresponding laryngeal region point cloud data, then extracts point clouds of the thyroid cartilage, cricoid cartilage, arytenoid cartilage, epiglottic cartilage and hyoid from the region point cloud data, and finally separates these parts. 5.The laryngeal rehabilitation motion quantification evaluation system of multimodal learning according to claim 1, wherein, The point cloud processing module extracts three-dimensional motion parameter features by using a convolution block, a pooling block, an up-sampling block and a multi-layer perception block, including: background noise removal and sampling by using a K-neighborhood screening algorithm; extracting point clouds of each key cartilage structure of the larynx from the region point cloud data; automatically calibrating point marks of each key part of the cartilage by using a point cloud key point detection algorithm; obtaining coordinate data of each key part of the cartilage; slicing the point cloud of the laryngeal region to obtain each frame of point cloud sequence, calibrating key points of each cartilage, and obtaining positions and motion trajectories of each cartilage structure by using a key point displacement interpolation algorithm and a key point coordinate conversion algorithm.

6. The multimodal learning-based laryngeal rehabilitation motion quantification evaluation system according to claim 1, wherein, The acoustic signal processing module extracts motion acoustic features by using a convolution block, a pooling block, an up-sampling block and a multi-layer perception block, including: extracting high-frequency spectrum features and correlation features of the original audio signals; extracting breathing training-related acoustic features, relaxation training-related acoustic features, resonance voice training-related acoustic features and airflow vocalization training-related acoustic features; and deeply fusing the extracted features by using an LSTM block and a time sequence strategy. The topologically-aware multi-dimensional feature space mapping structure comprises: a multi-dimensional feature space construction unit for constructing a point cloud feature space and an acoustic feature space; a local structure preserving mapping unit for designing a feature mapping function to preserve the topological structure of the original feature space; The manifold alignment unit is configured to align two manifolds by defining distances of different feature spaces through Riemannian metrics. The unified feature representation unit is configured to map the aligned features to a unified feature representation space.

7. The multimodal learning-based laryngeal rehabilitation motion quantification evaluation system according to claim 1, wherein, The self-attention dynamic weight mechanism on the matrix manifold comprises: The feature matrix construction unit is configured to organize the unified feature representation into a feature matrix. The correlation matrix calculation unit is configured to calculate the inner product of the feature matrix to obtain a correlation matrix. The multi-head self-attention unit is configured to decompose the correlation matrix into multiple sub-matrices to calculate attention weights. The matrix manifold optimization unit is configured to define a gradient on the matrix manifold to guide the optimization of the weight matrix.

8. The multimodal learning-based laryngeal rehabilitation motion quantification evaluation system according to claim 1, wherein, The adaptive feature reconstruction system based on the probability field comprises: The probability field modeling unit is configured to regard the feature distribution as a probability field and establish a probability density function of the feature. The information entropy maximization unit is configured to calculate the information entropy of the feature distribution and maximize the information entropy under the constraint of maintaining useful information. The conditional random field construction unit is configured to establish a conditional dependency graph between features to describe the relationship between features. The adaptive feature reconstruction unit is configured to dynamically adjust the distribution and representation of the feature according to the characteristics of the probability field. 9.The laryngeal rehabilitation motion quantification evaluation system of multimodal learning according to claim 1, wherein, The quantitative evaluation result of the laryngeal rehabilitation action comprises: an action correctness discrimination result, an action start time, an action execution time, and an evaluation analysis result of an action correct execution time proportion; specialized evaluation indexes for four basic rehabilitation methods of breathing training, relaxation training, resonant voice training, and airflow sound production training; and a synergistic relationship analysis result between each laryngeal cartilage structure movement and the rehabilitation training method.

10. A method of laryngeal rehabilitation movement quantitative evaluation of multimodal learning, characterized in that, The method comprises the following steps: Collecting three-dimensional point cloud data streams of multiple viewing angles collected by a millimeter wave radar module and audio signals collected by a standard microphone, and then inputting the collected data into a trained deep learning neural network model to obtain a laryngeal rehabilitation action quantitative evaluation result through the trained deep learning neural network model; The deep learning neural network model is trained through a multi-modal data set, and the multi-modal data set comprises a three-dimensional point cloud data set and an acoustic signal set; The deep learning neural network model comprises three sub-networks: a point cloud processing module, an acoustic signal processing module, and a multi-modal action feature fusion module; The point cloud processing module is configured to obtain three-dimensional motion parameters of key parts of the thyroid cartilage, the cricoid cartilage, the arytenoid cartilage, the epiglottis cartilage, and the hyoid bone; The acoustic signal processing module is configured to extract features from the audio signals, including acoustic features related to the four basic rehabilitation methods of breathing training, relaxation training, resonant voice training, and airflow sound production training; The multi-modal action feature fusion module is configured to fuse the three-dimensional motion parameters obtained by the point cloud processing module and the results of the acoustic signal processing module at the feature level, thereby outputting an evaluation analysis result comprising an action correctness discrimination result, an action start time, an action execution time, and an action correct execution time proportion, and specialized evaluation indexes for the four basic rehabilitation methods. The multi-modal action feature fusion module comprises a multi-dimensional feature space mapping structure through topology perception, a self-attention dynamic weight mechanism on a matrix manifold and a self-adaptive feature reconstruction system based on a probability field, and realizes deep fusion of point cloud features and acoustic features.

Citation Information

Patent Citations

  • Swallowing detection method and device based on binocular vision, equipment and storage medium

    CN114515395A

  • Auxiliary scheme generation method and system for dysphagia rehabilitation

    CN120600228A