Parkinson's disease assessment method and system based on face video multi-modal physiological feature fusion
By using a multimodal physiological feature fusion method of facial videos, combined with rPPG signals and eye movement behavior parameters, the problem of insufficient sensitivity and specificity in Parkinson's disease diagnosis in existing technologies is solved, and accurate quantification of multi-system pathological characteristics of Parkinson's disease and full-cycle diagnosis and treatment support are achieved.
Patent Information
- Application Number
- CN202511278861.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies for diagnosing Parkinson's disease have insufficient sensitivity and specificity in single-modality detection methods, making it difficult to comprehensively capture multi-system pathological characteristics. In addition, there are challenges in integrating multimodal fusion technology and optimizing algorithms, which limits the clinical application of non-invasive detection.
Through a multimodal physiological feature fusion method based on facial video, combined with optical flow tracking and IPAST analysis, rPPG signals and eye movement behavior parameters are extracted, and feature fusion is performed using a deep learning network to output a continuous risk score for Parkinson's disease.
It has achieved comprehensive capture and accurate quantification of the multi-system pathological characteristics of Parkinson's disease, improved the sensitivity and specificity of diagnosis, and adapted to the full-cycle diagnosis and treatment needs of early screening and disease course monitoring.
Smart Images

Figure CN120809239A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to multiple technical fields such as computer vision and physiological signal processing, and in particular to a Parkinson's disease evaluation method and system based on multi-modal physiological feature fusion of face video. BACKGROUND
[0002] In recent years, early diagnosis and disease monitoring of neurodegenerative diseases (such as Parkinson's disease) have attracted much attention in the fields of geriatric medicine and neuroscience. Currently, the diagnosis of Parkinson's disease mainly relies on motor symptom scale evaluation (such as UPDRS-III) and molecular imaging examination, but the former is easily affected by subjective factors, and the latter is difficult to be widely applied in population screening and long-term monitoring due to high cost, complex operation and invasiveness. In recent years, non-invasive detection technology based on physiological signals (such as eye tracking, heart rate variability analysis) has provided a new idea for the auxiliary diagnosis of Parkinson's disease. However, the existing single modal detection method (such as only analyzing eye movement or only measuring heart rate) often fails to fully capture the multi-system pathological characteristics of Parkinson's disease, resulting in insufficient sensitivity and specificity of diagnosis. In addition, traditional methods usually rely on pre-set thresholds or linear models, which cannot effectively handle individual differences and disease heterogeneity, especially in the early stages or atypical cases. On the other hand, although multi-modal fusion (such as combining eye movement with autonomic nervous indicators) has more potential in theory, how to realize efficient data synchronization acquisition, dynamic feature extraction and cross-modal correlation analysis still faces major challenges in technical integration and algorithm optimization. These limitations seriously restrict the promotion and application efficiency of non-invasive technology in the clinical practice of Parkinson's disease.
[0003] In recent years, certain progress has been made in the task prediction method based on the fusion of multi-modal physiological features of facial video at home and abroad, but there are still deficiencies, for example: the patent with publication number CN119214619A proposes a heartbeat estimation method based on facial video RPPG signal, but the method is not accurate in heart rate variability (HRV) analysis, focuses on average heart rate, and has weak ability to capture high-frequency physiological fluctuations (such as respiratory sinus arrhythmia), which limits its application in autonomic nervous function assessment (such as Parkinson's disease). The patent with publication number CN117678998A proposes an rPPG algorithm based on adaptive projection plane and feature screening, which converts three-dimensional RGB signals into one-dimensional BVP signals containing pulse information. Although the band-pass filtering of the obtained BVP signal removes noise outside the normal heart rate range, it eliminates the interference caused by factors such as light changes and facial movements during the detection of heart rate based on facial video, but lacks multi-modal collaborative optimization and does not fuse with other physiological signals (such as eye movement, EEG), which cannot utilize cross-modal correlation to improve diagnostic specificity (such as the combination of scanning abnormalities and HRV in Parkinson's disease); the patent with publication number CN116343280A proposes a non-contact coronary heart disease evaluation device based on facial video, which comprehensively considers pulse features and facial features, respectively evaluates the uncertainty of the two model predictions, and selects the prediction result with smaller uncertainty as the final coronary heart disease detection result, which relatively more comprehensively evaluates the coronary heart disease related features, but there is a limitation that the correlation between facial features and coronary heart disease is insufficient, and the facial microvascular changes (such as pallor and cyanosis) of coronary heart disease patients may be confused by other factors (such as anemia and hypothermia), leading to false positives. SUMMARY
[0004] The purpose of the present application is to overcome the deficiencies of the prior art and provide a Parkinson's disease evaluation method and system based on the fusion of multi-modal physiological features of facial video.
[0005] The purpose of the present application is achieved by the following technical solutions: In a first aspect, the present application discloses a Parkinson's disease evaluation method based on the fusion of multi-modal physiological features of facial video, comprising the following steps: S1, based on a continuous facial video stream, rPPG signals are extracted through optical flow tracking and facial ROI region modeling technology; then the heart rate variability key parameters are calculated to obtain the heart rate variability features; S2, based on the IPAST analysis of the dynamic changes of the eye fixation points in the face video, key eye movement behavior parameters are extracted, and the eye movement behavior parameters are modeled in space and time by a convolution gating recurrent unit to obtain eye movement behavior features; the eye movement behavior parameters include a percentage of off-center fixation points in a forward saccade task, a percentage of off-center fixation points in a reverse saccade task, a number of forward saccades occurring too early before the appearance of a stimulus, a number of reverse saccades occurring too early before the appearance of a stimulus, a reaction time of a forward saccade, a reaction time of a reverse saccade, a fast and correct forward saccade generated between 90-130 milliseconds, a fast and incorrect reverse saccade generated between 90-130 milliseconds, a reverse saccade error within a normal delay range, a time point at which an autonomous reaction overpowers an automatic reaction, a speed of a correct forward saccade, and a speed of a correct reverse saccade; S3, inputting the heart rate variability features and the eye movement behavior features into a multi-modal fusion network structure, fusing the spliced vectors of the two by a transformer, and outputting a continuous risk score or a disease state label for assisting doctors in early Parkinson's risk assessment.
[0006] Based on the first aspect, step S1 specifically includes the following steps: S11, using a Dlib detector to detect 81 feature points of a face, then selecting three regions of a forehead and both cheeks to obtain an ROI region, and using a sparse optical flow tracking algorithm to stabilize the ROI region, for the first t frame image , calculating the pixel average value in the ROI region of the RGB channel by the formula , wherein , M represents the total number of frames, represents space, H represents height, W represents width, and C represents the number of channels, represents the pixel coordinates in the ROI region , and represents the number of pixels in the ROI region , , , , and constitute a three-dimensional color time sequence, and an rPPG signal is extracted; S12, using a skin orthogonal plane method POS to model the rPPG signal, projecting the skin reflection changes in the RGB space to a direction with a high heart rate signal-to-noise ratio to obtain an original rPPG signal sequence ; S13, performing wavelet denoising and band-pass filtering optimization processing on the original rPPG signal to eliminate the noise of the original rPPG signal, and obtaining an rPPG signal sequence ; S14, performing peak search on the rPPG signal sequence performing peak search on the rPPG signal sequence wherein N represents the total length of the RR interval sequence; S15, calculating time domain features according to the RR interval sequence, including standard deviation normal to normal interval SDNN and root mean square difference of adjacent heartbeat intervals RMSSD; the calculation formula is as follows: ; ; ; wherein represents the mean of the RR interval sequence, i is the index symbol in the RR interval sequence; S16, calculating frequency domain features according to the RR interval sequence, including low frequency feature LF and high frequency feature HF, and calculating the ratio LF / HF of the two; finally, the time domain features and the frequency domain features are spliced, and the heart rate variability feature vector is obtained through the long short-term memory network model LSTM , .
[0007] Based on the first aspect, step S2 specifically comprises the following steps: S21, analyzing the face video through IPAST, extracting 12 eye movement behavior parameters and forming eye movement parameter : , extracting 4 common factors with neurophysiological significance to form a common factor sequence : wherein represents the task disengagement ability, represents the impulse inhibition, represents the voluntary saccade generation, represents the saccade dynamics; S22, performing principal axis factor extraction and oblique rotation on the eye movement parameter , converting the 12 eye movement behavior parameters into 4 common factors, and constructing a common factor matrix for the 12 eye movement behavior parameters to satisfy wherein represents the loading matrix, represents the error term; S23, inputting the common factor sequence into a gated recurrent unit for spatiotemporal modeling, which is used to capture the dynamic evolution characteristics of the common factor sequence across time periods; first, the common factor sequence passes through three convolutional blocks, each of which contains a convolutional layer and a max pooling layer, and then passes through a gated recurrent unit GRU and a fully connected layer FC, and finally obtains an eye movement behavior feature vector : .
[0008] Based on the first aspect, step S3 specifically comprises the following steps: S31, inputting the heart rate variability feature vector and the eye movement behavior feature vector into a multi-modal physiological feature fusion network, the multi-modal physiological feature fusion network comprising an attention fusion layer Atten_Layer, a residual block Res_Block, an adaptive feature pooling layer AdaPool_Layer, and a regression prediction layer RL; S32, performing vector splicing on the heart rate variability feature vector and the eye movement behavior feature vector to obtain a fusion vector , and then inputting the fusion vector into the attention fusion layer Atten_Layer to adaptively learn the importance weight of the two feature vectors, dynamically fuse the heart rate variability feature and the eye movement behavior feature, and obtain an attention fusion vector : ; S33, processing the attention fusion vector through the residual block Res_Block to obtain residual features : , enhancing the expression ability of the model to complex features, and relieving the gradient vanishing problem; S34, automatically extracting key features : from the residual features through the adaptive feature pooling layer AdaPool_Layer; S35, finally processing the key features through the regression prediction layer RL to output the physiological state change and the activity prediction value of the Parkinson's disease patient : .
[0009] In a second aspect, the present application discloses a Parkinson's disease evaluation system based on multi-modal physiological feature fusion of face video, which is used for the Parkinson's disease evaluation method based on multi-modal physiological feature fusion of face video in any one of the above aspects, and comprises: a heart rate variability feature extraction module, configured to extract an rPPG signal according to a continuous face video stream, and further extract a heart rate variability feature according to the rPPG signal; an eye movement behavior feature extraction module, configured to analyze the dynamic change of eye fixation points in the face video based on IPAST, extract key eye movement behavior parameters, and perform spatio-temporal modeling on the eye movement parameters through a convolutional gated recurrent unit to obtain eye movement behavior features. The fusion analysis and auxiliary module inputs the extracted heart rate variability features and eye movement behavior features into a multi-modal fusion network structure, and outputs a continuity risk score or a disease state label, which is used for assisting doctors in early Parkinson risk assessment.
[0010] The present application has the following advantages: 1) By deeply integrating various physiological signals collected by video, a multi-modal feature collaborative analysis framework is constructed, which not only breaks through the subjective limitations of traditional scale assessment, but also overcomes the one-sided defects of single biomarker detection, realizes the comprehensive capture and accurate quantification of the multi-system pathological characteristics of Parkinson's disease, and adapts to the whole cycle of diagnosis and treatment needs from early screening to disease monitoring. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 Fig. 1 is a Parkinson's assessment network architecture based on HRV signal and eye movement signal according to an embodiment of the present application; Figure 2 Fig. 2 is a structure diagram of an HRV feature extraction module according to an embodiment of the present application; Figure 3 Fig. 3 is a structure diagram of an eye movement behavior feature extraction module according to an embodiment of the present application. DETAILED DESCRIPTION
[0012] The technical solutions of the present application will be described in detail below with reference to the embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0013] The present application discloses a Parkinson's disease assessment method and system based on multi-modal physiological features of face video. By deeply integrating various physiological signals collected by video, a multi-modal feature collaborative analysis framework is constructed, which not only breaks through the subjective limitations of traditional scale assessment, but also overcomes the one-sided defects of single biomarker detection, realizes the comprehensive capture and accurate quantification of the multi-system pathological characteristics of Parkinson's disease. This new method will significantly improve the sensitivity and specificity of the system, adapt to the whole cycle of diagnosis and treatment needs from early screening to disease monitoring, and the Parkinson's assessment network architecture based on HRV signal and eye movement signal is as shown in Fig. 1. Figure 1As shown, the method combines heart rate variability (HRV) features converted from rPPG (remote photoplethysmography) signals with eye movement behavior parameters for multi-system assessment of Parkinson's disease. By analyzing facial videos, it simultaneously extracts core features of autonomic nervous system function and eye movement control systems, thereby achieving a comprehensive perception of Parkinson's disease's motor and non-motor symptoms. This method provides an efficient and objective means for disease assessment and personalized intervention. The method specifically includes the following steps: S1. Based on a continuous face video stream, rPPG signals are extracted through optical flow tracking and facial ROI modeling technology. Then, key parameters of heart rate variability are calculated to obtain heart rate variability features (HRV features). S2. Analyze the dynamic changes of eye gaze points in face videos based on IPAST, extract key eye movement behavior parameters, and perform spatiotemporal modeling of the eye movement behavior parameters through convolutional gated recurrent units to obtain eye movement behavior characteristics; the eye movement behavior parameters include the percentage of off-center gaze points in the forward saccade task, the percentage of off-center gaze points in the reverse saccade task, the number of forward saccades that occur prematurely before the stimulus appears, the number of reverse saccades that occur prematurely before the stimulus appears, the reaction time of forward saccades, the reaction time of reverse saccades, rapid correct forward saccades generated between 90-130 milliseconds, rapid incorrect reverse saccades generated between 90-130 milliseconds, reverse saccade errors within a normal latency range, the time point when the autonomous response begins to overwhelm the automatic response, the speed of correct forward saccades, and the speed of correct reverse saccades; S3. Input the heart rate variability characteristics and eye movement behavior characteristics into the multimodal fusion network structure, fuse the concatenated vectors of the two through the transformer, and output a continuous risk score or disease status label to assist doctors in early Parkinson's risk assessment.
[0014] Specifically, step S1 includes the following steps: S11, use Dlib detector to detect 81 feature points of the face, then select the forehead and cheeks to get the ROI area, and use sparse optical flow tracking algorithm to stabilize the ROI area. t Frame Image , through the formula Calculate its ROI area in RGB channel The average value of pixels in ,in , M represents the total number of frames, Indicates space, H indicates height, W indicates width, and C indicates the number of channels. Indicates ROI area The pixel coordinates in Indicates ROI area the number of pixels in the image, , 、 and constitute a three-dimensional color time sequence, and the rPPG signal is extracted; S12, model the rPPG signal using a Plane Orthogonal to Skin (POS) method, project the skin reflection changes in the RGB space onto a direction with a high heart rate signal-to-noise ratio, and obtain an original rPPG signal sequence ; S13, perform wavelet denoising and band-pass filtering optimization processing on the original rPPG signal to eliminate the noise of the original rPPG signal, and obtain an rPPG signal sequence ; S14, perform peak value search on the rPPG signal sequence , and calculate an RR interval sequence , where N represents the total length of the RR interval sequence; S15, calculate time domain features according to the RR interval sequence, including a standard deviation normal-to-normal interval SDNN and a root mean square difference RMSSD of adjacent heartbeat intervals; the calculation formulas are as follows: ; ; ; wherein represents the mean of the RR interval sequence, i is an index symbol in the RR interval sequence; S16, calculate frequency domain features according to the RR interval sequence, including a low frequency feature LF and a high frequency feature HF, and calculate the ratio LF / HF of the two; finally, the time domain features and the frequency domain features are spliced, and a heart rate variability feature vector is obtained through a long short-term memory network model LSTM , .
[0015] Specifically, step S2 specifically includes the following steps: S21, analyze the face video through IPAST, extract 12 eye movement behavior parameters and form an eye movement parameter : , extract 4 public factors with neurophysiological significance, and form a public factor sequence : , wherein represents the task disengagement ability, represents the impulse inhibition, represents the voluntary saccade generation, represents the saccade dynamics; S22, perform principal component analysis on the eye movement parameter Perform principal axis factor extraction and oblique rotation to transform the 12 eye movement behavior parameters into 4 common factors, and construct a common factor matrix for the 12 eye movement behavior parameters to satisfy ,in represents the load matrix, represents the error term; S23, the common factor sequence Input to the gated recurrent unit for spatiotemporal modeling to capture the common factor sequence Dynamic evolution characteristics across time periods; first, the common factor sequence After three convolution blocks, each of which contains a convolution layer and a maximum pooling layer, and then passes through a gated recurrent unit GRU and a fully connected layer FC, the eye movement behavior feature vector is finally obtained. : .
[0016] Specifically, step S3 includes the following steps: S31, the heart rate variability feature vector and eye movement behavior feature vector Input into the multimodal physiological feature fusion network, which includes an attention fusion layer Atten_Layer, a residual block Res_Block, an adaptive feature pooling layer AdaPool_Layer and a regression prediction layer RL; S32, the heart rate variability feature vector and eye movement behavior feature vector Perform vector stitching Get the fusion vector , and then fuse the vector Input into the attention fusion layer Atten_Layer, adaptively learn the importance weights of the two feature vectors, dynamically fuse the heart rate variability features and eye movement behavior features, and obtain the attention fusion vector : ; S33, attention fusion vector through residual block Res_Block Process and obtain residual features : , enhance the model's ability to express complex features and alleviate the gradient disappearance problem; S34, through the adaptive feature pooling layer AdaPool_Layer from the residual features Automatically extract key features : ; S35, finally through the regression prediction layer RL Key features Performing processing to output a physiological state change and an activity prediction value of a Parkinson's disease patient : .
[0017] The Parkinson's disease assessment system based on the fusion of multi-modal physiological features of face videos is used for the Parkinson's disease assessment method based on the fusion of multi-modal physiological features of face videos, and comprises the following steps: A heart rate variability feature (HRV feature) extraction module extracts an rPPG signal based on a continuous face video stream by using an optical flow tracking and face ROI region modeling technology. Then, key parameters of HRV, such as a standard deviation normal to normal interval (SDNN), a root mean square difference of adjacent heartbeat intervals (RMSSD), a low-frequency to high-frequency ratio (LF / HF), and the like, are calculated by using a frequency domain and time domain analysis algorithm to reflect the function state of an autonomic nervous system. To enhance the stability and individual difference adaptability of HRV extraction, a wavelet denoising and band-pass filtering optimization algorithm is introduced, and the adaptation capability to a low signal-to-noise ratio signal is improved by using time series modeling; a structure diagram of the HRV feature extraction module is as shown in Figure 2 ; An eye movement behavior feature extraction module extracts 12 key eye movement behavior parameters based on IPAST analysis of dynamic changes of eye fixation points in a face video, including a percentage of off-center fixation points in a positive saccade task, a percentage of off-center fixation points in a negative saccade task, a number of positive saccades occurring too early before the appearance of a stimulus, a number of negative saccades occurring too early before the appearance of a stimulus, a reaction time of a positive saccade, a reaction time of a negative saccade, a number of fast and correct positive saccades occurring between 90 and 130 milliseconds, a number of fast and incorrect negative saccades occurring between 90 and 130 milliseconds, a number of negative saccade errors within a normal delay range, a time point at which an autonomous reaction starts to overwhelm an automatic reaction, a speed of a correct positive saccade, and a speed of a correct negative saccade. Then, the 12 eye movement parameters are subjected to spatiotemporal modeling by using a convolution gate recurrent unit to obtain eye movement behavior features; a structure diagram of the eye movement behavior feature extraction module is as shown in Figure 3 ; A fusion analysis and assistance module inputs the above-mentioned heart rate variability features and eye movement behavior features into a multi-modal fusion network structure, fuses the splicing vectors of the two by using a transformer, and realizes prediction of Parkinson's disease multi-system involvement features. Finally, a continuous risk score or a disease state label is output, which can be used to assist doctors in judging early Parkinson's risk or tracking intervention.
[0018] The present application is based on facial video data, accurately extracts rPPG signals and further converts into HRV parameters; uses interleaved prosaccade and antisaccade task (IPAST) to extract eye movement behavior characteristics, identifies cognitive and motor disorder characteristics related to Parkinson's disease; fuses two types of key physiological indicators, realizes intelligent evaluation and risk discrimination of Parkinson's disease through modeling.
[0019] The above only describes the preferred embodiments of the present application, and it should be understood that the present application is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, and can be used in various other combinations, modifications and environments, and can be modified within the scope of the concepts described herein, by the above-mentioned teaching or related art or knowledge. The modifications and changes made by those skilled in the art without departing from the spirit and scope of the present application shall be within the scope of protection of the appended claims of the present application.
Claims
1. A Parkinson's disease assessment method based on the fusion of multimodal physiological features of facial videos, characterized by: The following steps are involved: S1. Based on a continuous face video stream, rPPG signals are extracted through optical flow tracking and facial ROI region modeling technology; then key parameters of heart rate variability are calculated to obtain heart rate variability characteristics; S2. Analyze the dynamic changes of eye gaze points in face videos based on IPAST, extract key eye movement behavior parameters, and perform spatiotemporal modeling of the eye movement behavior parameters through convolutional gated recurrent units to obtain eye movement behavior characteristics; the eye movement behavior parameters include the percentage of off-center gaze points in the forward saccade task, the percentage of off-center gaze points in the reverse saccade task, the number of forward saccades that occur prematurely before the stimulus appears, the number of reverse saccades that occur prematurely before the stimulus appears, the reaction time of forward saccades, the reaction time of reverse saccades, rapid correct forward saccades generated between 90-130 milliseconds, rapid incorrect reverse saccades generated between 90-130 milliseconds, reverse saccade errors within a normal latency range, the time point when the autonomous response begins to overwhelm the automatic response, the speed of correct forward saccades, and the speed of correct reverse saccades; S3. Input the heart rate variability characteristics and eye movement behavior characteristics into the multimodal fusion network structure, fuse the concatenated vectors of the two through the transformer, and output a continuous risk score or disease status label to assist doctors in early Parkinson's risk assessment.
2. The Parkinson's disease assessment method based on multimodal physiological feature fusion of face video according to claim 1 is characterized in that: Step S1 specifically includes the following steps: S11, use Dlib detector to detect 81 feature points of the face, then select the forehead and cheeks to get the ROI area, and use sparse optical flow tracking algorithm to stabilize the ROI area. t Frame Image , through the formula Calculate its ROI area in RGB channel The average value of pixels in ,in , M represents the total number of frames, Indicates space, H indicates height, W indicates width, and C indicates the number of channels. Indicates ROI area The pixel coordinates in Indicates ROI area The number of pixels in , 、 and A three-dimensional color time series was constructed to extract the rPPG signal; S12. Use the skin orthogonal plane method (POS) to model the rPPG signal, project the skin reflectance changes in the RGB space to a direction with a high heart rate signal-to-noise ratio, and obtain the original rPPG signal sequence. ; S13, perform wavelet denoising and bandpass filtering optimization processing on the original rPPG signal to eliminate the noise of the original rPPG signal and obtain the rPPG signal sequence ; S14. rPPG signal sequence Perform peak search and calculate RR interval sequence , where N represents the total length of the RR interval sequence; S15. Calculate the time domain features based on the RR interval sequence, including the standard deviation of the normal-to-normal interval SDNN and the root mean square difference (RMSSD) between adjacent heartbeat intervals. The calculation formula is as follows: ; ; ;in represents the mean of the RR interval series, i is the index in the RR interval sequence; S16. Calculate the frequency domain features based on the RR interval sequence, including the low-frequency feature LF and the high-frequency feature HF, and calculate the ratio LF / HF between the two; finally, concatenate the time domain features and the frequency domain features, and obtain the heart rate variability feature vector through the long short-term memory network model LSTM , .
3. The Parkinson's disease assessment method based on multimodal physiological feature fusion of face video according to claim 2 is characterized in that: Step S2 specifically includes the following steps: S21. Analyze facial videos through IPAST, extract 12 eye movement behavior parameters and form eye movement parameters : , extract 4 common factors with neurophysiological significance and form a common factor sequence : ,in Indicates task disengagement capability, Indicates impulse suppression, Indicates voluntary saccade generation, represents the dynamics of saccades; S22. Eye movement parameters Perform principal axis factor extraction and oblique rotation to transform the 12 eye movement behavior parameters into 4 common factors, and construct a common factor matrix for the 12 eye movement behavior parameters to satisfy ,in represents the load matrix, represents the error term; S23, the common factor sequence Input to the gated recurrent unit for spatiotemporal modeling to capture the common factor sequence Dynamic evolution characteristics across time periods; first, the common factor sequence After three convolution blocks, each of which contains a convolution layer and a maximum pooling layer, and then passes through a gated recurrent unit GRU and a fully connected layer FC, the eye movement behavior feature vector is finally obtained. : .
4. The Parkinson's disease assessment method based on multimodal physiological feature fusion of face video according to claim 3 is characterized in that: Step S3 specifically includes the following steps: S31, the heart rate variability feature vector and eye movement behavior feature vector Input into the multimodal physiological feature fusion network, which includes an attention fusion layer Atten_Layer, a residual block Res_Block, an adaptive feature pooling layer AdaPool_Layer and a regression prediction layer RL; S32, the heart rate variability feature vector and eye movement behavior feature vector Perform vector stitching Get the fusion vector , and then fuse the vector Input into the attention fusion layer Atten_Layer, adaptively learn the importance weights of the two feature vectors, dynamically fuse the heart rate variability features and eye movement behavior features, and obtain the attention fusion vector : ; S33, attention fusion vector through residual block Res_Block Process and obtain residual features : , enhance the model's ability to express complex features and alleviate the gradient disappearance problem; S34, through the adaptive feature pooling layer AdaPool_Layer from the residual features Automatically extract key features : ; S35, finally through the regression prediction layer RL Key features Process and output the physiological state changes and activity prediction values of Parkinson's patients : .
5. A Parkinson's disease assessment system based on the fusion of multimodal physiological features of facial videos, used in the Parkinson's disease assessment method based on the fusion of multimodal physiological features of facial videos according to any one of claims 1 to 4, characterized in that: include: A heart rate variability feature extraction module is used to extract rPPG signals based on continuous face video streams and further extract heart rate variability features based on the rPPG signals; The eye movement behavior feature extraction module uses IPAST to analyze the dynamic changes of eye gaze points in facial videos, extract key eye movement behavior parameters, and perform spatiotemporal modeling of eye movement parameters through convolutional gated recurrent units to obtain eye movement behavior features. The fusion analysis and auxiliary module inputs the extracted heart rate variability features and eye movement behavior features into the multimodal fusion network structure, and outputs a continuous risk score or disease status label to assist doctors in early Parkinson's risk assessment.
Citation Information
Patent Citations
Non-contact coronary heart disease evaluation device based on face video
CN116343280A
Non-contact heart rate detection method based on adaptive projection plane and feature screening
CN117678998A
Heartbeat estimation method and device based on face video RPPG signal
CN119214619A
Application of portable three-dimensional eye tracker in cognitive function evaluation
CN119856928A
Long time sequence rPPG signal heart rate variability estimation method and system
CN119969986A
Cited By
Physiological signal visual enhancement method and system based on video decoupling and recombination
CN121810517A
A method and system for visual enhancement of physiological signals based on video decoupling and reconstruction
CN121810517B