A method and system for constructing an intelligent recognition model for Parkinson's disease

By collecting and analyzing facial image video and speech data of Parkinson's patients, an intelligent recognition model is constructed, which solves the problem of insufficient utilization of facial skin muscle characteristics and speech information in the prior art, and achieves accurate identification and evaluation of Parkinson's disease.

CN119887776BActive Publication Date: 2025-09-02SHENZHEN RAPHA BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510376622.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-09-02
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

Existing Parkinson's recognition technology fails to effectively utilize facial skin muscle characteristics and speech information, resulting in low recognition accuracy and long time.

Method used

By collecting facial image videos of Parkinson's patients, analyzing and extracting facial images frame by frame, identifying facial skin features, combining the pronunciation characteristics of speech data, quantifying the degree of facial stiffness, and building an intelligent recognition model.

Benefits of technology

Accurate identification and evaluation of Parkinson's disease is achieved, which improves identification accuracy and shortens identification time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887776B_ABST
    Figure CN119887776B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image recognition technology, and in particular to a method and system for constructing an intelligent recognition model for Parkinson's disease. The method comprises the following steps: collecting facial image videos of Parkinson's patients; analyzing the facial image videos frame by frame, extracting each frame from the facial image videos, and generating frame-by-frame facial images; identifying facial skin features in the frame-by-frame facial images; performing texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; and calculating wrinkle layering on the skin texture wrinkle data to obtain skin wrinkle layering. The present invention uses data processing and image recognition technologies to identify facial skin and muscle features of Parkinson's patients, and correlates and maps the patient's voice information with the presented facial skin and muscle features to construct an intelligent Parkinson's disease recognition model, thereby improving the accuracy of Parkinson's disease recognition and shortening the recognition time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a method and system for constructing an intelligent recognition model for Parkinson's disease. Background Art

[0002] Parkinson's disease patients mainly exhibit characteristics such as resting tremor, muscle rigidity, bradykinesia, postural instability, and gait abnormalities. Early studies mainly used traditional machine learning algorithms, such as decision trees, support vector machines, and Bayesian classifiers, to analyze and classify gait and speech data of Parkinson's patients. These methods showed certain recognition capabilities when the amount of data was limited, but were limited by the complexity of feature extraction and the generalization ability of the model. With the development of medical imaging technology, researchers began to use technologies such as magnetic resonance imaging (MRI) to construct recognition models by analyzing changes in brain structure in Parkinson's patients. In recent years, research has gradually shifted to the integration of multimodal data, including clinical data, imaging data, genetic data, etc., to more comprehensively identify Parkinson's disease and its subtypes. For example, the Deep Phenotype Progression Embedding (DPPE) model combined with long short-term memory (LSTM) units was used to analyze the longitudinal clinical records of Parkinson's patients and identify different disease progression patterns. However, existing Parkinson's disease recognition technology fails to identify the patient's facial skin and muscle features, and it is difficult to associate and map the patient's voice information with the presented facial skin and muscle features, resulting in low accuracy in Parkinson's disease recognition and long recognition time. Summary of the Invention

[0003] Based on this, it is necessary to provide a method and system for constructing an intelligent recognition model for Parkinson's disease to solve at least one of the above technical problems.

[0004] To achieve the above objectives, a method for constructing an intelligent recognition model for Parkinson's disease is provided, the method comprising the following steps:

[0005] Step S1: collecting a facial image video of a Parkinson's disease patient; analyzing the facial image video frame by frame, and extracting each frame image in the facial image video to generate a facial frame-by-frame image;

[0006] Step S2: identifying facial skin features of the facial image frame by frame; performing texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; performing wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determining skin texture smoothness based on the skin wrinkle layering; combining the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data;

[0007] Step S3: collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data; determining facial expression loss features from facial texture index data according to the prosodic feature data, and quantifying facial stiffness data based on the facial expression loss features;

[0008] Step S4: performing Parkinson's disease status assessment based on the facial texture index data and the facial stiffness degree data to obtain Parkinson's disease assessment data; and constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

[0009] By capturing facial video of Parkinson's patients and analyzing and extracting frame-by-frame facial images, the present invention comprehensively captures the dynamic changes in facial expressions, providing rich and continuous image data for subsequent analysis and ensuring the integrity and accuracy of facial feature extraction. The present invention also identifies facial skin features and performs texture wrinkle detection on the frame-by-frame facial images, generating skin texture wrinkle data. This data is then used to measure skin wrinkle layering and determine skin texture smoothness, ultimately combining these data to generate facial texture index data. This process quantifies subtle changes in facial skin, providing objective skin feature indicators for early identification of Parkinson's disease and helping to more accurately reflect the patient's facial muscle condition. Furthermore, the present invention collects speech data from Parkinson's patients and identifies their prosodic features. Based on this data, facial expression loss features are identified and facial stiffness data is quantified. Combining speech prosodic features with facial texture index data enables a comprehensive multi-dimensional assessment of the patient's facial expression state, overcoming the limitations of single-modality data and improving the accuracy of facial stiffness assessment in Parkinson's disease. Parkinson's disease status is assessed based on the facial texture index data and facial stiffness data, generating Parkinson's disease assessment data, which is then used to construct an intelligent Parkinson's disease recognition model. This model integrates multimodal data features to accurately identify and assess Parkinson's disease. Therefore, the present invention uses data processing and image recognition technologies to identify the facial skin and muscle features of Parkinson's patients. It then correlates and maps the patient's voice information with the facial skin and muscle features presented to construct an intelligent Parkinson's disease recognition model, thereby improving the accuracy of Parkinson's disease recognition and shortening the recognition time.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: using a camera to collect data at a frame rate of 30 frames per second, with a camera resolution of no less than 1920×1080 pixels, and a collection time of no less than 3 minutes each time, with the camera and the patient's face kept at a distance of 50-70 cm;

[0012] Step S12: Multi-angle capture is adopted, including front, left and right sides, with at least 1 minute of video captured from each angle to obtain a facial image video of the Parkinson's disease patient;

[0013] Step S13: extracting each frame of the facial image video at a rate of 30 frames per second, and performing grayscale processing on each frame to obtain grayscale image format data;

[0014] Step S14: unify the grayscale image format data into a uniform image resolution and save it as a frame-by-frame facial image.

[0015] The present invention uses a frame rate of 30 frames per second and a resolution of no less than 1920×1080 pixels for video capture, ensuring the clarity and smoothness of facial image videos and fully recording the subtle dynamic changes in facial expressions of Parkinson's patients. Furthermore, the acquisition time of no less than 3 minutes per session and the 50-70 cm distance between the camera and the patient's face ensure the stability and consistency of data acquisition, providing high-quality raw data for subsequent analysis. The multi-angle acquisition method uses frontal, left, and right side views, and captures at least 1 minute of video from each view, fully covering all areas of the Parkinson's patient's face and avoiding information omissions caused by a single viewing angle. This multi-angle acquisition strategy can capture the differences in facial expressions from different angles, providing more comprehensive facial image video data for subsequent analysis. Each frame of the facial image video is extracted at a rate of 30 frames per second and grayscaled to obtain grayscale image format data. Grayscale processing removes color interference, highlighting facial texture and structural features, while reducing data complexity and improving the efficiency and accuracy of subsequent processing. Unifying the image resolution of grayscale image data and saving them as facial frame-by-frame images ensures data consistency and standardization. This unified image resolution facilitates subsequent processing and analysis, while standardized facial frame-by-frame images provide high-quality, uniformly formatted input data for building intelligent recognition models, laying the foundation for their accuracy and reliability.

[0016] Preferably, the step S2 of identifying facial skin features of the facial frame-by-frame images includes:

[0017] Identify facial regions in frame-by-frame facial images;

[0018] Extract skin texture features of the facial area, identify the texture direction and texture density of the skin texture features, and generate skin texture data;

[0019] Perform skin shape recognition on the facial area according to texture direction and texture density to generate skin shape data;

[0020] Extract skin wrinkle features in the facial area, identify the length, depth and distribution direction of skin wrinkle features, and obtain skin wrinkle data;

[0021] Extracting skin sagging features of the facial area, identifying the sagging amount of the skin sagging features, and obtaining skin sagging data;

[0022] Skin texture data, skin shape data, skin wrinkle data and skin sagging data are used to generate facial skin features.

[0023] By accurately identifying facial regions in frame-by-frame facial images, this method ensures targeted and accurate feature extraction, avoids interference from non-facial areas, and provides a clear analysis target for intelligent Parkinson's disease recognition. Extracting skin texture features, particularly texture direction and density, can quantify subtle changes in facial skin. This provides an important objective indicator for early identification of Parkinson's disease, as facial muscle stiffness in Parkinson's patients can lead to abnormal skin texture. Skin shape recognition based on texture direction and density can further reveal the dynamics of facial muscles and changes in facial contours. This multi-dimensional analysis facilitates a more comprehensive assessment of Parkinson's disease patients' facial features, providing rich data support for intelligent recognition models. Extracting skin wrinkle features, including wrinkle length, depth, and distribution, can reflect facial muscle mobility and changes in skin elasticity. Parkinson's disease patients often exhibit facial stiffness and reduced expression, and changes in wrinkle features can be used to quantify this pathological condition. Extracting skin sagging features, particularly the amount of sagging, can reflect changes in facial muscle tension. The relaxation and drooping of facial muscles in Parkinson's patients is one of their typical symptoms. The quantification of this feature provides important pathological information for intelligent recognition models. Integrating multiple skin feature data to generate comprehensive facial skin features can provide multi-dimensional and comprehensive input data for Parkinson's intelligent recognition models.

[0024] Preferably, the step S2 of detecting facial skin features by texture wrinkles includes:

[0025] Detect the texture lines of facial skin features and record the distance between texture lines. When the distance between texture lines is less than 1 mm, it is marked as wrinkle distance data.

[0026] Detect the texture depth of facial skin features. When the texture depth exceeds 0.5 mm, it is marked as wrinkle depth data.

[0027] Detect the texture continuity of facial skin features and record the interruption position of texture continuity. If the interruption position occurs more than twice, it is marked as wrinkle interruption data.

[0028] The skin texture density of facial skin features was detected and the number of texture lines per unit area was recorded. When the texture density exceeded 10 per square centimeter, it was marked as wrinkle density data.

[0029] This invention quantifies subtle changes in facial skin by accurately detecting the spacing between texture lines and marking wrinkle spacing data. Parkinson's patients often experience abnormally reduced skin texture spacing due to facial muscle stiffness. This data provides important pathological features for intelligent recognition models, helping to improve the model's ability to identify Parkinson's. Texture depth detection reflects microstructural changes on the skin surface. Facial skin of Parkinson's patients often develops deeper wrinkles due to muscle stiffness. Marking wrinkle depth data provides the model with a quantitative indicator of skin changes, further enriching the model's input features. Discontinuities in texture continuity reflect the irregularity of skin folds. Folds in Parkinson's patients' facial skin often exhibit interruptions due to muscle movement disorders. Marking wrinkle interruption data captures this irregularity, providing the model with more comprehensive skin feature information and helping to improve recognition accuracy. Texture density detection quantifies the number of texture lines per unit area. Texture density in Parkinson's patients' facial skin can be significantly increased due to muscle stiffness. Marking wrinkle density data provides the model with quantitative information on skin texture changes, further enhancing the model's ability to identify Parkinson's.

[0030] Preferably, the step S2 of calculating the wrinkle layering of the skin texture wrinkle data includes:

[0031] Quantifying the pixel grayscale values ​​of skin texture wrinkle data, dividing the pixel grayscale values ​​into depth levels, and obtaining texture wrinkle level data;

[0032] Determine the wrinkle direction of the texture wrinkle grade data and measure the change angle of the wrinkle direction. When the change angle is less than 30 degrees, it is marked as a normal wrinkle area; when the change angle exceeds 30 degrees, it is marked as a complex wrinkle area.

[0033] The number of wrinkle layers in normal wrinkle areas is identified and recorded as normal wrinkle layers; the number of wrinkle layers in complex wrinkle areas is identified and recorded as complex wrinkle layers;

[0034] The stacking ratio of normal wrinkle layers and complex wrinkle layers is calculated to obtain the skin wrinkle stacking ratio.

[0035] By quantifying the pixel grayscale values ​​of skin texture wrinkle data and classifying it into depth levels, the present invention can convert subtle changes in skin texture into quantifiable texture wrinkle level data. This quantification method provides an objective skin texture depth metric for intelligent Parkinson's disease recognition, helping to distinguish skin texture changes between normal and pathological states. By determining wrinkle direction and measuring the angle of change, it can distinguish between normal and complex wrinkle areas. Because the wrinkle direction variations in the facial skin of Parkinson's patients may be more complex, this differentiation method can more accurately identify lesion features, providing more discriminative input data for the model. Identifying the number of wrinkle layers in different wrinkle areas can further refine the characteristic analysis of skin texture. Distinguishing between normal and complex wrinkle layers can reflect the layered changes in the facial skin of Parkinson's patients, providing richer texture feature information for intelligent recognition models. The skin wrinkle layering ratio, calculated by calculating the layering ratio, can comprehensively reflect the complexity and layering of skin wrinkles. This metric can provide more representative skin texture features for intelligent Parkinson's disease recognition, helping to improve the accuracy and reliability of the model for Parkinson's disease.

[0036] Preferably, determining the skin texture smoothness based on the skin wrinkle layering in step S2 includes:

[0037] Mark the wrinkle contours of skin wrinkles and calculate the length of wrinkle contours;

[0038] Segment the skin wrinkle overlapping area according to the wrinkle contour line length to obtain the wrinkle overlapping segmentation area;

[0039] Identify the wrinkle layer ripple features of the wrinkle stacking segmentation area, and determine the ripple bulge degree according to the wrinkle layer ripple features;

[0040] The regional skin texture ridge degree is mapped based on the ripple ridge degree, and the skin texture smoothness is calculated based on the regional skin texture ridge degree.

[0041] The present invention can accurately quantify the geometric characteristics of skin wrinkles by marking wrinkle contours and calculating their length. This process provides basic data for subsequent wrinkle analysis and can effectively distinguish normal skin texture from abnormal wrinkles caused by Parkinson's disease. Layered region segmentation based on wrinkle contour length can divide complex skin texture areas into sub-regions with different characteristics. This segmentation method can more carefully analyze changes in skin texture and provide more discriminating feature data for the intelligent identification of Parkinson's disease. By identifying the ripple characteristics of the wrinkle layer and determining the degree of ripple ridge, the microstructural changes of skin wrinkles can be further quantified. Determining the degree of ripple ridge provides richer texture features for the intelligent recognition model, which helps to improve the model's accuracy for Parkinson's disease. Mapping the degree of ripple ridge and measuring the skin texture smoothness can comprehensively reflect the overall state of the skin texture. The measurement of skin texture smoothness provides a key quantitative indicator for the intelligent recognition model of Parkinson's disease, allowing for a more comprehensive assessment of the patient's facial skin condition.

[0042] Preferably, the step S3 of collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data includes:

[0043] The microphone sampling rate was set to 44.1-54.1kHz, and the microphone was used to continuously collect speech data from Parkinson's patients.

[0044] The speech data is segmented, with each segment being 5 seconds long and 2 seconds apart, to obtain speech segment data;

[0045] Calculate the fundamental frequency jitter of the speech segment data, and calculate the average value, standard deviation and variation range of the fundamental frequency jitter to generate fundamental frequency jitter data;

[0046] Calculating the intonation intensity of the speech segment data and calculating the average value of the intonation intensity to generate intonation intensity data;

[0047] Calculate the speech duration of the speech segment data, and count the variation range of the speech duration to generate speech duration data;

[0048] The fundamental frequency jitter data, intonation intensity data and speech duration data are used as prosodic feature data.

[0049] By setting a high sampling rate (44.1-54.1kHz), this invention ensures the acquisition of high-quality, high-resolution speech signals, fully capturing the subtle changes in the speech of Parkinson's patients. This high-quality speech data acquisition provides a reliable foundation for subsequent feature extraction and model building. Segmenting the speech data divides the continuous speech signal into statistically significant short segments. This segmentation method effectively captures transient features and prosodic changes in the speech signal while reducing its complexity, facilitating subsequent feature extraction and analysis. Pitch jitter is a common pathological feature of Parkinson's speech, manifesting as unstable and subtle fluctuations in the fundamental frequency. By calculating the mean, standard deviation, and range of pitch jitter, this pathological change in the speech signal can be quantified, providing key prosodic feature data for intelligent recognition models. Intonation intensity reflects the energy and loudness variations of the speech signal. Parkinson's patients often exhibit monotonous intonation and subtle variations in intensity. By calculating the average value of intonation intensity, the patient's speech intensity characteristics can be effectively quantified, further enriching the model's input data. The range of variation in speech duration can reflect the changes in rhythm and fluency of Parkinson's patients' pronunciation. This feature extraction can capture pathological phenomena such as pauses and dragging in the patient's speech, providing the model with important information about speech prosody. Combining these three feature data into prosodic feature data can comprehensively reflect the pathological changes in the speech of Parkinson's patients. This multi-dimensional feature extraction method provides rich input for the intelligent recognition model, helping to improve the model's recognition accuracy and robustness for Parkinson's disease.

[0050] Preferably, the step S3 of determining the facial expression loss feature from the facial texture index data according to the prosody feature data, and quantifying the facial stiffness degree data based on the facial expression loss feature includes:

[0051] detecting the patient's facial movements in the pronunciation state according to the prosodic feature data, and recording the movement amplitude data of the patient's facial movements;

[0052] The movement amplitude data is divided into amplitude levels to obtain the facial movement amplitude level;

[0053] Perform facial expression mapping on the prosody feature data based on the facial movement amplitude level to obtain facial expression features;

[0054] Marking the facial texture index data for texture change missingness according to facial expression features to obtain texture change missingness data; and determining facial detail missingness data according to the texture change missingness data;

[0055] Perform facial expression presentation feature recognition on facial detail missing data to generate facial expression presentation data; determine facial expression missing features based on the facial expression presentation data;

[0056] Quantifying missing features of facial expressions to generate missing feature quantification data;

[0057] The missing feature quantification data is mapped to the degree of facial stiffness to generate facial stiffness data.

[0058] By analyzing the correlation between prosodic feature data and facial movements, this method can capture the facial muscle activity of Parkinson's patients during speech production. Recording facial movement amplitude data provides quantitative muscle movement information for subsequent analysis, helping to reveal the intrinsic connection between facial muscle stiffness caused by Parkinson's disease and speech disorders. Grading the movement amplitude data can transform complex facial movement changes into quantifiable indicators. This grading method provides standardized input for subsequent feature mapping and analysis, further improving the model's ability to recognize Parkinson's facial features. By mapping facial movement amplitude levels with prosodic feature data, pathological information in speech can be combined with facial expression changes. This multimodal feature fusion method can more comprehensively reflect the pathological state of Parkinson's patients, providing richer feature information for intelligent recognition models. By labeling texture changes based on facial expression features, it can accurately identify the loss of facial texture changes caused by Parkinson's disease. This process further refines the analysis of facial features, providing the model with quantitative data on changes in facial details, helping to improve the model's recognition accuracy for Parkinson's disease. Further analysis of the data with missing facial details can identify abnormal facial expressions in Parkinson's patients. Identifying facial expression loss features provides key pathological indicators for the model, helping to more accurately assess the degree of facial stiffness. By quantifying facial expression loss features and mapping them to facial stiffness, complex facial pathological changes can be converted into quantifiable indicators. Facial stiffness data provides important input for intelligent Parkinson's disease recognition models, helping to improve their recognition performance and reliability.

[0059] Preferably, step S4 includes the following steps:

[0060] Step S41: performing wrinkle quantitative analysis on the facial texture index data to obtain Parkinson's disease wrinkle quantitative data; performing stiffness shape analysis on the facial stiffness degree data to obtain Parkinson's disease stiffness shape data;

[0061] Step S42: performing Parkinson's disease status assessment based on the Parkinson's disease wrinkle quantification data and the Parkinson's disease stiffness shape data to obtain Parkinson's disease assessment data;

[0062] Step S43: constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data to obtain an intelligent recognition pre-model; and training the intelligent recognition pre-model based on the Parkinson's disease assessment data to obtain an intelligent recognition training model.

[0063] Step S44: performing model cross-validation evaluation on the intelligent recognition training model to obtain training model evaluation data; adjusting model parameters of the intelligent recognition training model using the training model evaluation data to obtain a Parkinson's disease intelligent recognition model.

[0064] The present invention performs wrinkle quantification analysis on facial texture index data, converting facial skin texture changes into quantifiable Parkinson's wrinkle quantification data. Furthermore, facial stiffness data is analyzed for stiffness shape to obtain Parkinson's stiffness shape data. This data provides objective quantitative indicators of Parkinson's pathological characteristics, laying the foundation for further status assessment. Parkinson's status assessment based on Parkinson's wrinkle quantification data and Parkinson's stiffness shape data can comprehensively reflect the patient's pathological characteristics. This multi-dimensional assessment method can more comprehensively reflect the severity of Parkinson's disease in patients, providing accurate input for the construction of an intelligent recognition model. A Parkinson's disease intelligent recognition pre-model is constructed based on the Parkinson's disease assessment data, and the pre-model is trained based on the assessment data to obtain an intelligent recognition training model. Through model training, the model learns the pathological characteristic patterns of Parkinson's disease, thereby improving the model's recognition ability for Parkinson's disease. Furthermore, data-driven learning during training effectively enhances the model's generalization ability. Cross-validation evaluation of the intelligent recognition training model comprehensively evaluates the model's performance and stability. The training model evaluation data obtained through cross-validation can provide a scientific basis for adjusting model parameters. Adjusting model parameters based on this evaluation data can further optimize model performance, ultimately resulting in an intelligent Parkinson's disease recognition model with high accuracy and strong generalization capabilities.

[0065] The present invention also provides a system for constructing an intelligent recognition model for Parkinson's disease, which is used in the above-mentioned method for constructing an intelligent recognition model for Parkinson's disease. The system for constructing an intelligent recognition model for Parkinson's disease includes:

[0066] A facial image video acquisition module is used to collect facial image videos of Parkinson's patients; the facial image video is analyzed frame by frame, and each frame image in the facial image video is extracted to generate a frame-by-frame facial image;

[0067] The facial texture feature recognition module is used to identify facial skin features in facial images frame by frame; perform texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; perform wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determine skin texture smoothness based on the skin wrinkle layering; combine the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data;

[0068] A facial stiffness quantification module is used to collect speech data from Parkinson's patients and identify rhythmic feature data of the speech data; determine facial expression loss features based on facial texture index data according to the rhythmic feature data, and quantify facial stiffness data based on the facial expression loss features;

[0069] The Parkinson's disease intelligent recognition model construction module is used to evaluate the Parkinson's disease status based on facial texture index data and facial stiffness data to obtain Parkinson's disease assessment data; and to build a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

[0070] The present invention collects facial image videos of Parkinson's patients and generates facial frame-by-frame images, providing high-quality raw data for subsequent analysis and ensuring the integrity and accuracy of facial feature extraction; identifies skin texture features of facial frame-by-frame images, generates skin texture fold data, wrinkle layering, and texture smoothness, and merges them into facial texture index data to achieve quantitative analysis of facial skin features of Parkinson's patients; collects speech data and identifies rhythmic features, determines facial expression loss features based on facial texture index data, quantifies facial stiffness data, and achieves multimodal fusion analysis of speech and facial features; evaluates Parkinson's status based on facial texture index data and facial stiffness data, constructs an intelligent recognition model, and achieves accurate status assessment of Parkinson's disease. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 A schematic flow chart of a method for constructing an intelligent recognition model for Parkinson's disease;

[0072] Figure 2 for Figure 1 Detailed implementation steps of step S1 in FIG.

[0073] Figure 3 for Figure 1 Detailed implementation steps of step S4 in FIG.

[0074] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0075] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

[0076] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.

[0077] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0078] To achieve this, please refer to Figures 1 to 3 A method and system for constructing an intelligent recognition model for Parkinson's disease, the method comprising the following steps:

[0079] Step S1: collecting a facial image video of a Parkinson's disease patient; analyzing the facial image video frame by frame, and extracting each frame image in the facial image video to generate a facial frame-by-frame image;

[0080] Step S2: identifying facial skin features of the facial image frame by frame; performing texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; performing wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determining skin texture smoothness based on the skin wrinkle layering; combining the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data;

[0081] Step S3: collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data; determining facial expression loss features from facial texture index data according to the prosodic feature data, and quantifying facial stiffness data based on the facial expression loss features;

[0082] Step S4: performing Parkinson's disease status assessment based on the facial texture index data and the facial stiffness degree data to obtain Parkinson's disease assessment data; and constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

[0083] By capturing facial video of Parkinson's patients and analyzing and extracting frame-by-frame facial images, the present invention comprehensively captures the dynamic changes in facial expressions, providing rich and continuous image data for subsequent analysis and ensuring the integrity and accuracy of facial feature extraction. The present invention also identifies facial skin features and performs texture wrinkle detection on the frame-by-frame facial images, generating skin texture wrinkle data. This data is then used to measure skin wrinkle layering and determine skin texture smoothness, ultimately combining these data to generate facial texture index data. This process quantifies subtle changes in facial skin, providing objective skin feature indicators for early identification of Parkinson's disease and helping to more accurately reflect the patient's facial muscle condition. Furthermore, the present invention collects speech data from Parkinson's patients and identifies their prosodic features. Based on this data, facial expression loss features are identified and facial stiffness data is quantified. Combining speech prosodic features with facial texture index data enables a comprehensive multi-dimensional assessment of the patient's facial expression state, overcoming the limitations of single-modality data and improving the accuracy of facial stiffness assessment in Parkinson's disease. Parkinson's disease status is assessed based on the facial texture index data and facial stiffness data, generating Parkinson's disease assessment data, which is then used to construct an intelligent Parkinson's disease recognition model. This model integrates multimodal data features to accurately identify and assess Parkinson's disease. Therefore, the present invention uses data processing and image recognition technologies to identify the facial skin and muscle features of Parkinson's patients. It then correlates and maps the patient's voice information with the facial skin and muscle features presented to construct an intelligent Parkinson's disease recognition model, thereby improving the accuracy of Parkinson's disease recognition and shortening the recognition time.

[0084] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a flow chart showing the steps of a method for constructing an intelligent recognition model for Parkinson's disease according to the present invention. In this example, the method for constructing an intelligent recognition model for Parkinson's disease includes the following steps:

[0085] Step S1: collecting a facial image video of a Parkinson's disease patient; analyzing the facial image video frame by frame, and extracting each frame image in the facial image video to generate a facial frame-by-frame image;

[0086] In this embodiment of the present invention, a high-resolution camera (with a resolution of no less than 1920×1080 pixels) is used to capture a video of the patient's face at a frame rate of 30 frames per second. The camera should be positioned directly in front of the patient to ensure a clear and unobstructed facial image in the video. Uniform lighting should also be ensured to prevent image quality from being affected by lighting variations. The captured video data is then imported into a video processing system for preliminary preprocessing, including noise removal and brightness and contrast adjustments to enhance image quality. Furthermore, the video is cropped to remove irrelevant background information, retaining only the facial region. Using video processing software, images are extracted frame by frame according to the video's frame rate. Specifically, the frame-by-frame extraction parameters are set, and each frame of the video is extracted sequentially and stored as a separate image file. During the extraction process, the image resolution and color depth are maintained consistent to ensure accuracy in subsequent processing. After extraction, all images are arranged in chronological order to form a frame-by-frame facial image sequence. Each image file should include a timestamp to ensure accurate correspondence to a specific time point in the video during subsequent analysis. Furthermore, the extracted images are uniformly formatted to ensure consistent image size, facilitating subsequent feature extraction and analysis.

[0087] Step S2: identifying facial skin features of the facial image frame by frame; performing texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; performing wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determining skin texture smoothness based on the skin wrinkle layering; combining the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data;

[0088] In this embodiment of the present invention, facial images are grayscaled frame by frame, converting them from RGB color format to grayscale format to highlight texture information and reduce data size. Subsequently, the grayscale image is smoothed using a Gaussian filter with a kernel size of 5×5 and a standard deviation of 1.5 to remove noise and ensure accuracy in subsequent processing. The smoothed grayscale image is processed using the Canny edge detection algorithm to extract wrinkle edges in the skin texture. In the Canny algorithm, a low threshold of 50 and a high threshold of 150 are set, and wrinkle edge information is extracted through dual-threshold detection. The detection result is a binary image, in which wrinkle edges are represented by white pixels and non-edge areas by black pixels. Furthermore, facial landmark detection technology is used to locate wrinkle-concentrated areas such as the forehead, eye area, and mouth corners, and these areas are analyzed to improve wrinkle detection accuracy. The grayscale standard deviation of the detected wrinkle areas is calculated to quantify the density and depth of the wrinkles. The specific operation is as follows: first, the grayscale value distribution within the wrinkle area is statistically analyzed, and the mean and standard deviation of the grayscale values ​​in this area are calculated. A larger standard deviation indicates a higher degree of wrinkle layering, meaning denser or deeper wrinkles. Skin texture smoothness is further assessed based on wrinkle layering data. Skin texture smoothness is quantified by calculating the grayscale variance of the wrinkle region. A smaller grayscale variance indicates a smoother skin texture; conversely, a larger variance indicates a rougher skin texture. This metric reflects the overall smoothness of skin texture. Wrinkle layering and skin texture smoothness are combined to generate a facial texture index. Specifically, the wrinkle layering data is multiplied by a weighting factor of 0.6, and the skin texture smoothness data is multiplied by a weighting factor of 0.4. The two are then added together to produce the final facial texture index.

[0089] Step S3: collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data; determining facial expression loss features from facial texture index data according to the prosodic feature data, and quantifying facial stiffness data based on the facial expression loss features;

[0090] In this embodiment of the present invention, speech data from Parkinson's disease patients is first collected. Patients are asked to complete speech tasks such as sustained vowels (such as / a / , / i / , / u / ), repeated syllables (such as / pa / , / ta / ), and situational dialogues (such as reading aloud a specified sentence). The speech acquisition equipment uses a high-quality microphone to ensure a quiet recording environment to reduce background noise interference. The collected speech data undergoes preprocessing, including noise reduction and normalization, to improve the quality and consistency of the speech signal. Next, prosodic feature recognition is performed on the collected speech data. Acoustic analysis tools are used to extract prosodic features from the speech, including parameters such as intonation, intensity, and duration. Prosodic features of the speech are quantified by calculating metrics such as the range of intonation fluctuation, the stability of intensity, and the rate of change of duration. For example, Parkinson's disease patients often exhibit monotonous intonation, minimal intensity variation, and significant duration variation. Subsequently, facial texture index data is analyzed based on the prosodic feature data to identify features of facial expression loss. The specific operation is to conduct a correlation analysis between speech prosodic features and facial texture indicators to identify features of stiff or missing facial expressions caused by abnormal speech prosody. For example, when speech prosodic features show a single tone and large variations in tone length, this may correspond to stiff facial expressions or a lack of natural dynamic changes. Finally, the facial stiffness data is quantified based on the facial expression loss features. By setting a threshold, the facial expression loss features are mapped to the degree of facial stiffness. For example, when the facial expression loss features exceed the preset threshold, the facial stiffness is judged to be high; conversely, the stiffness is low. The facial stiffness data is expressed in numerical form and used for subsequent Parkinson's disease identification analysis.

[0091] Step S4: performing Parkinson's disease status assessment based on the facial texture index data and the facial stiffness degree data to obtain Parkinson's disease assessment data; and constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

[0092] In this embodiment of the present invention, Parkinson's disease status is first assessed based on the facial texture index data and facial stiffness data obtained in step S3. The facial texture index data and facial stiffness data are fused and weighted summed to generate Parkinson's disease assessment data. Specific parameters are set as follows: a weight of 0.5 for the facial texture index data and a weight of 0.5 for the facial stiffness data. This weighted summation yields comprehensive assessment data, which is used to quantify the patient's Parkinson's disease status. Subsequently, an intelligent Parkinson's disease recognition model is constructed based on the Parkinson's disease assessment data. The collected multimodal data (including speech data and facial texture data) is preprocessed to extract key features. Prosodic features are extracted from the speech data, while texture and stiffness features are extracted from the facial texture data. These features are normalized to ensure data consistency. The processed data is then divided into a training set and a test set, with 70% of the data used as the training set and 30% as the test set. The training set is used to train the model, using the support vector machine (SVM) algorithm as the classifier. In the SVM model, the penalty parameter C was set to 9.6057, and the kernel parameter g was set to 0.5261. During training, 10-fold cross-validation was used to evaluate model performance and ensure generalization. Finally, the trained model was validated using a test set to evaluate metrics such as accuracy, sensitivity, and F1 score. Both accuracy and sensitivity were high, and the F1 score was used to comprehensively evaluate model performance. In this way, an intelligent Parkinson's disease recognition model based on multimodal data was constructed.

[0093] As an example of the present invention, refer to Figure 2 As shown, in this example, step S1 includes:

[0094] Step S11: using a camera to collect data at a frame rate of 30 frames per second, with a camera resolution of no less than 1920×1080 pixels, and a collection time of no less than 3 minutes each time, with the camera and the patient's face kept at a distance of 50-70 cm;

[0095] Step S12: Multi-angle capture is adopted, including front, left and right sides, with at least 1 minute of video captured from each angle to obtain a facial image video of the Parkinson's disease patient;

[0096] Step S13: extracting each frame of the facial image video at a rate of 30 frames per second, and performing grayscale processing on each frame to obtain grayscale image format data;

[0097] Step S14: unify the grayscale image format data into a uniform image resolution and save it as a frame-by-frame facial image.

[0098] In an embodiment of the present invention, first, a high-definition camera is used for video capture. The camera's frame rate is set to 30 frames per second and the resolution is 1920×1080 pixels to ensure that the captured images are clear and rich in details. The acquisition time is no less than 3 minutes each time, and the distance between the camera and the patient's face is maintained between 50-70 cm to ensure the integrity and consistency of the facial image. During the acquisition process, the camera is fixed on a tripod to avoid image blurring due to device shaking. Next, a multi-angle acquisition method is adopted, including the front, left side, and right side. At least 1 minute of video is captured at each angle to ensure that the patient's facial features are fully captured from different perspectives. During the acquisition process, the patient maintains a natural expression and avoids large movements or changes in expression to reduce interference factors. Subsequently, the captured facial image video is extracted frame by frame. Each frame of the video is extracted in sequence at a rate of 30 frames per second, and each frame is grayscaled. Grayscale processing is achieved by converting color images into grayscale images. Specifically, each pixel value of the RGB image is converted to a grayscale value. The grayscale value is calculated as follows: Grayscale value = 0.299 × R + 0.587 × G + 0.114 × B, where R, G, and B represent the pixel values ​​of the red, green, and blue channels, respectively. Finally, the grayscale image format data is uniformly processed, and the image resolution is adjusted to match the acquisition resolution, that is, 1920 × 1080 pixels. Bilinear interpolation is used during resolution adjustment to ensure image clarity during scaling. The processed images are saved as facial frame-by-frame images in PNG format.

[0099] Preferably, the step S2 of identifying facial skin features of the facial frame-by-frame images includes:

[0100] Identify facial regions in frame-by-frame facial images;

[0101] Extract skin texture features of the facial area, identify the texture direction and texture density of the skin texture features, and generate skin texture data;

[0102] Perform skin shape recognition on the facial area according to texture direction and texture density to generate skin shape data;

[0103] Extract skin wrinkle features in the facial area, identify the length, depth and distribution direction of skin wrinkle features, and obtain skin wrinkle data;

[0104] Extracting skin sagging features of the facial area, identifying the sagging amount of the skin sagging features, and obtaining skin sagging data;

[0105] Skin texture data, skin shape data, skin wrinkle data and skin sagging data are used to generate facial skin features.

[0106] In this embodiment of the present invention, face detection technology is used to process captured facial images frame by frame to locate the facial region. OpenFace software or similar facial detection tools are used to identify facial key points (such as the eyes, nose, and mouth) in the image and extract a rectangular region encompassing the entire face. The extracted facial region is then grayscaled, converting the color image to a grayscale image to highlight texture information. Gray-level co-occurrence matrix (GLCM) analysis is used to calculate texture direction and density. Texture direction is determined by analyzing the principal directions of the GLCM, while texture density is quantified by counting the frequency of texture repetition. Skin texture data is generated based on texture direction and density. Texture direction is expressed as an angle (e.g., 0°, 45°, 90°), and texture density is expressed in pixels for subsequent analysis. Skin shape features are identified by combining texture features with facial key point location information. The smoothness and symmetry of the facial contour are assessed by calculating the relative distances and angles between key points to generate skin shape data. An edge detection algorithm (such as the Canny algorithm) is used to identify wrinkle edges and calculate their length, depth, and distribution. Wrinkle length is measured by measuring the length of the wrinkle edge, depth is calculated by analyzing the grayscale variation difference within the wrinkle area, and distribution direction is determined by analyzing the direction of the wrinkles, resulting in skin wrinkle data. Skin laxity is quantified by calculating the vertical displacement of key facial points (such as the eyelids and corners of the mouth). For example, the amount of eyelid droop or mouth corner droop is measured and expressed in pixels to obtain skin laxity data. Skin texture data, skin shape data, skin wrinkle data, and skin laxity data are integrated to form facial skin features. After normalization, these data serve as input features for the Parkinson's disease intelligent recognition model.

[0107] Preferably, the step S2 of detecting facial skin features by texture wrinkles includes:

[0108] Detect the texture lines of facial skin features and record the distance between texture lines. When the distance between texture lines is less than 1 mm, it is marked as wrinkle distance data.

[0109] Detect the texture depth of facial skin features. When the texture depth exceeds 0.5 mm, it is marked as wrinkle depth data.

[0110] Detect the texture continuity of facial skin features and record the interruption position of texture continuity. If the interruption position occurs more than twice, it is marked as wrinkle interruption data.

[0111] The skin texture density of facial skin features was detected and the number of texture lines per unit area was recorded. When the texture density exceeded 10 per square centimeter, it was marked as wrinkle density data.

[0112] In this embodiment of the present invention, facial images are grayscaled frame by frame, converting color images into grayscale images to highlight texture information. An edge detection algorithm (such as the Canny algorithm) is used to identify skin texture lines and calculate the spacing between texture lines. If the spacing between texture lines is less than 1 mm, the area is recorded as wrinkle spacing data. Texture spacing is detected by calculating the pixel distance between adjacent texture lines. The pixel distance corresponding to 1 mm is converted based on the image resolution. Local binary pattern (LBP) analysis is performed on the grayscale image to extract texture depth information. Texture depth is quantified by calculating the grayscale gradient of the texture area. If the texture depth exceeds 0.5 mm, the area is marked as wrinkle depth data. Texture depth quantification is based on the intensity of grayscale changes. The grayscale change threshold corresponding to 0.5 mm is determined through experimental calibration. A texture tracking algorithm is used to analyze the continuity of texture lines and record the locations of texture interruptions. If the interruption occurs more than twice, the area is marked as wrinkle interruption data. Texture continuity is detected by analyzing the directional consistency of texture lines. The interruption locations are identified based on directional abrupt changes or grayscale discontinuities. The number of texture lines per unit area is counted to calculate texture density. When the texture density exceeds 10 lines per square centimeter, the area is marked as wrinkle density data. The number of texture lines per unit area is calculated by counting and converting the texture lines within a predefined skin area. The threshold of 10 lines per square centimeter was set based on the skin characteristics of Parkinson's patients.

[0113] Preferably, the step S2 of calculating the wrinkle layering of the skin texture wrinkle data includes:

[0114] Quantifying the pixel grayscale values ​​of skin texture wrinkle data, dividing the pixel grayscale values ​​into depth levels, and obtaining texture wrinkle level data;

[0115] Determine the wrinkle direction of the texture wrinkle grade data and measure the change angle of the wrinkle direction. When the change angle is less than 30 degrees, it is marked as a normal wrinkle area; when the change angle exceeds 30 degrees, it is marked as a complex wrinkle area.

[0116] The number of wrinkle layers in normal wrinkle areas is identified and recorded as normal wrinkle layers; the number of wrinkle layers in complex wrinkle areas is identified and recorded as complex wrinkle layers;

[0117] The stacking ratio of normal wrinkle layers and complex wrinkle layers is calculated to obtain the skin wrinkle stacking ratio.

[0118] In this embodiment of the present invention, facial images are grayscaled frame by frame, converting color images into grayscale images. Image processing software or tools are used to quantitatively analyze each pixel in the grayscale image. Texture wrinkles are classified into different depth levels based on the range of pixel grayscale values. For example, areas with grayscale values ​​between 0 and 50 are labeled as shallow wrinkles, areas with grayscale values ​​between 51 and 150 are labeled as medium wrinkles, and areas with grayscale values ​​between 151 and 255 are labeled as deep wrinkles. Image analysis tools such as OpenCV or MATLAB are used to calculate the principal direction of the wrinkle region. Wrinkle direction information is quantified by calculating the gradient direction of wrinkle lines or detecting line direction using the Hough transform. When the angle of change in wrinkle direction is less than 30 degrees, the region is labeled as a normal wrinkle region; when the angle of change exceeds 30 degrees, it is labeled as a complex wrinkle region. Wrinkle layer counts are performed for both normal and complex wrinkle regions. An edge detection algorithm (such as the Canny algorithm) is used to extract wrinkle edges, and the number of wrinkle layers within each region is counted. For normal wrinkle areas, the number of wrinkle layers is recorded as the normal wrinkle layer number; for complex wrinkle areas, the number of wrinkle layers is recorded as the complex wrinkle layer number. The normal and complex wrinkle layers are compared and analyzed. The ratio of the complex wrinkle layer number to the normal wrinkle layer number is calculated to obtain the skin wrinkle layering degree. This ratio reflects the complexity of skin wrinkles and can be used in subsequent Parkinson's disease identification and analysis.

[0119] Preferably, determining the skin texture smoothness based on the skin wrinkle layering in step S2 includes:

[0120] Mark the wrinkle contours of skin wrinkles and calculate the length of wrinkle contours;

[0121] Segment the skin wrinkle overlapping area according to the wrinkle contour line length to obtain the wrinkle overlapping segmentation area;

[0122] Identify the wrinkle layer ripple features of the wrinkle stacking segmentation area, and determine the ripple bulge degree according to the wrinkle layer ripple features;

[0123] The regional skin texture ridge degree is mapped based on the ripple ridge degree, and the skin texture smoothness is calculated based on the regional skin texture ridge degree.

[0124] In an embodiment of the present invention, an image processing tool (such as OpenCV) is used to process a skin wrinkle layering image. First, wrinkle contours are extracted using the Canny edge detection algorithm. The extracted contours are then labeled using the findContours function, and the length of each wrinkle contour is calculated using the arcLength function. Contour length is used as a quantitative indicator of wrinkle contour length, and the contour length data for each wrinkle region is recorded. The wrinkle layering region is then segmented based on the wrinkle contour length. A length threshold is set; for example, when the wrinkle contour length exceeds a preset value (such as 10 pixels), the region is classified as an independent wrinkle layering region. The wrinkle region is segmented using a semantic segmentation algorithm (such as U-Net or ENet), and the segmented regions are labeled as wrinkle layering segmentation regions. The segmentation results are saved as binary images, where white areas represent wrinkle layering regions and black areas represent non-wrinkle regions. The segmented wrinkle layering regions are then analyzed for ripple characteristics. The texture features of wrinkle areas are analyzed using local binary patterns (LBP) or gray-level co-occurrence matrices (GLCMs). The degree of undulation within the wrinkle layer is identified by calculating the grayscale gradient of the texture features. When the grayscale gradient exceeds a certain threshold (e.g., a grayscale value change exceeding 50), the area is marked as undulating. The degree of undulation is quantified by the magnitude of the grayscale gradient, and the undulation data for each overlapping wrinkle region are recorded. Skin texture undulation mapping is performed for each overlapping wrinkle region based on the undulation degree. Regions with higher undulation degrees are marked as high-undulation regions, and those with lower undulation degrees are marked as low-undulation regions. Skin texture smoothness is assessed by calculating the area percentage of high-undulation regions. When the area percentage of high-undulation regions exceeds a certain threshold (e.g., 20%), the skin texture smoothness of that region is considered low.

[0125] Preferably, the step S3 of collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data includes:

[0126] The microphone sampling rate was set to 44.1-54.1kHz, and the microphone was used to continuously collect speech data from Parkinson's patients.

[0127] The speech data is segmented, with each segment being 5 seconds long and 2 seconds apart, to obtain speech segment data;

[0128] Calculate the fundamental frequency jitter of the speech segment data, and calculate the average value, standard deviation and variation range of the fundamental frequency jitter to generate fundamental frequency jitter data;

[0129] Calculating the intonation intensity of the speech segment data and calculating the average value of the intonation intensity to generate intonation intensity data;

[0130] Calculate the speech duration of the speech segment data, and count the variation range of the speech duration to generate speech duration data;

[0131] The fundamental frequency jitter data, intonation intensity data and speech duration data are used as prosodic feature data.

[0132] In this embodiment of the present invention, a microphone sampling rate of 44.1-54.1kHz is used to continuously collect speech data from Parkinson's disease patients in a quiet environment using a high-quality microphone. The collected speech data includes sustained vowels (such as / a / , / i / , / u / ), repeated syllables (such as / pa / , / ta / ), and contextual dialogue, ensuring the integrity and clarity of the speech signal. The collected speech data is segmented, with each segment length set to 5 seconds and an interval of 2 seconds between segments. Using speech processing software or tools, the continuous speech signal is segmented into multiple 5-second segments to generate speech segment data. Jitter is then calculated for the segmented speech data. First, the fundamental frequency curve of the speech signal is extracted, and then the relative jitter value of the fundamental frequency curve is calculated. The specific operation includes smoothing the fundamental frequency curve and then interpolating it. The residual between the original fundamental frequency curve and the smoothed curve is calculated to obtain the fundamental frequency jitter value. The mean, standard deviation, and variation range of the fundamental frequency jitter are calculated to generate fundamental frequency jitter data. Intonation intensity analysis is performed on the speech segment data. The short-time energy of the speech signal is calculated to determine the average intonation intensity of each speech segment. Short-time energy is calculated by calculating the sum of the squared amplitudes of all sampling points in each frame of the speech signal and taking the average as the intonation intensity. This ultimately generates intonation intensity data. Speech duration analysis is performed on the speech segment data. The actual duration of each speech segment is calculated to determine the variation range of the speech duration. The speech duration calculation formula is: speech duration (seconds) = (speech data length / sampling rate). The duration data of each speech segment is recorded and its variation range is calculated to generate speech duration data. The fundamental frequency jitter data, intonation intensity data, and speech duration data are integrated into prosodic feature data. This data reflects the pitch and prosodic characteristics of the speech signal.

[0133] Preferably, the step S3 of determining the facial expression loss feature from the facial texture index data according to the prosody feature data, and quantifying the facial stiffness degree data based on the facial expression loss feature includes:

[0134] detecting the patient's facial movements in the pronunciation state according to the prosodic feature data, and recording the movement amplitude data of the patient's facial movements;

[0135] The movement amplitude data is divided into amplitude levels to obtain the facial movement amplitude level;

[0136] Perform facial expression mapping on the prosody feature data based on the facial movement amplitude level to obtain facial expression features;

[0137] Marking the facial texture index data for texture change missingness according to facial expression features to obtain texture change missingness data; and determining facial detail missingness data according to the texture change missingness data;

[0138] Perform facial expression presentation feature recognition on facial detail missing data to generate facial expression presentation data; determine facial expression missing features based on the facial expression presentation data;

[0139] Quantifying missing features of facial expressions to generate missing feature quantification data;

[0140] The missing feature quantification data is mapped to the degree of facial stiffness to generate facial stiffness data.

[0141] In an embodiment of the present invention, face detection technology (such as a deep learning-based face detection algorithm) is used to process captured facial images frame by frame to locate key facial points (such as eyebrows and mouth corners). By analyzing the displacement changes of key points during pronunciation, facial movements during pronunciation are detected. The displacement amplitudes of the key points are calculated and quantified into movement amplitude data in pixels. Based on the magnitude of the movement amplitude data, facial movement amplitudes are classified into different levels. For example, areas with movement amplitudes less than 10 pixels are labeled as low, 10-20 pixels as medium, and greater than 20 pixels as high. The facial movement amplitude levels are correlated with prosodic feature data (such as fundamental frequency jitter and voice intensity). For example, when fundamental frequency jitter is high and the facial movement amplitude is low, it is mapped to a tense or stiff expression; when voice intensity is high and the facial movement amplitude is high, it is mapped to an active expression. The facial expression features are then compared with facial texture indicators (such as skin texture smoothness and wrinkle layering). When facial expression features indicate stiffness, areas with high wrinkle stacking or low skin texture smoothness in the facial texture index data are marked as texture change missing data. The texture change missing data is analyzed to identify areas with missing facial detail. For example, when the wrinkle stacking exceeds a preset threshold (e.g., more than 15 wrinkle layers per square centimeter) and the skin texture smoothness falls below a threshold (e.g., smoothness less than 0.3), the area is identified as facial detail missing data. Facial expression recognition technology (e.g., a convolutional neural network-based facial expression recognition model) is used to analyze the facial detail missing data and identify expression presentation features. For example, facial muscle stiffness or a monotonous expression can be identified during pronunciation. Expression missing features are analyzed based on the facial expression presentation data. For example, if the facial expression presentation data shows a monotonous expression with a lack of dynamic change during pronunciation, this is identified as a facial expression missing feature. The facial expression missing features are then quantitatively analyzed. For example, the area percentage and duration of the expression missing region are calculated to generate quantitative data for the missing features. Quantitative data is expressed as a percentage, such as if the facial area lacks expression, accounting for 30% of the total facial area. Based on the quantitative data of missing features, the degree of facial stiffness is categorized as mild, moderate, or severe. For example, when the facial area lacks expression, it is labeled as mild stiffness; 20%-40% is labeled as moderate stiffness; and more than 40% is labeled as severe stiffness.

[0142] More importantly, mapping the missing feature quantification data to the degree of facial stiffness includes:

[0143] Detect missing feature quantification data for missing action duration;

[0144] Detect missing motion amplitude values ​​of missing feature quantization data;

[0145] Determine the frequency of missing action occurrence based on the missing action duration and missing action amplitude values;

[0146] The missing action duration, missing action amplitude value, and missing action occurrence frequency were marked as facial stiffness indicators respectively;

[0147] The facial stiffness indexes were weighted and calculated to obtain the degree of facial stiffness.

[0148] In this embodiment of the present invention, time series analysis is performed on the quantified missing feature data to extract the duration of the missing motion. By analyzing the timestamps of the facial expression data, the start and end times of each missing motion are calculated, thereby obtaining the duration of the missing motion. For example, if a missing motion begins at the 5th second and lasts until the 10th second, the duration of the motion is 5 seconds. Amplitude analysis is also performed on the quantified missing feature data to extract the amplitude of the missing motion. The displacement of key facial points (such as eyebrows and mouth corners) during the motion is calculated to quantify the amplitude of the missing motion. For example, if a key point moves 15 pixels from its initial position during the motion, the amplitude of the motion is 15 pixels. The number of occurrences of the missing motion per unit time is counted to determine its frequency. For example, if a missing motion occurs 6 times during a 30-second observation period, the frequency of the motion is 0.2 times per second. By comparing the duration and amplitude of the missing motion, its potential impact on Parkinson's disease can be further assessed. These three parameters are labeled as facial stiffness indicators for subsequent analysis. The duration of a missing action reflects its persistence, the amplitude of a missing action reflects its intensity, and the frequency of a missing action reflects its frequency. Together, these metrics constitute the quantitative characteristics of facial stiffness. The facial stiffness indices are weighted to determine the final degree of facial stiffness. For example, the weight of the duration of a missing action is set to 0.4, the weight of the amplitude of a missing action is set to 0.3, and the weight of the frequency of a missing action is set to 0.3.

[0149] As an example of the present invention, refer to Figure 3 As shown, in this example, step S4 includes:

[0150] Step S41: performing wrinkle quantitative analysis on the facial texture index data to obtain Parkinson's disease wrinkle quantitative data; performing stiffness shape analysis on the facial stiffness degree data to obtain Parkinson's disease stiffness shape data;

[0151] Step S42: performing Parkinson's disease status assessment based on the Parkinson's disease wrinkle quantification data and the Parkinson's disease stiffness shape data to obtain Parkinson's disease assessment data;

[0152] Step S43: constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data to obtain an intelligent recognition pre-model; and training the intelligent recognition pre-model based on the Parkinson's disease assessment data to obtain an intelligent recognition training model.

[0153] Step S44: performing model cross-validation evaluation on the intelligent recognition training model to obtain training model evaluation data; adjusting model parameters of the intelligent recognition training model using the training model evaluation data to obtain a Parkinson's disease intelligent recognition model.

[0154] In an embodiment of the present invention, image processing technology is used to perform wrinkle quantification analysis on collected facial texture index data. By calculating the grayscale gradient and texture density of the wrinkle area, wrinkle depth and distribution density are quantified. For example, wrinkle depth can be measured by the intensity of grayscale changes, and texture density can be calculated by the number of wrinkle lines per unit area. This ultimately yields Parkinson's wrinkle quantification data. Shape analysis is performed on the facial stiffness data to identify the contours and shape characteristics of the stiff areas. The shape characteristics of the stiff areas are quantified by calculating the area, perimeter, and shape factors (such as circularity or rectangularity) of the stiff areas. This ultimately yields Parkinson's stiffness shape data. Parkinson's disease status is assessed based on the Parkinson's disease wrinkle quantification data and the stiffness shape data. The wrinkle quantification data and stiffness shape data are comprehensively analyzed, and a threshold (e.g., wrinkle depth exceeding 0.5 mm, stiffness area percentage exceeding 10%) is set to determine the patient's Parkinson's disease status. This ultimately yields Parkinson's disease assessment data. Based on this Parkinson's disease assessment data, an intelligent recognition pre-model is constructed. A support vector machine (SVM) was selected as the classifier, with a Gaussian radial basis kernel (RBF) kernel function, a penalty parameter C set to 1.0, and a kernel parameter γ set to 0.1. The evaluation data was divided into a training set and a test set, with the training set used for model training and the test set used for model validation. The training set was used to train the intelligent recognition pre-model. Model performance was optimized by adjusting model parameters (such as the SVM C value and γ value). The intelligent recognition training model was ultimately obtained and cross-validated. A 10-fold cross-validation method was used to calculate the model's accuracy, recall, and F1 score. Finally, evaluation data for the training model was obtained. Based on this training model evaluation data, parameters of the intelligent recognition training model were adjusted. For example, the SVM C value was adjusted to 2.0 and the γ value to 0.05 to improve the model's generalization ability and prediction accuracy. The resulting Parkinson's disease intelligent recognition model was obtained.

[0155] More importantly, step S42 includes the following steps:

[0156] Step S421: Calculate the four parameters of the wrinkle region texture feature contrast, correlation, energy, and homogeneity of the Parkinson's disease wrinkle quantification data, and calculate the mean and standard deviation of each parameter to obtain a wrinkle quantification feature vector;

[0157] Step S422: Calculate the joint angle change data of the Parkinson's disease stiff shape data, adjust the range of the joint angle change data to between 0 and 1, calculate the maximum angle change range, average angle change rate, and standard deviation of the angle change for each joint within 10 seconds, and obtain the stiff shape feature vector;

[0158] Step S423: setting the weight of the wrinkle quantization feature vector to 0.4 and the weight of the rigid shape feature vector to 0.6; and performing weighted summation on the wrinkle quantization feature vector and the rigid shape feature vector to obtain Parkinson's disease state score data;

[0159] Step S424: performing a score level evaluation on the Parkinson's disease status score data to obtain Parkinson's disease evaluation data.

[0160] In this embodiment of the present invention, the gray-level co-occurrence matrix (GLCM) method is used to calculate the texture features of Parkinson's disease wrinkle quantification data. Specific parameters include contrast, correlation, energy, and homogeneity. Contrast reflects texture clarity, correlation measures the linear relationship between texture elements, energy indicates texture uniformity, and homogeneity reflects texture consistency. The mean and standard deviation of each parameter are calculated. For example, the mean contrast is 0.45 and the standard deviation is 0.12; the mean correlation is 0.80 and the standard deviation is 0.05; the mean energy is 0.30 and the standard deviation is 0.08; and the mean homogeneity is 0.75 and the standard deviation is 0.06. This ultimately generates a wrinkle quantification feature vector. Motion capture technology or video analysis methods are used to record the angular changes of key facial points (such as eyebrows and mouth corners). The joint angle change data is normalized to a range between 0 and 1. The maximum angular change range, average angular change rate, and standard deviation of angular change for each joint within 10 seconds are calculated. For example, the maximum angle change range is 0.8, the average angle change rate is 0.05, and the standard deviation of the angle change is 0.02. Finally, a rigid shape feature vector is formed; the weight of the wrinkle quantization feature vector is set to 0.4, and the weight of the rigid shape feature vector is set to 0.6; the wrinkle quantization feature vector and the rigid shape feature vector are weighted and summed to obtain the Parkinson's disease state score data. For example, the value of the wrinkle quantization feature vector is 0.5, and the value of the rigid shape feature vector is 0.7; the Parkinson's disease state score data is graded. A threshold interval is set, for example: a score less than 0.5 is mild, a score between 0.5 and 0.7 is moderate, and a score greater than 0.7 is severe; based on the score interval, a state score of 0.62 is assessed as a moderate Parkinson's disease state, and the Parkinson's disease assessment data is finally obtained.

[0161] The present invention also provides a system for constructing an intelligent recognition model for Parkinson's disease, which is used in the above-mentioned method for constructing an intelligent recognition model for Parkinson's disease. The system for constructing an intelligent recognition model for Parkinson's disease includes:

[0162] A facial image video acquisition module is used to collect facial image videos of Parkinson's patients; the facial image video is analyzed frame by frame, and each frame image in the facial image video is extracted to generate a frame-by-frame facial image;

[0163] The facial texture feature recognition module is used to identify facial skin features in facial images frame by frame; perform texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; perform wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determine skin texture smoothness based on the skin wrinkle layering; combine the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data;

[0164] A facial stiffness quantification module is used to collect speech data from Parkinson's patients and identify rhythmic feature data of the speech data; determine facial expression loss features based on facial texture index data according to the rhythmic feature data, and quantify facial stiffness data based on the facial expression loss features;

[0165] The Parkinson's disease intelligent recognition model construction module is used to evaluate the Parkinson's disease status based on facial texture index data and facial stiffness data to obtain Parkinson's disease assessment data; and to build a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

[0166] The present invention collects facial image videos of Parkinson's patients and generates facial frame-by-frame images, providing high-quality raw data for subsequent analysis and ensuring the integrity and accuracy of facial feature extraction; identifies skin texture features of facial frame-by-frame images, generates skin texture fold data, wrinkle layering, and texture smoothness, and merges them into facial texture index data to achieve quantitative analysis of facial skin features of Parkinson's patients; collects speech data and identifies rhythmic features, determines facial expression loss features based on facial texture index data, quantifies facial stiffness data, and achieves multimodal fusion analysis of speech and facial features; evaluates Parkinson's status based on facial texture index data and facial stiffness data, constructs an intelligent recognition model, and achieves accurate status assessment of Parkinson's disease.

[0167] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0168] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing an intelligent recognition model for Parkinson's disease, characterized in that: The following steps are involved: Step S1: collecting facial image videos of Parkinson's disease patients; Analyze the facial image video frame by frame, extract each frame of the facial image video, and generate a frame-by-frame facial image; Step S2: identifying facial skin features of the facial image frame by frame; performing texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; performing wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determining skin texture smoothness based on the skin wrinkle layering; combining the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data; Step S3: collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data; determining facial expression loss features from facial texture index data according to the prosodic feature data, and quantifying facial stiffness data based on the facial expression loss features; Step S4: performing Parkinson's disease status assessment based on the facial texture index data and the facial stiffness degree data to obtain Parkinson's disease assessment data; and constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

2. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: using a camera to collect data at a frame rate of 30 frames per second, with a camera resolution of no less than 1920×1080 pixels, and a collection time of no less than 3 minutes each time, with the camera and the patient's face kept at a distance of 50-70 cm; Step S12: Multi-angle capture is adopted, including front, left and right sides, with at least 1 minute of video captured from each angle to obtain a facial image video of the Parkinson's disease patient; Step S13: extracting each frame of the facial image video at a rate of 30 frames per second, and performing grayscale processing on each frame to obtain grayscale image format data; Step S14: unify the grayscale image format data into a uniform image resolution and save it as a frame-by-frame facial image.

3. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: The step S2 of identifying facial skin features of the facial frame-by-frame images includes: Identify facial regions in frame-by-frame facial images; Extract skin texture features of the facial area, identify the texture direction and texture density of the skin texture features, and generate skin texture data; Perform skin shape recognition on the facial area according to texture direction and texture density to generate skin shape data; Extract skin wrinkle features in the facial area, identify the length, depth and distribution direction of skin wrinkle features, and obtain skin wrinkle data; Extracting skin sagging features of the facial area, identifying the sagging amount of the skin sagging features, and obtaining skin sagging data; Skin texture data, skin shape data, skin wrinkle data and skin sagging data are used to generate facial skin features.

4. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: The step S2 of detecting facial skin features by texture wrinkles includes: Detect the texture lines of facial skin features and record the distance between texture lines. When the distance between texture lines is less than 1 mm, it is marked as wrinkle distance data. Detect the texture depth of facial skin features. When the texture depth exceeds 0.5 mm, it is marked as wrinkle depth data. Detect the texture continuity of facial skin features and record the interruption position of texture continuity. If the interruption position occurs more than twice, it is marked as wrinkle interruption data. The skin texture density of facial skin features is detected, and the number of texture lines per unit area is recorded. When the texture density exceeds 10 per square centimeter, it is marked as wrinkle density data.

5. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: The step S2 of calculating the wrinkle layering degree of the skin texture wrinkle data includes: Quantifying the pixel grayscale values ​​of skin texture wrinkle data, dividing the pixel grayscale values ​​into depth levels, and obtaining texture wrinkle level data; Determine the wrinkle direction of the texture wrinkle grade data and measure the change angle of the wrinkle direction. When the change angle is less than 30 degrees, it is marked as a normal wrinkle area; when the change angle exceeds 30 degrees, it is marked as a complex wrinkle area. The number of wrinkle layers in normal wrinkle areas is identified and recorded as normal wrinkle layers; the number of wrinkle layers in complex wrinkle areas is identified and recorded as complex wrinkle layers; The stacking ratio of normal wrinkle layers and complex wrinkle layers is calculated to obtain the skin wrinkle stacking ratio.

6. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: Determining the skin texture smoothness based on the skin wrinkle layering in step S2 includes: Mark the wrinkle contours of skin wrinkles and calculate the length of wrinkle contours; Segment the skin wrinkle overlapping area according to the wrinkle contour line length to obtain the wrinkle overlapping segmentation area; Identify the wrinkle layer ripple features of the wrinkle stacking segmentation area, and determine the ripple bulge degree according to the wrinkle layer ripple features; The regional skin texture ridge degree is mapped based on the ripple ridge degree, and the skin texture smoothness is calculated based on the regional skin texture ridge degree.

7. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: The step S3 of collecting speech data of a Parkinson's disease patient and identifying prosodic feature data of the speech data includes: The microphone sampling rate was set to 44.1-54.1kHz, and the microphone was used to continuously collect speech data from Parkinson's patients. The speech data is segmented, with each segment being 5 seconds long and 2 seconds apart, to obtain speech segment data; Calculate the fundamental frequency jitter of the speech segment data, and calculate the average value, standard deviation and variation range of the fundamental frequency jitter to generate fundamental frequency jitter data; Calculating the intonation intensity of the speech segment data and calculating the average value of the intonation intensity to generate intonation intensity data; Calculate the speech duration of the speech segment data, and count the variation range of the speech duration to generate speech duration data; The fundamental frequency jitter data, intonation intensity data and speech duration data are used as prosodic feature data.

8. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: Determining the facial expression loss feature from the facial texture index data according to the prosody feature data in step S3, and quantifying the facial stiffness degree data based on the facial expression loss feature includes: detecting the patient's facial movements in the pronunciation state according to the prosodic feature data, and recording the movement amplitude data of the patient's facial movements; The movement amplitude data is divided into amplitude levels to obtain the facial movement amplitude level; Perform facial expression mapping on the prosody feature data based on the facial movement amplitude level to obtain facial expression features; Marking the facial texture index data for texture change missingness according to facial expression features to obtain texture change missingness data; and determining facial detail missingness data according to the texture change missingness data; Perform facial expression presentation feature recognition on facial detail missing data to generate facial expression presentation data; determine facial expression missing features based on the facial expression presentation data; Quantifying missing features of facial expressions to generate missing feature quantification data; The missing feature quantification data is mapped to the degree of facial stiffness to generate facial stiffness data.

9. The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: performing wrinkle quantitative analysis on the facial texture index data to obtain Parkinson's disease wrinkle quantitative data; performing stiffness shape analysis on the facial stiffness degree data to obtain Parkinson's disease stiffness shape data; Step S42: performing Parkinson's disease status assessment based on the Parkinson's disease wrinkle quantification data and the Parkinson's disease stiffness shape data to obtain Parkinson's disease assessment data; Step S43: constructing a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data to obtain an intelligent recognition pre-model; and training the intelligent recognition pre-model based on the Parkinson's disease assessment data to obtain an intelligent recognition training model. Step S44: performing model cross-validation evaluation on the intelligent recognition training model to obtain training model evaluation data; adjusting model parameters of the intelligent recognition training model using the training model evaluation data to obtain a Parkinson's disease intelligent recognition model.

10. A system for constructing an intelligent recognition model for Parkinson's disease, characterized in that: The method for constructing an intelligent recognition model for Parkinson's disease according to claim 1 is used to construct a system for constructing an intelligent recognition model for Parkinson's disease, comprising: A facial image video acquisition module is used to collect facial image videos of Parkinson's patients; the facial image video is analyzed frame by frame, and each frame image in the facial image video is extracted to generate a frame-by-frame facial image; The facial texture feature recognition module is used to identify facial skin features in facial images frame by frame; perform texture wrinkle detection on the facial skin features to generate skin texture wrinkle data; perform wrinkle layering measurement on the skin texture wrinkle data to obtain skin wrinkle layering; determine skin texture smoothness based on the skin wrinkle layering; combine the skin wrinkle layering and skin texture smoothness as texture indices to obtain facial texture index data; A facial stiffness quantification module is used to collect speech data from Parkinson's patients and identify prosodic feature data of the speech data; determine facial expression loss features based on facial texture index data according to the prosodic feature data, and quantify facial stiffness data based on the facial expression loss features; The Parkinson's disease intelligent recognition model construction module is used to evaluate the Parkinson's disease status based on facial texture index data and facial stiffness data to obtain Parkinson's disease assessment data; and to build a Parkinson's disease intelligent recognition model based on the Parkinson's disease assessment data.

Citation Information

Patent Citations

  • Method for detecting facial expression addiction of Parkinson patient

    CN111210415A

  • Parkinson's disease early detection method and system based on multi-mode deep learning

    CN118136232A