Alzheimer disease early screening method fusing multi-modal biological characteristics

By integrating multimodal biometrics, processing and analyzing facial, iris, and voice data, and combining an adaptive weight allocation algorithm, the adaptability of early Alzheimer's disease screening methods in contexts of fluctuating states and grassroots settings was addressed, achieving highly accurate screening results.

CN121867696APending Publication Date: 2026-04-17HAINAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HAINAN UNIV
Filing Date
2026-01-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing early screening methods for Alzheimer's disease are weak in resisting interference when faced with fluctuations in the patient's condition, and are not well adapted to grassroots settings. This leads to the intrusion of non-pathological noise and the masking or misjudgment of early and mild pathological features. Furthermore, it is difficult to conduct effective screening using simple equipment.

Method used

A multimodal biometrics fusion approach is adopted, which acquires dynamic sequence images of the face, static images of the iris, and target topic speech description data. Adaptive Wiener filtering, correction algorithms, and spectral subtraction are used to process noise, and features such as fixation duration, eye saccade response delay time, and amplitude of involuntary facial muscle tremors are extracted. Combined with a pathological association adaptive weight allocation algorithm and a multimodal fusion classification model, a screening report is generated.

Benefits of technology

It significantly reduces the risk of misdiagnosis and missed diagnosis of early, mild pathological features, improves the accuracy and reliability of screening results, meets the application needs of grassroots scenarios, and achieves effective screening in resource-limited environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121867696A_ABST
    Figure CN121867696A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of early screening of neurodegenerative diseases, and particularly relates to an Alzheimer's disease early screening method fusing multi-modal biological characteristics. The method comprises the following steps: firstly, acquiring face dynamic, iris static and voice audio data of a to-be-screened object, and processing through a modal specific noise reduction algorithm to generate a standardized feature set; secondly, extracting cognitive state characteristics such as staying duration of a gazing target, comparing the cognitive state characteristics with a preset threshold value, and outputting a suitability judgment result and adjusting guidance; and finally, generating a weighted target feature set through an adaptive weight distribution algorithm according to a judgment result, and inputting the weighted target feature set into a pre-training model to complete cross-modal fusion and risk level output. According to the method, a state adaptation judgment and dynamic weight adjustment mechanism is additionally arranged, so that non-pathological interference is effectively filtered, basic simple acquisition equipment is adapted, the precision and standardization of the whole screening process are realized, and the early screening accuracy and the basic practicability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of early screening technology for neurodegenerative diseases, specifically to an early screening method for Alzheimer's disease that integrates multimodal biomarkers. Background Technology

[0002] Alzheimer's disease is a neurological disorder characterized by progressive cognitive decline and neurodegenerative changes. Its pathological process is insidious and irreversible. Early symptoms often include mild memory fluctuations, visuospatial perception deviations, or decreased language coherence, which can be easily confused with normal aging symptoms. The effectiveness of disease intervention is closely related to the timing of early identification. Early detection can effectively slow the progression and improve the quality of life for patients. Therefore, developing scientific early screening methods for Alzheimer's disease is of great significance for achieving convenient and large-scale early identification of the disease and reducing the social medical burden.

[0003] However, changes in the state of early-stage Alzheimer's patients, such as fatigue and mood swings, can cause non-pathological fluctuations in their biometrics. Existing technologies use a fixed pattern of static feature stitching combined with direct analysis, without setting up a cognitive state adaptation judgment step or quantitatively correcting feature interference for state fluctuations. This results in non-pathological noise being mixed into the analysis, causing early mild pathological features to be masked or misjudged. At the same time, existing technologies rely on professional equipment such as eye trackers to collect features, without adapting to simple equipment such as ordinary cameras and mobile phone microphones at the grassroots level. There is neither a dedicated solution to process the low-quality data collected by these devices, nor a feature completion mechanism under equipment limitations, making it difficult for the technology to be implemented in grassroots scenarios with limited resources.

[0004] Therefore, this invention proposes an early screening method for Alzheimer's disease that integrates multimodal biometrics. Summary of the Invention

[0005] To address the technical problems mentioned in the background section regarding the weak resistance to fluctuations in patient condition and insufficient adaptability to grassroots settings in existing early screening methods for Alzheimer's disease, the present invention aims to provide an early screening method for Alzheimer's disease that integrates multimodal biometrics.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] Early screening methods for Alzheimer's disease that integrate multimodal biometrics include:

[0008] S1: Obtain the raw multimodal biometric data of the target object and process the data to generate a standardized feature set;

[0009] S2: Based on the standardized feature set, extract three types of cognitive state features: fixation duration, eye saccade response delay time, and unconscious facial muscle tremor amplitude. Compare these features with preset screening and adaptation thresholds, and output the cognitive state judgment result and corresponding adjustment guidance.

[0010] S3: Based on the cognitive state determination result, the pathological association adaptive weight allocation algorithm is used to quantify the degree of interference of patient state fluctuations on each modality feature and dynamically adjust the weights to generate a weighted target feature set that is adapted to the current state.

[0011] S4: Input the weighted target feature set into the pre-trained multimodal fusion classification model, and then... The image-based feature extraction algorithm parses spatial pathological features, combines them with linguistic pathological features to complete cross-modal fusion, and outputs a screening report.

[0012] Furthermore, the multimodal biometric raw data in S1 includes dynamic sequence images of the face, static images of the iris, and target topic speech description data;

[0013] To address motion blur interference in the dynamic face image sequence, an adaptive Wiener filtering algorithm is used; to address uneven illumination noise in the static iris image, an adaptive Wiener filtering algorithm is used. The correction algorithm is used to suppress environmental noise interference in the target topic speech description data. First, a short-time Fourier transform is performed on the target topic speech description data to obtain the corresponding speech power spectrum. Then, spectral subtraction is used for noise reduction.

[0014] Subsequently, from the deblurred dynamic sequence image of the face Extract gaze stability features The static iris image after illumination correction Extracting iris tremor frequency features From the denoised speech power spectrum Extract semantic coherence features from the corresponding speech content ;

[0015] The processed original multimodal biometric data Convert to 256-dimensional feature vector And associated with the three types of core pathological features extracted accordingly. Structured splicing is performed to generate a 771-dimensional standardized feature set. ,and ;in, Represents the feature vector of a dynamic sequence of human faces; Represents the feature vector of a static iris image; This represents the speech power spectrum feature vector.

[0016] Furthermore, the specific method for extracting and calculating cognitive state features in S2 is as follows: based on the feature vector of the dynamic sequence image of the face. Analyze pupil position trajectory data, using The algorithm locates the pupil center coordinates in each frame and calculates the cumulative time when the pupil center falls within three preset 100×100 pixel visual target areas as the fixation duration. The pupil movement speed is calculated using first-order difference, and the starting and arriving frames of saccades are identified with a speed threshold of 500 pixels / second. The average delay of at least three effective saccades is calculated as the eye saccade response delay time. Ten key facial feature points are extracted, and the mean standard deviation of the displacement of each feature point during the acquisition period is calculated as the amplitude of the involuntary facial muscle tremors.

[0017] Furthermore, the cognitive state determination and adjustment guidance in S2 specifically includes: setting a preset screening adaptation threshold, wherein the adaptation threshold for the duration of gaze on the target is: The threshold for adaptation of eye saccade response delay time is: The threshold for the amplitude of involuntary facial muscle tremors ;

[0018] When all three features—the duration of fixation on the target, the delay in eye saccade response, and the amplitude of involuntary facial muscle tremors—meet their corresponding threshold requirements, the system is considered to be in a compatible state and a quantitative value for the feature is output. If any one of the features fails to meet the threshold, the system is considered to be in an incompatible state. Adjustment guidelines are generated for different incompatible types, including rest duration, environmental adjustments, and suggestions for eye activities. If two or more features are incompatible, it is recommended to restart the complete screening process the next day.

[0019] Furthermore, the core logic of the pathological association adaptive weight allocation algorithm in S3 is: to use the standardized feature set... The features are divided into three categories: dynamic facial sequence image features, static iris image features, and speech / audio features; a gaze duration influence coefficient is also defined. Influence coefficient of eye saccade response delay parameter The influence coefficient of the amplitude of involuntary facial muscle tremors The influence coefficients of the three states and their corresponding value ranges are used to quantify the degree of interference of the three cognitive state features on the weights of the three feature groups. Correlation with the threshold of fixation duration Correlation with eye saccade response delay time adaptation threshold It is associated with the threshold for the amplitude of involuntary facial muscle tremors.

[0020] Furthermore, the process of generating the weighted target feature set in S3 is as follows: based on the influence coefficient of the fixation target dwell time. Influence coefficient of eye saccade response delay parameter The influence coefficient of the amplitude of involuntary facial muscle tremors Calculate the weight adjustment factor for each feature group;

[0021] Through formula The adjustment factor of the feature group of the dynamic sequence image of the face was calculated. ,in The enhanced interference cancellation coefficient; obtained through the formula The adjustment factor of the feature group of the static iris image was calculated. ,in For weakened interference cancellation coefficient; for tone feature group adjustment factor Less affected by temporary states, maintaining the baseline weight;

[0022] Based on the calculated feature group weight adjustment factor , and The final weights of the feature groups in the dynamic sequence of facial images were calculated after normalization. Final weights of iris static image feature groups Final weights of speech audio feature groups The weights of the three types of feature groups;

[0023] Finally, the standardized feature set is... The three types of feature groups are multiplied by their corresponding final weights to achieve feature weighting, and then concatenated in the original dimensional order to generate the weighted target feature set. ,and ;in, This is a feature set related to dynamic facial sequence images, including gaze stability features. and basic feature vectors ; This is a set of features related to static iris images, including iris tremor frequency features. and basic feature vectors ; This is a group of speech and audio-related features, including semantic coherence features. and basic feature vectors .

[0024] Furthermore, based on the weighted target feature set The feature encoding and spatial pathological feature parsing in S4 specifically involve: processing the weighted target feature set... Weighted facial dynamic feature group The face is encoded into a 128-dimensional face encoding vector through a 3-layer convolutional neural network. For weighted static iris feature groups It is encoded into a 64-dimensional iris coding vector through a 2-layer fully connected network. For weighted speech audio feature groups It is encoded into a 128-dimensional speech coding vector through a single-layer long short-term memory network. ;

[0025] use Image feature extraction algorithms, from the face encoding vector The gaze trajectory dispersion is calculated. Coordination with facial movements Two types of indicators; from the iris coding vector The irregularity of iris texture is calculated. With the coefficient of variation of tremor frequency Two types of indicators.

[0026] Furthermore, the cross-modal fusion and screening report generation in S4 specifically involves: from the speech coding vector Extract semantic break frequency With vocabulary richness Two types of core indicators are used to construct spatial pathological feature vectors. Language pathology feature vector The two types of vectors are mapped to 128 dimensions through a fully connected layer. The spatial and linguistic feature weights are calculated based on the attention weight formula, and the weighted sum is used to obtain a 256-dimensional global feature vector. ;

[0027] global feature vectors Input the classification layer of the pre-trained multimodal fusion classification model, and then... The function outputs the risk probability and combines it with a preset threshold to determine the risk level, while also marking abnormal pathological features. The screening report includes six core modules: basic information, review of cognitive status determination, weight allocation results, pathological feature analysis, risk level conclusion, and data traceability.

[0028] Compared with the prior art, the advantages of the present invention are as follows:

[0029] 1. This invention adds a cognitive state suitability determination step before multimodal feature analysis. By extracting features such as fixation target dwell time and eye saccade response delay time and comparing them with preset thresholds, it accurately identifies whether the patient's current state is suitable for screening. At the same time, based on the determination results, a pathological association adaptive weight allocation algorithm is used to quantify the degree of interference of state fluctuations on facial, iris, and voice features and dynamically adjust feature weights, effectively filtering non-pathological noise, significantly reducing the risk of misjudgment and missed judgment of early minor pathological features, and improving the accuracy and reliability of screening results.

[0030] 2. This invention designs a specific solution for simple devices such as ordinary cameras and mobile phone microphones in basic scenarios, using adaptive Wiener filtering, Modality-specific noise reduction algorithms, such as correction and spectral subtraction, are used to process motion-blurred face images, unevenly lit iris images, and noisy speech data, respectively, ensuring the effectiveness of features in low-quality data. At the same time, through a multimodal feature structured splicing and dynamic weight allocation mechanism, when a certain type of modality data cannot be obtained due to equipment limitations, it can be effectively screened by supplementing it with other modal features, meeting the application needs of resource-limited scenarios such as grassroots communities and homes, and helping the technology to be deployed on a large scale.

[0031] 3. In the feature processing stage, this invention standardizes multimodal data into a feature set of a unified dimension, combined with... Image feature extraction algorithms and Technologies such as attention mechanisms are used to accurately quantify key pathological indicators such as gaze trajectory dispersion, iris texture irregularity, and semantic breakage frequency; in the report generation stage, abnormal features are annotated based on the reference range of healthy individuals, and through... The function outputs a clear risk level determination, forming a standardized process from data collection, status determination, feature fusion to result output. This not only provides traceable quantitative evidence for clinical diagnosis, but also improves the consistency of screening results in different scenarios. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a schematic diagram of the workflow of the method of the present invention;

[0034] Figure 2 This is a schematic diagram of the cognitive state adaptability determination and adjustment guidance workflow of the present invention;

[0035] Figure 3 This is a schematic diagram illustrating the multimodal feature weighted fusion and risk level determination of the present invention. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] To achieve the above objectives, this invention provides an early screening method for Alzheimer's disease that integrates multimodal biometrics, such as... Figures 1-3 As shown, the method includes:

[0038] S1: Obtain the raw multimodal biometric data of the target object and process the data to generate a standardized feature set.

[0039] S101: The multimodal biometric raw data includes dynamic sequence images of the face, static images of the iris, and target topic speech description data;

[0040] Use resolution ≥720 A standard camera is aimed at the entire facial area of ​​the person to be screened, and continuously captures 10 seconds of dynamic facial images at a frame rate of 30 frames per second to obtain the dynamic sequence of facial images. ;

[0041] After completing the acquisition of the dynamic facial image sequence, a camera of the same specifications is used. The shooting angle and focus are adjusted to focus on the iris region of the subject's eye. Once the iris image is clear, the camera is triggered to capture a single shot, continuously acquiring a 2-second dynamic iris image sequence at a frame rate of 30 frames per second. The frame with the highest clarity is selected as the static iris image. And based on the 2-second dynamic sequence, the displacement sequence of the iris edge sampling points is extracted;

[0042] With a sampling rate ≥16 Using the phone's microphone, a specified topic description instruction is given to the person to be screened (such as "Describe your daily routine from leaving home to arriving at your destination"). Simultaneously, the microphone's recording function is activated, continuously collecting 30 seconds of voice response content to obtain the target topic voice description data. ;

[0043] S102: To address motion blur interference in the aforementioned dynamic facial image sequence, an adaptive Wiener filtering algorithm is employed for processing. The calculation formula is as follows:

[0044]

[0045] in, To produce a dynamic sequence of deblurred facial images; This is a regularization parameter used to balance the deblurring effect with the preservation of image details; This is a blur kernel function used to quantify the distribution of motion blur in dynamic facial image sequences. , The regularization coefficient is . For fuzzy kernel gradient;

[0046] In this embodiment, based on the practical description of adaptive Wiener filtering in "Digital Image Processing (Third Edition)", the regularization parameter... The range of values ​​is set to ;

[0047] S103: To address the uneven illumination noise in the static iris image, the following measures are taken: The correction algorithm performs suppression, and the calculation formula is: ;in, This is a static image of the iris after illumination correction. For correction factors; This indicates the original iris image. pixel values Exponential operations are used to adjust the brightness uniformity of the original iris image;

[0048] In this embodiment, based on "Computer Vision: Algorithms and Applications" The correction parameters are set, and the correction coefficient is... Set the value range to ;when When it is active, it brightens darker areas in the image while suppressing overly bright areas, improving details in dark areas; when This means that the corrected iris image is completely identical to the original iris image. The correction will not produce any adjustment effect on the image; it is equivalent to not performing any illumination correction. This process darkens brighter areas in the image while enhancing contrast in dark areas and optimizing details in bright areas, ultimately resulting in a more uniformly illuminated corrected static iris image. ;

[0049] S104: To address environmental noise interference in the target topic speech description data, firstly, the target topic speech description data... Perform a short-time Fourier transform to obtain the speech description data of the target topic. Corresponding speech power spectrum Then, spectral subtraction is used for denoising, and the calculation formula is:

[0050]

[0051] in, The speech power spectrum corresponding to the denoised target topic speech description data; The environmental noise power spectrum is obtained by collecting audio data during quiet periods in the scene to be screened and calculating it through frequency domain analysis;

[0052] After denoising, the speech power spectrum is... The signal is converted back to the time domain speech signal through inverse short-time Fourier transform, and then processed by sentence segmentation and part-of-speech tagging.

[0053] S105: Based on the processed multimodal biometric raw data, calculate three core pathological association features related to early Alzheimer's disease: gaze stability features, iris tremor frequency features, and semantic coherence features.

[0054] S1051: From the deblurred dynamic sequence image of the face Extract gaze stability features The calculation formula is:

[0055]

[0056] in, The data collection duration (in seconds); for The distance (in pixels) of the gaze point relative to the target region at any given time; the gaze point passes through... Algorithm detects and determines pupil position; The time interval between adjacent fixations (in seconds);

[0057] In this embodiment, based on the "Technical Specifications for Eye Tracking in Cognitive Function Assessment", the data collection duration is... The time interval between adjacent gaze points is 10 seconds. It is 1 / 30 of a second, matching the acquisition frame rate of 30 frames per second;

[0058] S1052: From the illumination-corrected static image of the iris Extracting iris tremor frequency features The calculation formula is:

[0059]

[0060] in, This represents the number of sampling points at the edge of the iris. For the first The jitter frequency (in Hz) at the nth sampling point is the i-th The main frequency is obtained by fast Fourier transform of the displacement sequence of each sampling point within a 2-second dynamic sequence;

[0061] In this embodiment, based on the "Technical Specification for Iris Biometric Acquisition and Processing", the number of iris edge sampling points is 300, which are evenly distributed between the inner and outer edges of the iris;

[0062] S1053: From the power spectrum of the denoised speech Extract semantic coherence features from the corresponding speech content The calculation formula is: ;in, The number of semantically coherent statements used to describe the topic; The total number of statements describing the topic;

[0063] In this embodiment, the number of semantically coherent statements when describing the topic... The judgment criteria are grammatical completeness of the statement, logical coherence, and no deviation from the specified topic, through pre-training. The model calculates the semantic similarity between sentences, and a similarity of ≥0.7 is considered coherent. The similarity threshold of 0.7 refers to the recommended standard in the "Guidelines for Semantic Coherence Assessment in Natural Language Processing" 2024; the total number of sentences describing the topic... Determined by combining punctuation mark segmentation with speech pause detection;

[0064] S106: Process the original multimodal biometric data. Convert it to a 256-dimensional feature vector using existing feature extraction methods. And associated with the three types of core pathological features extracted accordingly. Structured splicing is performed to generate a 771-dimensional standardized feature set. ,and ;in, This represents the 256-dimensional feature vector corresponding to the deblurred dynamic sequence image of a face; This represents the 256-dimensional feature vector corresponding to the illumination-corrected static image of the iris; This represents the 256-dimensional feature vector corresponding to the power spectrum of the denoised speech.

[0065] S2: Based on the standardized feature set, extract three types of cognitive state features: gaze duration, eye saccade response delay, and facial muscle involuntary tremor amplitude. Compare these features with preset screening and adaptation thresholds, and output the cognitive state judgment result and corresponding adjustment guidance.

[0066] The duration of fixation on the target and the delay in eye saccade response are associated with impairment of attention and visuospatial function in patients with cognitive impairment, and the amplitude of involuntary facial muscle tremors is associated with abnormal neuronal function in the brain. All three characteristics are potential early biomarkers for Alzheimer's disease.

[0067] S201: Based on feature vectors of dynamic facial sequence images Analyze pupil position trajectory data, using The algorithm locates the pupil center coordinates in each frame of the image. , The frame number is... This corresponds to a 10-second acquisition duration and a frame rate of 30 frames per second; simultaneously, three visual target areas are preset. , , Each region is The rectangular regions of pixels are evenly distributed across the captured image, consistent with the image capture scene in S1; the duration of gaze on the target is... Defined as the cumulative time it takes for the pupil center to fall within any target area, the formula is:

[0068]

[0069] in, For frame status identifier, if the first Frame pupil center Landing in the target area , , If any one of them is in the middle, then ,otherwise ; The duration of a single frame (matching the frame rate captured in S101);

[0070] In this embodiment, based on standard parameter settings in the field of visual data acquisition, the duration of a single frame is... The value is 1 / 30 of a second;

[0071] S202: The eye saccade response delay time is defined as the motion process in which the pupil center rapidly moves from one target area to another, based on the pupil center coordinates. Identify the scanning start frame and the target area arrival frame, and calculate the scanning delay time;

[0072] First, the pupillary motion velocity is calculated using the first-order difference. :

[0073]

[0074] in, For the first Frame and the Pupil motion speed between frames (unit: pixels / second); , For the first Frame pupil center coordinates; , For the first Frame pupil center coordinates;

[0075] When the pupil movement speed greater than the speed threshold At that time, it is determined to be the starting frame of the saccade motion. When the pupil center first enters the new target area after a saccade, it is determined to be an arriving frame. The saccade response delay time is the average delay of multiple saccade processes, calculated as follows:

[0076]

[0077] in, The time delay for eye saccades (in seconds); The number of valid scans identified during the data collection period; For the first The arrival frame number of the next scan; For the first The starting frame number of the next scan;

[0078] In this embodiment, based on the quantitative conclusion in "Eye-Tracking Technology Psychology" that saccadic motion speed can reach 500 degrees / second, and combined with the conversion relationship between screen pixels and viewing angle, 500 pixels / second is a typical dividing value within the range, which can effectively distinguish saccadic motion from normal fixation fluctuations. Therefore, the speed threshold is... Values Pixels per second; Based on the general guidelines for cognitive assessment of neurodegenerative diseases, eye movement parameters must be measured at least three times to reduce random errors and ensure data reliability. Therefore, the number of effective saccades is... If the data is collected less than 3 times, it is considered invalid and step S1 must be repeated as instructed. When re-collecting, the visual target area must be maintained. , , The location and size are consistent with the initial data collection.

[0079] S203: Based on the feature vector of the dynamic sequence image of the face Extracting key facial feature points (based on The algorithm and the 68-point facial feature point detection model select 10 easily observable tremors, such as the corners of the eyes, mouth, and nose, and denoted as... (The feature point selection logic is consistent with the feature extraction logic of face image processing in S1); calculate the standard deviation of displacement of each feature point during the acquisition period, which is used as the tremor amplitude of the corresponding feature point. The calculation formula is:

[0080]

[0081] in, The amplitude of involuntary facial muscle tremors (unit: pixels); , For the first The feature point at the th ... Frame coordinates; For the first 1 feature point in 300 frames Average coordinates; For the first 1 feature point in 300 frames Average coordinates;

[0082] S204: Based on cognitive status data from 1000 healthy individuals (aged 45-75 years, with no history of cognitive impairment, neurological diseases, or mental illnesses) and 500 early-stage Alzheimer's disease patients (clinically diagnosed, disease duration ≤1 year, age matched to the healthy population), statistical analysis using 95% confidence intervals was conducted, combined with... Curve validation (ensuring sensitivity ≥85% and specificity ≥80%), while also using amyloid protein. The test results were subjected to a consistency check. (Value ≥ 0.75) to define the screening and fitting thresholds for the three types of features, ensuring that the thresholds are highly correlated with the early pathological features of Alzheimer's disease, specifically:

[0083] The threshold for the duration of gaze on the target is: ,and , To adapt to the lower threshold; To adapt to the upper limit of the threshold; below The suggestion is that attention is easily distracted, and the level is higher than [a certain percentage]. The message indicates a slow response, indicating a non-compatible state.

[0084] The threshold for adaptation of eye saccade response delay time is higher than The prompt indicates an abnormal salivation function, indicating a non-fit state; the amplitude of involuntary facial muscle tremors is within the fit threshold. higher than The indication is abnormal facial muscle control, suggesting a mismatch.

[0085] In this embodiment, based on the typical dwell time range of the gaze target in healthy individuals in the early cognitive screening study of Alzheimer's disease, the boundary value of the 95% confidence interval is taken as the fitting threshold. Therefore, the lower limit of the fitting threshold is... The value is 3.5, which is the upper limit of the adaptation threshold. The value is set to 7.5; based on the normal upper limit of eye saccade delay in healthy adults as defined in "Application of Eye-Tracking Technology in Cognitive Assessment," the eye saccade response delay time adaptation threshold is set. The value is set to 0.3 seconds; based on quantitative research on facial dynamic features, the statistical upper limit of the amplitude of facial micro-tremors in healthy individuals is determined, and the threshold value for the amplitude of involuntary facial muscle tremors is set. The value is 2.0 pixels;

[0086] S205: The quantization results of all three types of features meet the corresponding adaptation threshold requirements. , , If the system is determined to be in an adapted state, it will output synchronously. , and The specific numerical values; if the quantification result of any feature does not meet the adaptation threshold requirements, it is judged as an unsuitable state, the current screening process is terminated, and targeted adjustment guidelines are generated;

[0087] The specific adjustment guidelines are as follows:

[0088] The second indicates a loss of attention. The instructions suggest taking a 10-15 minute break to avoid environmental interference (such as turning off electronic devices and keeping the screening environment quiet). After that, repeat step S1 to collect raw multimodal biometric data.

[0089] A second indicates excessive concentration or slow reaction. The instructions suggest resting for 5-8 minutes to relax the eye muscles (look into the distance for 3-5 minutes) to relieve visual fatigue. After that, repeat step S1 to collect raw multimodal biometric data.

[0090] like The second indicates the saccade response delay. The instructions suggest resting for 15 minutes and performing simple eye exercises (such as rotating the eyeballs left and right and blinking for 1 minute each) to avoid fatigue affecting cognitive response. After that, repeat step S1 to collect raw multimodal biometric data.

[0091] like The pixel indicates significant facial tremors. The instructions suggest resting for 20 minutes, adjusting your posture, maintaining emotional stability, avoiding tension or body shaking, and repeating step S1 to collect raw multimodal biometric data once your condition has stabilized.

[0092] If two or more characteristics are not a match, the guidance states that the current state is not suitable for screening requirements, and there are multiple interferences such as fatigue and emotional fluctuations. It is recommended to restart the complete screening process (starting from S1) the next day when the mental state is good and the environment is quiet.

[0093] S206: Adaptation status output judgment result and the duration of gaze on the target. The aforementioned eye saccade response delay time and the amplitude of the involuntary facial muscle tremors The system quantifies three types of features; it outputs judgment results and corresponding adjustment guidelines for non-fit states, clearly informing users of the process nodes that need to be re-executed, thus ensuring the validity of subsequent screening data.

[0094] S3: Based on the cognitive state determination result, the pathological association adaptive weight allocation algorithm is used to quantify the degree of interference of patient state fluctuations on each modality feature and dynamically adjust the weights to generate a weighted target feature set that is adapted to the current state.

[0095] S301: Let the standardized feature set be... It includes three types of modal core features and corresponding basic feature vectors, namely, the feature groups related to dynamic facial sequence images. Iris static image related feature group Speech and audio related feature groups ;

[0096] The facial dynamic sequence image related feature group Includes gaze stability features and facial dynamic sequence image feature vector The iris static image related feature group Includes iris tremor frequency characteristics and iris static image feature vector The speech audio related feature group Includes semantic coherence features and speech power spectrum feature vector ;Right now Furthermore, the three feature groups have the same initial weight in the standardized feature set, all of which are... ;

[0097] S302: Define the influence coefficient of fixation target dwell time. Influence coefficient of eye saccade response delay parameter The influence coefficient of the amplitude of involuntary facial muscle tremors The three types of state influence coefficients and their corresponding value ranges are used to quantify the duration of fixation on the target. Eye saccade reaction delay time and the amplitude of involuntary facial muscle tremors The degree of interference with the weights of the three feature groups;

[0098] S3021: Influence coefficient of the fixation target dwell time Adaptation threshold for the duration of gaze on the target Related, and , This is the lower limit of the coefficient. This represents the upper limit of the coefficient; the calculation formula is:

[0099]

[0100] Indicates when hour, Follow Decrease and increase, and the upper limit does not exceed This indicates that distraction enhances the interference with feature groups in dynamic facial sequence images; when hour, Follow Increases and decreases, and the lower limit is not lower than This indicates that slow reaction time increases the interference to the feature group of dynamic facial sequence image. At this time, a smaller coefficient corresponds to a lower proportion of interference weight.

[0101] In this embodiment, based on the statistical results of visual dwell time interference in individuals with mild cognitive impairment from the study "Visual Behavioral Characteristics of Early Cognitive Impairment in Alzheimer's Disease", the lower limit of the coefficient is... The value is 0.6, which is the upper limit of the coefficient. The value is 1.4;

[0102] S3022: Influence coefficient of the eye saccade response delay parameter Adaptation threshold with the eye saccade response delay time Related, and , This is the lower limit of the coefficient. This represents the upper limit of the coefficient; the calculation formula is:

[0103]

[0104] Indicates when hour, (No interference); when hour, Follow Increases linearly; when hour, Fixed as (Maximum interference);

[0105] In this embodiment, based on the correlation analysis data between delayed eye saccades and cognitive function in "Ocular Motion Abnormalities in Neurodegenerative Diseases", the lower limit of the coefficient is... The value is 0.6, which is the upper limit of the coefficient. The value is 1.4;

[0106] S3023: Regarding the amplitude of the aforementioned involuntary facial muscle tremors Influence coefficient The threshold for matching the amplitude of the involuntary facial muscle tremors. Related, and , This is the lower limit of the coefficient. This represents the upper limit of the coefficient; the calculation formula is:

[0107]

[0108] Indicates when hour, (No interference); when hour, Follow Increases linearly; when hour, Fixed as (Maximum interference);

[0109] In this embodiment, based on the quantitative conclusions of the pathological interference of facial muscle tremors in "Analysis of Facial Micro-expression Features in Early Alzheimer's Disease", the lower limit of the coefficient is... The value is 0.6, which is the upper limit of the coefficient. The value is 1.4;

[0110] S303: Based on the influence coefficient of the fixation target dwell time The influence coefficient of the eye saccade response delay parameter The influence coefficient of the amplitude of the involuntary facial muscle tremors Three types of state influence coefficients are introduced, along with a preset enhanced interference cancellation coefficient. Weakened interference cancellation coefficient Compared with the default weight benchmark value The priority is adjusted by quantifying the weights of each feature group using a formula;

[0111] The gaze stability characteristics of the facial dynamic sequence image feature group are strongly correlated with cognitive state, requiring enhanced anti-interference weights and corresponding adjustment factors for the facial dynamic sequence image feature group. The calculation formula is: , The value is determined by the correlation between facial features and cognitive state; the iris tremor features in the static iris image feature group are easily affected by temporary states such as fatigue and emotions, so their weight should be weakened, corresponding to the adjustment factor of the static iris image feature group. The calculation formula is: , The value of is determined by the degree to which the feature is affected by temporary states; the semantic coherence features of the speech audio feature group are highly correlated with early language function decline and are less affected by temporary states, therefore the baseline weight is maintained, and the corresponding speech feature group adjustment factor is adjusted accordingly. Set directly ;

[0112] S304: To ensure that the sum of the weights of the three feature groups is 1, and to meet the normalization requirements of subsequent feature weighting calculations, the weights of each feature group are adjusted based on the calculated weight adjustment factors. , and The final weights of the three feature groups are calculated using a normalization formula.

[0113] Final weights of facial dynamic sequence image feature groups The calculation formula is: Final weights of iris static image feature groups The calculation formula is: Final weights of speech audio feature groups The calculation formula is: To ensure ;

[0114] when The assessment results show that the subject to be screened is in a stable state (i.e. , , When ), the state influence coefficient At this time, the weight adjustment factor , , Substituting into the normalization formula automatically satisfies the condition. The preset ratio matches the pathological characteristics of language function decline in the early stages of Alzheimer's disease;

[0115] S305: Standardize the feature set The three types of feature groups are multiplied by their corresponding final weights to achieve feature weighting, and then concatenated in the original dimensional order to generate the weighted target feature set. ,and ;in, Weighted facial dynamic feature groups; Weighted static iris feature set; For weighted speech audio feature groups;

[0116] The final generated weighted target feature set It still has a 771-dimensional feature vector, and the weight ratio of each dimension of the feature has been dynamically optimized according to the cognitive state.

[0117] S4: Input the weighted target feature set into the pre-trained multimodal fusion classification model, analyze the spatial pathological features through the G06V10 class image feature extraction algorithm, combine the language pathological features to complete cross-modal fusion, and output a screening report.

[0118] S401: For the weighted target feature set For different modal features, a dedicated encoder is used to unify dimensions and enhance semantics;

[0119] For the weighted facial dynamic feature group Through a 3-layer convolutional neural network Extract spatial temporal features and output a 128-dimensional face encoding vector. The kernel sizes are as follows: The step size is 1, and the activation function is adopted. For the weighted static iris feature group Through a Layer 2 fully connected network Mapped to a 64-dimensional iris coding vector The number of neurons in the hidden layer is 256 and 128, respectively, and the activation function is... For the weighted speech audio feature group Through a single-layer long short-term memory network Capture semantic temporal features and output a 128-dimensional speech coding vector. , The number of hidden layer units is 256. The ratio is set to 0.3 to prevent overfitting;

[0120] S402: Adopted Image feature extraction algorithms respectively process the face encoding vector and the iris encoding vector Spatial pathological features were analyzed to quantify structural abnormality indicators associated with early Alzheimer's disease;

[0121] S4021: For the face encoding vector Calculate the gaze trajectory dispersion Coordination with facial movements Two types of indicators, specifically:

[0122] The gaze trajectory dispersion The formula used to measure the degree to which the pupil's movement trajectory deviates from the normal distribution is: ;in, This represents the total number of fixations. , For the first The coordinates of the gaze point; , This is the average of the coordinates of all gaze points; The higher the value, the more discrete the gaze trajectory, indicating a higher risk of visuospatial function impairment;

[0123] Facial movement coordination Based on the motion correlation of 10 key facial feature points, the calculation formula is as follows: ;in, The number of key facial feature points (in this embodiment) ); , The first , The motion sequence of feature points (composed of coordinates from 300 frames); for and The Pearson correlation coefficient;

[0124] In this embodiment, based on the sampling parameter requirements for "Dynamic Visual Function Detection" in the 2024 version of the "Application Specifications of Eye-tracking Technology in Cognitive Assessment", the total number of fixation points is... The value is 300, corresponding to a 10-second acquisition duration and a frame rate of 30 frames per second; based on the statistical standard definition, referring to the "Linear Correlation Measurement" chapter in "Applied Statistics (5th Edition)," the Pearson correlation coefficient is... The value is [-1] 1]; Based on the "Research on the Correlation between Facial Motor Features and Neuronal Function in the Brain", the facial motor coordination The range of values ​​is [0, 1, 2, 3, 4, 5, 6, 7, 8 [1] The smaller the value, the worse the coordination of facial muscle movements, suggesting a higher risk of abnormal neuronal function in the brain;

[0125] S4022: For the iris coding vector Calculate the irregularity of iris texture With the coefficient of variation of tremor frequency Two types of indicators;

[0126] The irregularity of the iris texture The iris texture grayscale values ​​are quantified by the ratio of the standard deviation to the mean, and the calculation formula is as follows: ;in, A still image of the iris after illumination correction. The grayscale matrix; for Standard deviation; for The mean; The higher the value, the more irregular the iris texture, which is associated with early neurodegenerative changes;

[0127] The coefficient of variation of the tremor frequency The formula used to measure the degree of fluctuation in the frequency of iris tremor is: ;in, This is a set of tremor frequencies from 300 sampling points at the edge of the iris; for Standard deviation; for The mean; The higher the value, the more unstable the iris tremor, indicating a higher risk of abnormal brainstem neural control function;

[0128] S403: For the speech coding vector Extract semantic break frequency With vocabulary richness Two types of core indicators;

[0129] The semantic break frequency The formula used to count the proportion of semantically incoherent sentences in speech descriptions is as follows: ;in, Total number of statements (consistent with the definition in S1053); The number of semantically coherent sentences (through pre-training) The model determines that the similarity is ≥0.7. The higher the value, the more chaotic the language logic, and the higher the risk of language center function decline;

[0130] The vocabulary richness By vocabulary type - token ratio ( Quantification, the calculation formula is: ;in, The number of distinct words in the speech description (after deduplication); The total number of words in the speech description (including repetitions); The value range is The smaller the value, the higher the degree of vocabulary deficiency, which is associated with a decline in early memory retrieval function;

[0131] S404: Spatial pathological feature vector Language pathology feature vector Merged into a 256-dimensional global feature vector Specifically:

[0132] First, through a fully connected layer... (4-dimensional) mapping to 128-dimensional, The (2D) mapping is converted to 128D to obtain the mapped spatial pathological feature vector. With the mapped language pathology feature vector Subsequently, attention weights were calculated based on the correlation between the two types of features and the early pathology of Alzheimer's disease. The calculation formula is as follows:

[0133]

[0134]

[0135] in, Attention weights for spatial pathological features For attention weights of language pathology features, and satisfying the following conditions: ; The weight matrix (dimension 128×128) is used for spatial pathological feature mapping. The weight matrix (128×128) is used for mapping language pathology features; The bias term (dimension 1×128) is used when mapping spatial pathological features; The bias term (dimension 1×128) is used when mapping language pathology features; It is an exponential function;

[0136] Through formula The global feature vector is obtained by weighted summation. ;

[0137] S405: Transfer global feature vectors Input the classification layer of the pre-trained multimodal fusion classification model, and then... The function outputs the risk probability and, combined with a preset threshold, determines the risk level, while also annotating abnormal pathological features:

[0138] based on (Alzheimer's Disease Neuroimaging Initiative) 2018-2023 follow-up dataset, and the 2024 edition of the "Expert Consensus on Early Screening of Alzheimer's Disease in China," the risk levels mentioned include low risk output at the classification level. Medium risk High risk Three types of probabilities, satisfying Specifically: If The risk level was determined to be low, and there were no obvious early pathological features of Alzheimer's disease. Annual follow-up examination is recommended. or The risk level was determined to be medium, with indications of mild pathological abnormalities. A follow-up examination every 6 months is recommended, along with consideration of clinical cognitive scales (such as...). Further evaluation is needed; if The test result was deemed high-risk, indicating significant pathological abnormalities. It was recommended that the individual immediately go to the hospital for amyloid protein testing. Clinical diagnostic examinations such as cerebrospinal fluid testing;

[0139] The reference range for healthy individuals is based on data from 1000 individuals aged 50-70 years without cognitive impairment (95% confidence interval) combined with the "Baseline Study of Cognitive Biomarkers in Normal Populations" 2025. , , , , , If any indicator exceeds the above range, the screening report should clearly indicate the name of the indicator, the measured value, and the direction of deviation (e.g., "gaze trajectory dispersion"). Pixels above the healthy range [5.2, 8.7] indicate a risk of visuospatial function impairment.

[0140] The final screening report is output in a structured format, including six core modules: basic information, review of cognitive status assessment, weight allocation results, pathological feature analysis, risk level conclusion, and data traceability, ensuring that the information is clear and traceable.

[0141] The basic information includes the ID of the object to be screened, the collection time, and the device model (camera resolution, microphone sampling rate); the cognitive state judgment review includes the results of the adaptation / incompatibility judgment in S2 and the corresponding output data; the weight allocation results include the final weights of the three feature groups; the pathological feature analysis includes the measured values, health reference ranges, and abnormal annotations of six pathological indicators; the risk level conclusion includes the low / medium / high risk judgment results, corresponding probabilities, and clinical recommendations; the data traceability includes the original multimodal data storage path and algorithm parameter configuration, which facilitates subsequent data verification and model optimization.

[0142] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0143] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An early screening method for Alzheimer's disease that integrates multimodal biometrics, characterized in that, include: S1: Obtain the raw multimodal biometric data of the target object and process the data to generate a standardized feature set; S2: Based on the standardized feature set, extract three types of cognitive state features: fixation duration, eye saccade response delay time, and unconscious facial muscle tremor amplitude. Compare these features with preset screening and adaptation thresholds, and output the cognitive state judgment result and corresponding adjustment guidance. S3: Based on the cognitive state determination result, the pathological association adaptive weight allocation algorithm is used to quantify the degree of interference of patient state fluctuations on each modality feature and dynamically adjust the weights to generate a weighted target feature set that is adapted to the current state. S4: Input the weighted target feature set into the pre-trained multimodal fusion classification model, and then... The image-based feature extraction algorithm parses spatial pathological features, combines them with linguistic pathological features to complete cross-modal fusion, and outputs a screening report.

2. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 1, characterized in that, The multimodal biometric raw data in S1 includes dynamic sequence images of the face, static images of the iris, and speech description data of the target topic; To address motion blur interference in the aforementioned dynamic facial image sequence, an adaptive Wiener filtering algorithm is employed. To address the uneven illumination noise in the static iris image, the following approach is adopted. The correction algorithm is used to suppress environmental noise interference in the target topic speech description data. First, a short-time Fourier transform is performed on the target topic speech description data to obtain the corresponding speech power spectrum. Then, spectral subtraction is used for noise reduction. Subsequently, from the deblurred dynamic sequence image of the face Extract gaze stability features ; The static image of the iris after illumination correction Extracting iris tremor frequency features From the denoised speech power spectrum Extract semantic coherence features from the corresponding speech content ; The processed original multimodal biometric data Convert to 256-dimensional feature vector And associated with the three types of core pathological features extracted accordingly. Structured splicing is performed to generate a 771-dimensional standardized feature set. ,and ;in, Represents the feature vector of a dynamic sequence of human faces; Represents the feature vector of a static iris image; This represents the speech power spectrum feature vector.

3. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 1, characterized in that, The specific method for extracting and calculating cognitive state features in S2 is as follows: based on the feature vector of the dynamic sequence image of the face. Analyze pupil position trajectory data, using The algorithm locates the pupil center coordinates in each frame and calculates the cumulative time when the pupil center falls within three 100×100 pixel preset visual target areas as the fixation duration. Pupil movement velocity is calculated using first-order difference. The starting and arriving frames of saccades are identified with a velocity threshold of 500 pixels / second. The average delay of at least three effective saccades is calculated as the eye saccade response delay time. Ten key facial feature points are extracted, and the mean standard deviation of displacement of each feature point during the acquisition period is calculated as the amplitude of involuntary facial muscle tremors.

4. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 1, characterized in that, The cognitive state determination and adjustment guidance in S2 specifically includes: setting a preset screening adaptation threshold, wherein the adaptation threshold for the duration of fixation on the target is: The threshold for adaptation of eye saccade response delay time is: The threshold for the amplitude of involuntary facial muscle tremors ; When all three features—the duration of fixation on the target, the delay in eye saccade response, and the amplitude of involuntary facial muscle tremors—meet their corresponding threshold requirements, the system is considered to be in a compatible state and a quantitative value for the feature is output. If any one of the features fails to meet the threshold, the system is considered to be in an incompatible state. Adjustment guidelines are generated for different incompatible types, including rest duration, environmental adjustments, and suggestions for eye activities. If two or more features are incompatible, it is recommended to restart the complete screening process the next day.

5. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 1, characterized in that, The core logic of the pathological association adaptive weight allocation algorithm in S3 is: to use the standardized feature set... The features are divided into three categories: dynamic facial sequence image features, static iris image features, and speech / audio features; a gaze duration influence coefficient is also defined. Influence coefficient of eye saccade response delay parameter The influence coefficient of the amplitude of involuntary facial muscle tremors The influence coefficients of the three states and their corresponding value ranges are used to quantify the degree of interference of the three cognitive state features on the weights of the three feature groups. Correlation with the threshold of fixation duration Correlation with eye saccade response delay time adaptation threshold It is associated with the threshold for the amplitude of involuntary facial muscle tremors.

6. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 5, characterized in that, The process of generating the weighted target feature set in S3 is as follows: based on the influence coefficient of the fixation target dwell time. Influence coefficient of eye saccade response delay parameter The influence coefficient of the amplitude of involuntary facial muscle tremors Calculate the weight adjustment factor for each feature group; Through formula The adjustment factor of the feature group of the dynamic sequence image of the face was calculated. ,in The enhanced interference cancellation coefficient; obtained through the formula The adjustment factor of the feature group of the static iris image was calculated. ,in This is a weakened interference cancellation coefficient; Sound feature group adjustment factor Less affected by temporary states, maintaining the baseline weight; Based on the calculated feature group weight adjustment factor , and The final weights of the feature groups in the dynamic sequence of facial images were calculated after normalization. Final weights of iris static image feature groups Final weights of speech audio feature groups The weights of the three types of feature groups; Finally, the standardized feature set is... The three types of feature groups are multiplied by their corresponding final weights to achieve feature weighting, and then concatenated in the original dimensional order to generate the weighted target feature set. ,and ;in, This is a feature set related to dynamic facial sequence images, including gaze stability features. and basic feature vectors ; This is a set of features related to static iris images, including iris tremor frequency features. and basic feature vectors ; This is a group of speech and audio-related features, including semantic coherence features. and basic feature vectors .

7. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 6, characterized in that, Based on the weighted target feature set The feature encoding and spatial pathological feature parsing in S4 specifically involve: processing the weighted target feature set... Weighted facial dynamic feature group The face is encoded into a 128-dimensional face encoding vector through a 3-layer convolutional neural network. For weighted static iris feature groups It is encoded into a 64-dimensional iris coding vector through a 2-layer fully connected network. For weighted speech audio feature groups It is encoded into a 128-dimensional speech coding vector through a single-layer long short-term memory network. ; use Image feature extraction algorithms, from the face encoding vector The gaze trajectory dispersion is calculated. Coordination with facial movements Two types of indicators; from the iris coding vector The irregularity of iris texture is calculated. With the coefficient of variation of tremor frequency Two types of indicators.

8. The method for early screening of Alzheimer's disease by integrating multimodal biometrics according to claim 7, characterized in that, The cross-modal fusion and screening report generation in S4 specifically involves: from the speech encoding vector... Extract semantic break frequency With vocabulary richness Two types of core indicators are used to construct spatial pathological feature vectors. Language pathology feature vector The two types of vectors are mapped to 128 dimensions through a fully connected layer. The spatial and linguistic feature weights are calculated based on the attention weight formula, and the weighted sum is used to obtain a 256-dimensional global feature vector. ; global feature vectors Input the classification layer of the pre-trained multimodal fusion classification model, and then... The function outputs the risk probability and combines it with a preset threshold to determine the risk level, while also marking abnormal pathological features. The screening report includes six core modules: basic information, review of cognitive status determination, weight allocation results, pathological feature analysis, risk level conclusion, and data traceability.