Alzheimer disease early warning, evaluation and health management system and method based on artificial intelligence

Through multimodal data fusion and intelligent diagnostic strategies, an adaptive Alzheimer's disease early warning system is built, which solves the problem of insufficient multimodal feature processing and user interaction interface in the existing system, and achieves high accuracy and personalized Alzheimer's diagnosis, supporting clinicians' intuitive decision-making and model updates.

CN120388753APending Publication Date: 2025-07-29NANCHANG UNIV

Patent Information

Application Number
CN202510873511.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing Alzheimer's diagnosis system lacks multimodal feature processing methods, the single-dimensional classification model is fixed, it is difficult to update dynamically, and cannot meet the diagnostic needs of patients of different ages and stages of the disease. It also lacks convenient user interaction interface and visualization of diagnostic results, which limits practical application scenarios.

Method used

The multimodal data fusion strategy is adopted, and the Chinese voice, EEG signals, facial expression images and eye movement data are integrated through the data acquisition module. The single-dimensional classification model is built using the feature extraction and modeling module, and preliminary and final diagnosis is carried out through the intelligent diagnosis module. Combined with the Transformer cross-modal self-attention mechanism, the in-depth correlation analysis of multimodal data is realized. At the same time, the user interaction module is provided to support customized diagnostic thresholds and model updates.

Benefits of technology

It improves the accuracy and comprehensiveness of Alzheimer's diagnosis, adapts to the personalized diagnosis needs of patients of different ages and stages, enhances the clinical adaptability and operational convenience of the system, and supports the visualization of diagnostic results and dynamic update of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388753A_ABST
    Figure CN120388753A_ABST
Patent Text Reader

Abstract

The invention discloses a senile dementia early warning, evaluation and health management system and method based on artificial intelligence, and relates to the field of artificial intelligence and medical health. The system comprises a data acquisition module for acquiring Chinese speech, electroencephalogram signals, facial expression images and eye movement data of a patient; the data preprocessing module is used for preprocessing and storing various data; the feature extraction and modeling module is used for extracting a feature set and constructing a one-dimensional classification model; the intelligent diagnosis module inputs the feature set to a one-dimensional classification model to obtain a disease probability and a preliminary diagnosis result, and a final result is obtained through comprehensive diagnosis after a preliminary diagnosis threshold value is compared; and the user interaction module realizes diagnosis visualization, allows a user to define a preliminary diagnosis threshold value, imports data to retrain and updates a one-dimensional classification model. Through multi-modal data fusion and intelligent analysis, potential relations among different modal data are fully mined, and the accuracy of senile dementia diagnosis and the system adaptability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of the cross-integration of artificial intelligence and medical health, and specifically relates to an artificial intelligence-based Alzheimer's disease early warning, assessment and health management system and method. Background Art

[0002] Alzheimer's disease is a complex neurodegenerative disease, and early diagnosis and comprehensive management are crucial for slowing its progression. Currently, traditional diagnostic methods rely primarily on clinical scale assessments, imaging studies, or single biosignal analysis, which struggle to fully reflect the multidimensional pathological characteristics of Alzheimer's disease, leading to high rates of missed or misdiagnosis.

[0003] Existing Alzheimer's disease diagnosis systems lack multimodal feature processing methods specific to Alzheimer's disease. Their fixed, one-dimensional classification models are difficult to dynamically update based on new data, resulting in poor generalization across patients of different ages and disease stages. They also lack effective integration between initial and healthy diagnoses, and their threshold settings are rigid, failing to meet clinical needs for dynamic assessment of diagnostic confidence. Furthermore, most existing systems are laboratory prototypes, lack convenient user interfaces, and cannot visualize diagnostic results. They also lack support for personalized operations such as clinician-defined diagnostic thresholds and importing new data to retrain models, limiting their practical application scenarios. While multimodal fusion methods based on deep learning have shown promise with the advancement of artificial intelligence (AI), specific optimization for Alzheimer's disease remains a technological gap. Therefore, integrating multimodal data, constructing adaptive diagnostic models, and implementing clinically friendly interactive features remain key challenges in the intelligent management of Alzheimer's disease. To address this, we propose an AI-based system and method for early warning, assessment, and health management of Alzheimer's disease. Summary of the invention

[0004] The purpose of the present invention is to provide an artificial intelligence-based Alzheimer's disease early warning, assessment and health management system and method to solve the problems raised in the above background technology.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: An artificial intelligence-based Alzheimer's disease early warning, assessment and health management system, including a data acquisition module, a data preprocessing module, a feature extraction and modeling module, an intelligent diagnosis module and a user interaction module; The data acquisition module is used to collect Chinese speech, EEG signals, facial expression images and eye movement data of patients with Alzheimer's disease; The data preprocessing module is used to preprocess the Chinese speech, EEG signals, facial expression images and eye movement data of the Alzheimer's patients and store them in the database; The feature extraction and modeling module is used to perform feature extraction processing on the Chinese speech, EEG signals, facial expression images, and eye movement data of Alzheimer's patients after preprocessing, to obtain a speech feature set, an EEG feature set, an expression feature set, and an eye movement feature set. At the same time, one-dimensional classification models corresponding to the speech feature set, the EEG feature set, the expression feature set, and the eye movement feature set are respectively constructed; The intelligent diagnosis module is used to input the speech feature set, the EEG feature set, the expression feature set, and the eye movement feature set into their respective one-dimensional classification models to obtain the corresponding disease probabilities. According to the corresponding disease probabilities, a preliminary diagnosis result is obtained. At the same time, a preliminary diagnosis threshold is set, and the preliminary result is compared with the preliminary diagnosis threshold to determine whether to perform a health diagnosis process to obtain a final diagnosis result; The user interaction module is used to achieve diagnostic visualization according to the preliminary or final diagnosis result. At the same time, it allows customization of the preliminary diagnosis threshold, import of new data to retrain the one-dimensional classification model, and update of the one-dimensional classification model.

[0006] Preferably, in the data acquisition module, the process of collecting the Chinese speech, EEG signals, facial expression images, and eye movement data of Alzheimer's patients: Collect the Chinese speech of Alzheimer's patients through an AKG C300 microphone and a Focusrite sound card. The Chinese speech includes vowels, numbers from 0 to 9, phrases of daily sentences, and natural conversations; Collect the EEG signals of Alzheimer's patients with eyes open and eyes closed in a sitting and relaxed state through an 8-channel electroencephalogram device; Collect the facial expression images of Alzheimer's patients including angry, disgusted, afraid, happy, sad, and surprised through a 4K camera. The resolution requirement of the facial expression images is greater than or equal to 1024×768; Under the condition that Alzheimer's patients complete a specified visual task, collect the eye movement data of the fixation point segment, visual task duration, fixation time, and saccade speed of Alzheimer's patients through an infrared eye tracker.

[0007] Preferably, in the data preprocessing module, the process of preprocessing the Chinese speech, EEG signals, facial expression images, and eye movement data of Alzheimer's patients respectively: Perform pre-emphasis processing on the Chinese speech of Alzheimer's patients through a first-order FIR filter to enhance the high-frequency components of the Chinese speech and compensate for the high-frequency attenuation caused by vocal cord and lip radiation. For the pre-emphasized Chinese speech, using the short-term stationarity of the speech, it is segmented into 25ms speech frames through a sliding window mechanism, and the speech frames are windowed through a Hamming window to reduce spectral leakage at the frame boundary, obtaining the preprocessed Chinese speech of Alzheimer's patients; The Hamming window is: ; Among them, is the window function value of the Hamming window at the nth sampling point, is the total number of sampling points included in the window function, is the index of the discrete time series, representing that the current processing is the nth sampling point; The EEG signals of Alzheimer's patients are processed by a 50Hz notch filter to remove the interference of 50Hz power frequency on brain waves. The baseline mean value of the EEG signals after the working filtering process is calculated through the baseline mean formula , and according to the limit mean value, the EEG signals of Alzheimer's patients after preprocessing are obtained through the sample-by-sample correction formula for the EEG signals; The baseline mean formula is: ; Among them, is the baseline mean value, is the value of the original EEG signal at time t, and N is the number of samples; The sample-by-sample correction formula is: ; Among them, is the baseline mean value, is the value of the EEG signal after baseline correction at time t, is the value of the original EEG signal at time t; The facial expression images of Alzheimer's patients are calculated through the RGB image to grayscale image weighted average formula to obtain the grayscale facial expression images. The image size of the grayscale facial expression images is scaled to 224×224 by the bilinear interpolation method and histogram equalization processing is performed to obtain the preprocessed facial expression images of Alzheimer's patients; The RGB image to grayscale image weighted average formula is: ; Among them, is the grayscale value, and R, G, and B are the red, green, and blue component values of each pixel point in the color image respectively; The bilinear interpolation method is an algorithm commonly used in the process of image scaling; The histogram equalization processing is an image processing technology that enhances the image contrast by adjusting the grayscale distribution of the image; Perform fixation point screening on the eye movement data of Alzheimer's patients, traverse each fixation point, and delete the eye movement data with the coordinate values of (-999, -999) returned by the infrared eye tracker. Perform moving average filtering on the eye movement data after fixation point screening to reduce high-frequency noise in the eye movement trajectory, and obtain the preprocessed eye movement data of Alzheimer's patients; The moving average filtering process is a technique widely used in the field of signal processing; The wavelet denoising process is a denoising method applied in the field of signal processing.

[0008] Preferably, in the feature extraction and modeling module, the process of obtaining the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set: Fast Fourier Transform is an algorithm for efficiently calculating the discrete Fourier transform. Apply the Fast Fourier Transform to each frame of the Chinese speech of Alzheimer's patients after preprocessing to obtain the frequency domain amplitude spectrum. Pass the frequency domain amplitude spectrum through the Mel filter bank of the digital signal processing algorithm to obtain the Mel energy. At the same time, take the logarithm of the Mel energy and perform discrete cosine transform to obtain MFCC, that is, Mel Frequency Cepstral Coefficients. Calculate the MFCC through the MFCC first-order difference calculation formula to obtain the first-order difference MFCC, and thus obtain the speech feature set including MFCC and the first-order difference MFCC; The MFCC first-order difference calculation formula is: ; where, is the first-order difference eigenvalue of the MFCC of the t-th frame, t is the frame index, is the MFCC eigenvalue of the t + n-th frame, is the MFCC eigenvalue of the t - n-th frame, and n is an auxiliary index; Segment the EEG signals of Alzheimer's patients after preprocessing by the Welch method to obtain EEG signal segments. Calculate the periodogram through the periodogram formula for the EEG signal segments, take the average of all obtained periodograms to obtain the power spectral density PSD. At the same time, construct an m-dimensional vector and an m + 1-dimensional vector , and use the maximum absolute difference formula to calculate the vector distance between the m-dimensional vector or the m + 1-dimensional vector. Statistically calculate the proportion of vector distances between the m-dimensional vector or the m + 1-dimensional vector that are less than or equal to the distance threshold r to obtain the proportion of vector pairs between the m-dimensional vectors that meet the distance threshold condition and the proportion of vector pairs between the m + 1-dimensional vectors that meet the distance threshold condition , for and By calculating using the sample entropy formula, the sample entropy SE is obtained, and based on this, an EEG feature set including power spectral density and sample entropy is constructed; The periodogram formula is as follows: ; Among them, is the periodogram value at frequency f, N is the length of the signal data segment, is the discrete-time signal, is the complex exponential function; The maximum absolute difference formula is as follows: ; Among them, is the vector spacing between m-dimensional vectors, m is the dimension of the vector, is the vector the k-th element of and the vector the k-th element of the absolute value of the difference; The sample entropy formula is as follows: ; Among them, is the sample entropy, is the proportion of vector pairs that satisfy the distance threshold condition between m-dimensional vectors, is the proportion of vector pairs that satisfy the distance threshold condition between m+1-dimensional vectors; The Welch method is a classical method for estimating the power spectral density PSD of a signal; The facial expression images of Alzheimer's patients after preprocessing are input into the improved ResNet50 to obtain expression feature vectors, which are used to reflect the subtle differences in facial muscle movements, and based on this, an expression feature set is constructed; The Mel filter bank is a digital signal processing algorithm based on the auditory characteristics of the human ear; The improved ResNet50 is the ResNet50 network with the SE attention mechanism introduced; For the fixation time and saccade speed in the eye movement data of Alzheimer's patients after preprocessing, the average fixation time and average saccade speed are calculated respectively through the average calculation formula. At the same time, for the fixation point segments and visual task durations in the eye movement data of Alzheimer's patients after preprocessing, the number of fixation point segments during the visual task duration is counted, and the fixation point density is obtained through the fixation point density formula. Based on this, an eye movement feature set including fixation point density, average fixation time, and average saccade speed is constructed; The average calculation formula is as follows: ; Among them, is the average fixation time or average saccade speed, is the fixation time or saccade speed, and n is the number of data; The formula for the fixation point density is: ; where n is the number of fixation point segments within the statistical visual task duration, is the visual task duration.

[0009] Preferably, in the feature extraction and modeling module, the process of constructing respective one-dimensional classification models for the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set: The speech feature set, EEG feature set, facial expression feature set, and eye movement feature set are respectively divided according to a 7:3 ratio to obtain their respective corresponding training sets and test sets, and the true disease probability is manually labeled for their respective corresponding training sets and test sets , the respective corresponding training sets are input into their respective corresponding one-dimensional classification models for training, and the respective corresponding test sets are input into the trained respective corresponding one-dimensional classification models to obtain the predicted disease probabilities of their respective corresponding test sets , the mean square error MSE is calculated for the predicted disease probabilities of their respective corresponding test sets and the true disease probabilities of their respective corresponding test sets through the mean square error formula, and the mean square error MSE is compared with the mean square error threshold with a set value of 0.05; If the mean square error MSE is less than 0.05, the trained respective corresponding one-dimensional classification models are obtained; If the mean square error MSE is greater than or equal to 0.05, the true disease probabilities of their respective corresponding training sets and test sets are manually labeled again, and the respective corresponding one-dimensional classification models are trained again; The mean square error formula is: ; where N is the number of training sets, is the predicted disease probability of the i-th training set, is the true disease probability of the i-th training set, is the mean square error of the trained respective corresponding one-dimensional classification models; The respective corresponding one-dimensional classification models are as follows: The one-dimensional classification model corresponding to the speech feature set is ResNet18, which is a deep convolutional neural network architecture; the one-dimensional classification model corresponding to the EEG feature set is EEGNet, which is a lightweight convolutional neural network architecture specially designed for processing electroencephalogram signals, the one-dimensional classification model corresponding to the expression feature set is the improved ResNet50, and the one-dimensional classification model corresponding to the eye movement feature set is a three-dimensional convolutional network, which is developed on the basis of a two-dimensional convolutional neural network and is used to process data with a three-dimensional structure.

[0010] Preferably, in the intelligent diagnosis module, the process of obtaining the preliminary diagnosis result according to the respective corresponding disease probabilities: Pass the respective corresponding disease probabilities through the disease probability fusion formula to obtain the preliminary diagnosis result; The disease probability fusion formula is: ; Wherein, is the preliminary diagnosis result and the value is in [0, 1], refers to the weights of the respective corresponding one-dimensional classification models and the weights are all 0.2, and i is the respective corresponding one-dimensional classification model; The respective corresponding disease probabilities are: the disease probability corresponding to the speech feature set, the disease probability corresponding to the EEG feature set, the disease probability corresponding to the expression feature set, and the disease probability corresponding to the eye movement feature set.

[0011] Preferably, in the intelligent diagnosis module, the process of setting the preliminary diagnosis threshold and comparing the preliminary result with the preliminary diagnosis threshold: If the preliminary diagnosis result is greater than or equal to the preliminary diagnosis threshold with a value of 0.9, directly judge that the patient has late-stage Alzheimer's disease; If the preliminary diagnosis result is less than the preliminary diagnosis threshold with a value of 0.9, perform comprehensive diagnosis processing.

[0012] Preferably, in the intelligent diagnosis module, the process of obtaining the final diagnosis result: Perform standardization processing on the speech feature set, EEG feature set, expression feature set, and eye movement feature set respectively, and form an Alzheimer's disease feature data set with the standardized speech feature set, EEG feature set, expression feature set, and eye movement feature set; Input the Alzheimer's disease feature data set into the Transformer encoder to obtain a cross-modal self-attention feature vector, and process the cross-modal self-attention feature vector through a fully connected layer to obtain a diagnostic feature vector , calculate the diagnostic feature vector through the Softmax function to obtain the final diagnostic result including the probability of being healthy, the probability of early Alzheimer's disease, the probability of middle-stage Alzheimer's disease, and the probability of late-stage Alzheimer's disease; The Softmax function is: ; Among them, represents the probability that the sample belongs to class c, where c is healthy, early Alzheimer's disease, middle-stage Alzheimer's disease, or late-stage Alzheimer's disease, is the element corresponding to class c in the diagnostic feature vector taking the natural exponential operation, is or ; The Transformer encoder is the core component of the Transformer neural network; The fully connected layer is the most basic layer in the neural network.

[0013] Preferably, in the user interaction module, according to the preliminary or final diagnostic result, realize diagnostic visualization. At the same time, the process of allowing customization of the preliminary diagnostic threshold, importing data to retrain the one-dimensional classification model, and updating the one-dimensional classification model is as follows: Display the preliminary or final diagnostic result in the form of a probability distribution graph through the Tableau visualization platform. At the same time, through the data import interface, import the customized preliminary diagnostic threshold, new data, or a new one-dimensional classification model to realize changing the preliminary diagnostic threshold, retraining the one-dimensional classification model, or replacing the existing one-dimensional classification model; The new data is the Chinese speech, electroencephalogram signal, facial expression image, or eye movement data of new Alzheimer's disease patients; The Tableau visualization platform is a business intelligence and data visualization platform.

[0014] Due to the adoption of the above technical solutions, the technical progress achieved by the present invention compared with the prior art is: 1. The present invention realizes multi-modal data fusion and intelligent diagnosis strategies, improves the diagnostic efficiency. Through the data acquisition module, it integrates multi-modal information of Chinese speech, electroencephalogram signals, facial expressions, and eye movement data. After preprocessing by the preprocessing module, it uses the feature extraction and modeling module to construct a dedicated feature set and a one-dimensional classification model. The intelligent diagnosis module realizes in-depth correlation analysis of multi-modal data through preliminary probability fusion and the Transformer cross-modal self-attention mechanism, breaks through the one-sidedness of traditional single-modal diagnosis, significantly improves the accuracy and comprehensiveness of Alzheimer's disease diagnosis, and especially has a better ability to capture early subtle pathological features than the prior art.

[0015] 2. The present invention realizes the functions of adaptive model construction and clinical interaction, enhances the practicability of the system. The feature extraction and modeling module supports the dynamic training of single-dimensional classification models, and the user interaction module allows for customizing the diagnostic threshold, importing new data for retraining, and updating the model, solving the problems of fixed models and weak generalization ability in traditional systems, and can adapt to the personalized diagnostic needs of patients of different ages and disease stages. At the same time, the probability distribution visualization of the diagnostic results is realized through the Tableau platform, providing intuitive decision-making support for clinicians. Combined with the configurable diagnostic process, the clinical adaptability and operation convenience of the system are significantly improved, promoting the implementation of intelligent diagnostic technology in actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of the system function modules of the present invention; Figure 2 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0019] Embodiment 1, as Figure 1 described, an early warning, evaluation, and health management system for Alzheimer's disease based on artificial intelligence, including a data acquisition module, a data preprocessing module, a feature extraction and modeling module, an intelligent diagnosis module, and a user interaction module.

[0020] The data acquisition module is used to collect Chinese speech, electroencephalogram signals, facial expression images, and eye movement data of Alzheimer's patients; The data preprocessing module is used to preprocess the Chinese speech, electroencephalogram signals, facial expression images, and eye movement data of Alzheimer's patients respectively and store them in the database; The feature extraction and modeling module is used to perform feature extraction processing on the Chinese speech, electroencephalogram (EEG) signals, facial expression images, and eye movement data of Alzheimer's patients after preprocessing, to obtain a speech feature set, an EEG feature set, an expression feature set, and an eye movement feature set. At the same time, one-dimensional classification models corresponding to the speech feature set, the EEG feature set, the expression feature set, and the eye movement feature set are respectively constructed; The intelligent diagnosis module is used to input the speech feature set, the EEG feature set, the expression feature set, and the eye movement feature set into their respective corresponding one-dimensional classification models to obtain the corresponding disease probabilities. Based on the corresponding disease probabilities, a preliminary diagnosis result is obtained. At the same time, a preliminary diagnosis threshold is set, and the preliminary result is compared with the preliminary diagnosis threshold to determine whether to perform comprehensive diagnosis processing to obtain a final diagnosis result; The user interaction module is used to achieve diagnostic visualization based on the preliminary or final diagnosis result. At the same time, it allows customizing the preliminary diagnosis threshold, importing new data to retrain the one-dimensional classification model, and updating the one-dimensional classification model.

[0021] Further, the working principle of the present invention is illustrated by Example 1 below: Suppose in the geriatrics department of a large tertiary hospital, 100 suspected Alzheimer's patients are selected as the research objects, with an age range between 60 and 80 years old. These patients have all shown varying degrees of abnormal symptoms in cognition, behavior, or language, but have not been clearly diagnosed. At the same time, 50 healthy elderly people of matching age are selected as the control group.

[0022] With the help of an AKG C300 microphone and a Focusrite sound card, the Chinese speech of the patients is collected, covering vowels, numbers 0 to 9, daily sentence phrases, and natural conversations. For example, patient Li showed a lot of word repetition and sentence interruption phenomena in natural conversations, and these speech data were completely recorded; an 8-channel electroencephalogram device is used to collect the EEG signals of the patients when their eyes are open and closed in a sitting and relaxed state. During the collection process, ensure that the patients remain quiet and avoid external interference. For example, in the open-eye state of patient Zhang, the power of the α wave in the EEG signal is significantly lower than the normal range; through a 4K camera, facial expression images of the patients when they are angry, disgusted, afraid, happy, sad, and surprised are taken, with a resolution greater than or equal to 1024×768. When taking the pictures, guide the patients to make corresponding expressions and record the facial muscle movement characteristics. For example, when patient Wang expressed a happy expression, the movement amplitude of the facial muscles was small and the expression was slightly stiff; when the patients completed a specified visual task, such as identifying objects of different shapes, an infrared eye tracker was used to track the eye movement data of the fixation point segment, the visual task duration, the fixation time, and the saccade speed.

[0023] The Chinese speech is pre-emphasized using a first-order FIR filter, segmented into 25-ms speech frames through a sliding window mechanism, and windowed with a Hamming window. After processing, the high-frequency components of the speech signal are enhanced, and spectral leakage is reduced, providing clearer speech features for subsequent analysis. The power frequency interference is removed using a 50-Hz notch filter, the baseline mean is calculated through the baseline mean formula, and then corrected according to the per-sample correction formula. The corrected EEG signal can more accurately reflect the patient's neuroelectrical activity state. The facial expression image is grayscale-converted using the RGB image to grayscale image weighted average formula, scaled to 224×224 using bilinear interpolation, and histogram equalization is performed. The processed image has enhanced contrast, facilitating the extraction of facial expression features. The fixation point data is screened, abnormal data with coordinate values of (-999, -999) is deleted, and then moving average filtering is performed. After processing, the high-frequency noise in the eye movement trajectory is reduced, and the data is smoother and more stable.

[0024] The preprocessed Chinese speech is subjected to fast Fourier transform, Mel filter bank processing, logarithmic transform, and discrete cosine transform to obtain MFCCs. Then, the first-order difference MFCCs are obtained through the MFCC first-order difference calculation formula, and a speech feature set is constructed. For example, in the speech feature set of patient Chen, certain dimensional values of the MFCCs are significantly different from those of the healthy control group, reflecting the abnormality of the speech spectral envelope and dynamic changes. The Welch method is used to segment the EEG signal, the periodogram is calculated and averaged to obtain the power spectral density, a vector is constructed and the sample entropy is calculated, and an EEG feature set is constructed. Among them, in the EEG feature set of patient Liu, the energy distribution of the power spectral density in certain frequency bands is abnormal, and the sample entropy value is also lower than the normal range, indicating a reduction in the complexity of the EEG signal. The preprocessed facial expression image is input into the improved ResNet50 network to obtain an expression feature vector, and an expression feature set is constructed. It can be found from the expression feature set that the facial expression features of patient Huang deviate significantly from those of healthy people in the feature vectors reflecting the subtle differences in facial muscle movement. The average fixation time, average saccade speed, and fixation point density of the eye movement data are calculated, and an eye movement feature set is constructed. In the eye movement feature set, patient He has a shorter average fixation time and a lower fixation point density, indicating a defect in visual attention.

[0025] Each feature set was divided into a training set and a test set in a ratio of 7:3. The true disease probability was manually labeled, and the training set was input into the corresponding unidimensional classification model for training. The model performance was evaluated using the test set. The mean square error (MSE) was calculated using the mean square error formula and compared with the set mean square error threshold of 0.05. If the MSE was less than 0.05, the model training was successful. Otherwise, the model was relabeled and trained. After training, the accuracy of each unidimensional classification model on the test set was: 88% for the speech model, 92% for the EEG model, 75% for the expression model, and 78% for the eye movement model.

[0026] The disease probability output by each single-dimensional classification model is calculated using the disease probability fusion formula to obtain a preliminary diagnosis result. For example, the speech model output disease probability of patient Sun is 0.7, the EEG model is 0.8, the expression model is 0.6, and the eye movement model is 0.6, so the preliminary diagnosis result is 0.68; the preliminary diagnosis threshold is set to 0.9. Since Sun's preliminary diagnosis result of 0.68 is less than 0.9, a comprehensive diagnosis is performed, and each feature set is standardized to form an Alzheimer's disease feature data set. The Transformer encoder is input to obtain the cross-modal self-attention feature vector, and then the fully connected layer and Softmax function are used to calculate the final diagnosis result. Sun's final diagnosis result shows that his health probability is 0.1, the probability of early Alzheimer's disease is 0.4, the probability of mid-term Alzheimer's disease is 0.45, and the probability of late Alzheimer's disease is 0.05. The comprehensive judgment is mid-term Alzheimer's disease. Among 100 suspected patients, after preliminary diagnosis and comprehensive diagnosis, 25 were finally diagnosed with early-stage Alzheimer's disease, 35 with mid-stage Alzheimer's disease, 10 with late-stage Alzheimer's disease, and 30 were healthy. The diagnostic results were 90% consistent with the clinical gold standard.

[0027] Through the Tableau visualization platform, the preliminary and final diagnosis results are displayed in the form of probability distribution graphs. Doctors can intuitively see the probability of each patient belonging to different diagnostic categories. For example, the probability distribution graph of patient Zhou shows that he has a high probability of early Alzheimer's disease. Doctors can further analyze and judge based on this. At the same time, doctors can customize the preliminary diagnosis threshold according to clinical experience and actual needs, or import new data to retrain the single-dimensional classification model, and update the single-dimensional classification model to improve the diagnostic performance of the system.

[0028] Example 2, as Figure 2 As shown, an artificial intelligence-based Alzheimer's disease health management method includes the following steps: S001. Collect Chinese speech, EEG signals, facial expression images and eye movement data of patients with Alzheimer's disease; S002. Preprocess the Chinese speech, EEG signals, facial expression images and eye movement data of patients with Alzheimer's disease and store them in the database; S003. Feature extraction processing is respectively performed on the Chinese speech, EEG signals, facial expression images, and eye movement data of the pre-processed Alzheimer's patients to obtain a speech feature set, an EEG feature set, an expression feature set, and an eye movement feature set. At the same time, single-dimensional classification models corresponding to the speech feature set, the EEG feature set, the expression feature set, and the eye movement feature set are respectively constructed; S004. The speech feature set, the EEG feature set, the expression feature set, and the eye movement feature set are input into their respective single-dimensional classification models to obtain the corresponding disease probabilities. Based on the corresponding disease probabilities, a preliminary diagnosis result is obtained. At the same time, a preliminary diagnosis threshold is set, and the preliminary result is compared with the preliminary diagnosis threshold to determine whether to perform a health diagnosis process to obtain a final diagnosis result; S005. According to the preliminary or final diagnosis result, diagnostic visualization is realized. At the same time, it is allowed to customize the preliminary diagnosis threshold, import new data to retrain the single-dimensional classification model, and update the single-dimensional classification model.

[0029] Further, the working principle of the present invention is illustrated below by Example 2: The Chinese speech of Alzheimer's patients containing vowels, numbers, phrases, and natural conversations is collected by means of a microphone and a sound card. The EEG signals of the patients with their eyes open and closed in a sitting and relaxed state are collected by using an 8-channel EEG device. The facial expression images of the patients showing anger, disgust, fear, happiness, sadness, and surprise are captured by a 4K camera. The eye movement data including fixation segments, visual task durations, fixation times, and saccade speeds of the patients when completing a specified visual task are tracked by an infrared eye tracker.

[0030] The Chinese speech is pre-emphasized by a first-order FIR filter, and then segmented into 25-ms speech frames by using a sliding window mechanism, windowed by a Hamming window, and stored in a database; the power frequency interference is removed from the EEG signals by a 50-Hz notch filter, the baseline mean is calculated and stored after per-sample correction; the facial expression images are stored after being subjected to RGB-to-grayscale weighted average calculation, bilinear interpolation scaling to a size of 224×224, and histogram equalization processing; the abnormal data with coordinate values of (-999, -999) is deleted from the eye movement data by fixation point screening, and then stored in the database after sliding average filtering processing.

[0031] The preprocessed Chinese speech is subjected to fast Fourier transform, Mel filter bank processing, logarithmic transformation, and discrete cosine transform to obtain MFCCs. Then, a speech feature set is constructed through first-order difference calculation, and a one-dimensional classification model is constructed using ResNet18. The EEG signals are segmented using the Welch method, the power spectral density is calculated through the periodogram, vectors are constructed, and sample entropy is calculated to construct an EEG feature set. An EEGNet is used to construct a one-dimensional classification model. The facial expression images are input into an improved ResNet50 with an SE attention mechanism to extract expression feature vectors, construct a feature set, and build a corresponding model. The average fixation time and average saccade speed of the eye movement data are calculated, the number of fixation point segments within the visual task duration is counted, and the fixation point density is calculated through a formula to construct an eye movement feature set. A three-dimensional convolutional network is used to construct a one-dimensional classification model.

[0032] Each feature set is input into the corresponding one-dimensional classification model to obtain the probability of disease. The preliminary diagnosis result is calculated through a fusion formula (the weights of each model are all 0.2). The preliminary diagnosis threshold is set to 0.9. If the preliminary diagnosis result ≥ 0.9, it is directly determined as late-stage Alzheimer's disease. If it is < 0.9, the feature set is standardized and then composed into an Alzheimer's disease feature data set. The cross-modal self-attention feature vectors are obtained by inputting it into a Transformer encoder, processed through a fully connected layer, and then calculated through a Softmax function to obtain the final diagnosis result containing the probabilities of healthy, early-stage, mid-stage, and late-stage Alzheimer's disease. The preliminary or final diagnosis result is displayed in the form of a probability distribution graph through the Tableau visualization platform, supporting the customization of the preliminary diagnosis threshold through a data import interface, importing new data to retrain the one-dimensional classification model, or updating the existing model.

[0033] The above is the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of modifications or replacements, which should all be covered by the protection scope of the present invention.

Claims

1. An early warning, evaluation and health management system for Alzheimer's disease based on artificial intelligence, characterized in that, Including: A data acquisition module, which is used to acquire Chinese speech, electroencephalogram (EEG) signals, facial expression images and eye movement data of Alzheimer's patients; A data preprocessing module, which is used to preprocess the Chinese speech, EEG signals, facial expression images and eye movement data of Alzheimer's patients respectively and store them in a database; A feature extraction and modeling module, which is used to perform feature extraction processing on the preprocessed Chinese speech, EEG signals, facial expression images and eye movement data of Alzheimer's patients respectively to obtain a speech feature set, an EEG feature set, an expression feature set and an eye movement feature set. At the same time, one-dimensional classification models corresponding to the speech feature set, the EEG feature set, the expression feature set and the eye movement feature set are constructed respectively; An intelligent diagnosis module, which is used to input the speech feature set, the EEG feature set, the expression feature set and the eye movement feature set into their respective one-dimensional classification models to obtain the corresponding disease probabilities. According to the corresponding disease probabilities, a preliminary diagnosis result is obtained. At the same time, a preliminary diagnosis threshold is set, and the preliminary result is compared with the preliminary diagnosis threshold to determine whether to perform comprehensive diagnosis processing to obtain a final diagnosis result; A user interaction module, which is used to realize diagnostic visualization according to the preliminary or final diagnosis result. At the same time, it allows customization of the preliminary diagnosis threshold, import of new data to retrain the one-dimensional classification model and update the one-dimensional classification model.

2. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 1, characterized in that, In the data acquisition module, the process of acquiring Chinese speech, EEG signals, facial expression images and eye movement data of Alzheimer's patients: Acquiring Chinese speech of Alzheimer's patients including vowels, numbers, phrases and natural conversations through a microphone and a sound card; Acquiring EEG signals of Alzheimer's patients through an 8-channel electroencephalograph; Acquiring facial expression images of Alzheimer's patients including anger, disgust, fear, happiness, sadness and surprise through a 4K camera; Tracking and acquiring eye movement data of Alzheimer's patients including fixation segments, visual task durations, fixation times and saccade speeds through an infrared eye tracker.

3. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 2, wherein In the data preprocessing module, the process of preprocessing the Chinese speech, EEG signals, facial expression images and eye movement data of Alzheimer's patients respectively: Pre-emphasizing the Chinese speech of Alzheimer's patients through a first-order FIR filter, segmenting the pre-emphasized Chinese speech into 25-ms speech frames through a sliding window mechanism, and windowing the speech frames through a Hamming window to obtain the preprocessed Chinese speech of Alzheimer's patients; The electroencephalogram (EEG) signals of Alzheimer's patients are processed by a 50 Hz notch filter for working filtering, and the baseline mean value is calculated for the EEG signals after working filtering by the baseline mean formula. According to the limit mean value, the preprocessed EEG signals of Alzheimer's patients are obtained for the EEG signals by the per-sample correction formula. Calculating the grayscale facial expression images of Alzheimer's patients through the RGB image to grayscale image weighted average formula, scaling the image size of the grayscale facial expression images to 224×224 through bilinear interpolation, and performing histogram equalization processing to obtain the preprocessed facial expression images of Alzheimer's patients; Performing fixation point screening processing on the eye movement data of Alzheimer's patients, deleting the eye movement data with the coordinate values of (-999, -999) returned by the infrared eye tracker for the fixation points, and performing sliding average filtering processing on the eye movement data after fixation point screening processing to obtain the preprocessed eye movement data of Alzheimer's patients.

4. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 3, characterized in that, In the feature extraction and modeling module, the process of obtaining the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set: Perform fast Fourier transform on the preprocessed Chinese speech of Alzheimer's patients to obtain the frequency-domain amplitude spectrum. Pass the frequency-domain amplitude spectrum through the Mel filter bank to obtain Mel energy. At the same time, take the logarithm of the Mel energy and perform discrete cosine transform to obtain MFCC. Calculate the first-order difference MFCC by using the first-order difference calculation formula of MFCC, and thus obtain the speech feature set including MFCC and the first-order difference MFCC; The preprocessed EEG signals of Alzheimer's patients are segmented by the Welch method to obtain EEG signal segments. The EEG signal segments are calculated by the periodogram formula to obtain periodograms. The average value of all the obtained periodograms is taken to obtain the power spectral density. At the same time, based on the preprocessed EEG signals of Alzheimer's patients, m-dimensional vectors and m + 1-dimensional vectors are constructed, and the vector spacing between the m-dimensional vectors or m + 1-dimensional vectors is calculated using the maximum absolute difference formula. The proportion of the vector spacing between the m-dimensional vectors or m + 1-dimensional vectors that is less than or equal to the distance threshold r is statistically obtained to get the proportion of vector pairs that satisfy the distance threshold condition between the m-dimensional vectors and the proportion of vector pairs that satisfy the distance threshold condition between the m + 1-dimensional vectors , for and , they are calculated by the sample entropy formula to obtain the sample entropy SE. Based on this, an EEG feature set including the power spectral density and the sample entropy is constructed; Input the preprocessed facial expression images of Alzheimer's patients into the improved ResNet50 to obtain the facial expression feature vectors, and thus construct the facial expression feature set; Calculate the average fixation time and average saccade speed respectively for the fixation time and saccade speed in the eye movement data of the preprocessed Alzheimer's patients through the average calculation formula. At the same time, for the fixation point segments and visual task duration in the eye movement data of the preprocessed Alzheimer's patients, count the number of fixation point segments during the visual task duration, and obtain the fixation point density through the fixation point density formula. Thus, construct the eye movement feature set including the fixation point density, average fixation time, and average saccade speed; The improved ResNet50 is the ResNet50 network introduced with the SE attention mechanism.

5. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 4, wherein, In the feature extraction and modeling module, the process of constructing the one-dimensional classification models corresponding to the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set respectively: Divide the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set according to the ratio of 7:3 respectively to obtain their corresponding training sets and test sets. Manually label the true disease probability for their corresponding training sets and test sets respectively. Input their corresponding training sets into their corresponding one-dimensional classification models for training, and input their corresponding test sets into the trained corresponding one-dimensional classification models to obtain the predicted disease probability of their corresponding test sets. Calculate the mean squared error MSE for the predicted disease probability and the true disease probability of their corresponding test sets through the mean squared error formula, and compare the mean squared error MSE with the mean squared error threshold with a set value of 0.05; If the mean squared error MSE is less than 0.05, then obtain the trained corresponding one-dimensional classification model; If the mean squared error MSE is greater than or equal to 0.05, then re-manually label the true disease probability for their corresponding training sets and test sets respectively, and re-train their corresponding one-dimensional classification models; The corresponding one-dimensional classification models are respectively: the one-dimensional classification model corresponding to the speech feature set is ResNet18, the one-dimensional classification model corresponding to the EEG feature set is EEGNet, the one-dimensional classification model corresponding to the facial expression feature set is the improved ResNet50, and the one-dimensional classification model corresponding to the eye movement feature set is the three-dimensional convolutional network.

6. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 5, characterized in that, In the intelligent diagnosis module, the process of obtaining the preliminary diagnosis result according to the corresponding disease probability: Obtain the preliminary diagnosis result by passing the corresponding disease probability through the disease probability fusion formula; The respective corresponding disease probabilities are: the disease probability corresponding to the speech feature set, the disease probability corresponding to the EEG feature set, the disease probability corresponding to the facial expression feature set, and the disease probability corresponding to the eye movement feature set.

7. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 6, characterized in that, In the intelligent diagnosis module, the process of setting a preliminary diagnosis threshold and comparing the preliminary result with the preliminary diagnosis threshold: If the preliminary diagnosis result is greater than or equal to the preliminary diagnosis threshold with a value of 0.9, directly determine that the patient has late-stage Alzheimer's disease; If the preliminary diagnosis result is less than the preliminary diagnosis threshold with a value of 0.9, perform comprehensive diagnosis processing.

8. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 7, wherein, In the intelligent diagnosis module, the process of obtaining the final diagnosis result: Perform standardization processing on the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set respectively, and form an Alzheimer's disease feature data set with the standardized speech feature set, EEG feature set, facial expression feature set, and eye movement feature set; Input the Alzheimer's disease feature data set into the Transformer encoder to obtain a cross-modal self-attention feature vector, process the cross-modal self-attention feature vector through a fully connected layer to obtain a diagnostic feature vector, and calculate the diagnostic feature vector through the Softmax function to obtain a final diagnosis result including the healthy probability, early-stage Alzheimer's disease probability, middle-stage Alzheimer's disease probability, and late-stage Alzheimer's disease probability.

9. The early warning, assessment and health management system for Alzheimer's disease based on artificial intelligence according to claim 8, characterized in that, In the user interaction module, according to the preliminary or final diagnosis result, realize diagnostic visualization. At the same time, the process of allowing customization of the preliminary diagnosis threshold, importing data to retrain the one-dimensional classification model, and updating the one-dimensional classification model: Display the preliminary or final diagnosis result in the form of a probability distribution graph through the Tableau visualization platform. At the same time, through the data import interface, import the customized preliminary diagnosis threshold, new data, or a new one-dimensional classification model to realize changing the preliminary diagnosis threshold, retraining the one-dimensional classification model, or replacing the existing one-dimensional classification model; The new data is the Chinese speech, EEG signal, facial expression image, or eye movement data of a new Alzheimer's disease patient.

10. A health management method for Alzheimer's disease based on artificial intelligence, characterized in that, The method is used to implement an Alzheimer's disease early warning, assessment, and health management system based on artificial intelligence in claim 1. The method includes the following steps: S1. Collect the Chinese speech, EEG signal, facial expression image, and eye movement data of Alzheimer's disease patients; S2. Perform preprocessing on the Chinese speech, EEG signal, facial expression image, and eye movement data of Alzheimer's disease patients respectively, and store them in the database; S3. Perform feature extraction processing on the preprocessed Chinese speech, EEG signal, facial expression image, and eye movement data of Alzheimer's disease patients respectively to obtain a speech feature set, an EEG feature set, a facial expression feature set, and an eye movement feature set. At the same time, construct one-dimensional classification models corresponding to the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set respectively; S4. Input the speech feature set, EEG feature set, facial expression feature set, and eye movement feature set into their respective one-dimensional classification models to obtain their respective corresponding disease probabilities. Based on their respective corresponding disease probabilities, obtain a preliminary diagnosis result. At the same time, set a preliminary diagnosis threshold, compare the preliminary result with the preliminary diagnosis threshold, and determine whether to perform a health diagnosis process to obtain a final diagnosis result; S5. According to the preliminary or final diagnosis result, achieve diagnostic visualization. At the same time, allow customization of the preliminary diagnosis threshold, import of new data to retrain the one-dimensional classification model, and update of the one-dimensional classification model.

Citation Information

Patent Citations

  • Senile dementia monitoring system based on healthy service robot

    CN105078449A

  • Cognitive competence evaluation system and method based on eye movement and electroencephalogram characteristics

    CN110801237A

  • Medical diagnosis auxiliary system based on electroencephalogram signals and artificial intelligence classification

    CN117612710A

  • Dynamic emotion recognition method and system based on electroencephalogram signals and eye movement signals

    CN118490233A

  • Multi-modal language barrier screening system based on intelligent elderly assistant

    CN118588286A

Cited By

  • Dementia identification method based on electroencephalogram norm network and double-current attention fusion

    CN121606304A

  • A dementia recognition method based on electroencephalogram norm network and double-flow attention fusion

    CN121606304B