Early screening tool for cognitive impairment of diabetes
Through the AI-MCSS multimodal cognitive screening system, combined with facial expressions, voice features and clock drawing tests, blood sugar fluctuations are dynamically monitored, solving the problems of long time consumption and low recognition rate of existing tools, and achieving high-precision diagnosis and early warning of diabetic cognitive impairment.
Patent Information
- Application Number
- CN202510777784.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-05
AI Technical Summary
Existing diagnostic tools for diabetic cognitive impairment are time-consuming, easily affected by the patient's cultural level and the testing environment, have low recognition rates, and existing models fail to effectively integrate facial expression recognition and voice feature analysis, and lack consideration of individual blood sugar fluctuations, resulting in insufficient model accuracy.
Develop a highly sensitive and specific AI-MCSS multimodal cognitive screening system. By dynamically collecting information data, combining nested long-short-term memory models, optimizing model parameters, integrating facial behavior, voice features and clock drawing tests, and using multimodal data fusion and correction training, dynamic monitoring of blood sugar fluctuations can be achieved.
It improves the accuracy and sensitivity of identifying diabetic cognitive impairment, can provide accurate diagnosis and early warning in the case of blood sugar fluctuations, and enhances the adaptability and accuracy of the model.
Smart Images

Figure CN120600316A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an early screening tool for cognitive impairment, in particular to an early screening tool for diabetic cognitive impairment, and belongs to the field of artificial intelligence pathological identification. Background Art
[0002] Currently, cognitive assessments are commonly used to diagnose cognitive impairment in diabetes. However, these assessments are time-consuming, susceptible to patient literacy and testing environment, and have a low recognition rate for early-stage mild cognitive impairment. While our team's previously developed nomogram diagnostic model and Python-based machine learning model can effectively assess a patient's immediate cognitive status, they have the following limitations: they only reflect cognitive function at the time of a single assessment and lack long-term dynamic monitoring and management of a patient's cognitive status. They also fail to integrate important bio-behavioral indicators such as facial expression recognition and speech feature analysis, resulting in a relatively single-dimensional assessment of cognitive function.
[0003] The study found that the current artificial intelligence model is effective in identifying common cognitive impairment, but has low accuracy in identifying diabetes-related cognitive impairment. Therefore, developing new artificial intelligence models is urgent.
[0004] Finally, blood sugar levels fluctuate at different times and periods, before and after medication, and with varying medication habits, which inevitably impacts training effectiveness. Existing models don't account for these individual fluctuations, leading to limited improvements in model accuracy. Therefore, significantly improving model accuracy has become a pressing technical challenge. Summary of the Invention
[0005] 0. Core Explanation of the Invention Therefore, the present invention aims to establish a highly sensitive, highly specific, and intelligent artificial intelligence multimodal cognitive screening system (AI-MCSS) for accurate diagnosis and early warning of diabetes-related cognitive dysfunction. The core of the present invention is to: First, the model is trained by dynamically collecting information data, rather than randomly collecting data for a fixed subject for overall training. This is because patients' blood sugar levels vary at different times and with different medication habits, and the data states corresponding to different blood sugar levels will also show different test results. Second, the improvement of the AI-MCSS model for cognitive dysfunction adopts a nested model and uses the long-short-term memory model to perform correction training on the AI-MCSS model results to correct the AI-MCSS model parameters, so as to improve the model based on the fluctuation factors of the sample itself on the basis of a large sample, further improve the model accuracy, and achieve high sensitivity that can identify actual fluctuations in blood glucose data.
[0006] The specific plan is as follows:
[0007] 1. Overall Architecture The early screening tool for diabetic cognitive impairment includes an assessment test system for testing and evaluating the test subjects using two standardized cognitive assessment scales, MMSE and MoCA. The two standardized cognitive assessment scales, MMSE and MoCA, each consist of multiple completely different questionnaires, which are used to repeatedly test the test subjects at different times.
[0008] Facial behavior quantitative analysis system, used to analyze the facial feature values of the test subject, Multimodal acoustic feature intelligent analysis system, used to intelligently analyze and model the test subject's speech and output speech feature data. The clock drawing test intelligent evaluation system is used to give the test subject a clock drawing test, perform intelligent modeling, and obtain clock drawing data. On-site recording equipment is used to record the facial expressions and voices of the test subjects, and communicate with the facial behavior quantitative analysis system and the multimodal acoustic feature intelligent analysis system to upload the recorded data. The backend server is used to construct an AI-MCSS model based on the received facial feature values, voice feature data, clock drawing data, and acquired clinical information data, and to nest the correction model with the AI-MCSS model, and to further optimize the parameters of the AI-MCSS model in combination with the repeated retest results. The correction model is a pre-trained retest prediction model under different periods of blood glucose levels. Its input end inputs the fully connected features obtained by the fully connected layer of the AI-MCSS model based on the fusion features of the modal features acquired according to the facial feature values, voice feature data, clock drawing data, and acquired clinical information data. The output end is the corresponding retest prediction result, and two sets of prediction results before and after optimization are provided to the patient for reference.
[0009] Optionally, the on-site recording equipment is a high-definition camera or a high-definition recording equipment.
[0010] Optionally, the assessment scale is sent to the testee via the Internet and is filled out on the testee's computer. The assessment test system and the clock drawing test intelligent assessment system are set up in the testee's computer, and the facial behavior quantitative analysis system and the multimodal acoustic feature intelligent analysis system are set up in the medical institution's computer. 2. Data Acquisition
[0011] Specifically, optionally, the method for obtaining facial feature values is: using the OpenFace tool to extract 17 types of facial action units, three-axis head posture and gaze direction features, and counting their time series mean, peak value, jitter amplitude, and gaze stability feature parameters as facial feature values of facial video data.
[0012] Optionally, the method for acquiring speech feature data is: using 16kHz sampling, pre-emphasis and Hamming window processing, and extracting MFCC, fundamental frequency, formant, speech rate, pause rate and eGeMAPS emotional features as key features of the speech data.
[0013] The method for obtaining the clock drawing features is as follows: 1024×768 resolution input is converted into a grayscale image, and then binarized to highlight the content of the clock drawing; opening and closing operations are used to remove noise and enhance image quality.
[0014] The clinical information data were obtained by normalizing continuous variables using Z-score and interpolating missing values.
[0015] 3. Construction of the Basic AI-MCSS Model Optionally, a basic AI-MCSS model is constructed based on facial feature values, voice feature data, clock drawing data, and acquired clinical information data. The specific method includes a step of acquiring the data, and the following subsequent steps: S1 multimodal feature mapping steps, specifically including facial feature value feature mapping based on Bi-LSTM network, voice feature data feature mapping based on 1D-CNN+LSTM series model, clock drawing data feature mapping based on MobileNetV2, and acquired clinical information data encoding and feature mapping through multi-layer perceptron (MLP), four types of modal feature mapping; In the feature fusion step of the S2 modality, the feature map of step S1 is fed into the Transformer-style multi-head attention module to achieve weighted fusion of cross-modal information and model the semantic relevance between the modalities as the weight of weighted fusion. S3 embeds the fused features into a set of fully connected layers + SoftMax to output the final three classifications: normal cognition, MCI, and suspected MCI.
[0016] In clinical practice, some patients may miss one or more modalities due to equipment or operation limitations (such as failure to complete voice recording or clock drawing test).
[0017] To this end, in the model design, preferably, at least one of the following complement mechanisms is introduced: a mask mechanism, a migration complement network, and multi-path output fusion.
[0018] The specific training method of the basic AI-MCSS model is: internal validation is based on 3,000-10,000 retrospective case data, with 70% of the training set used for model parameter learning; 15% of the validation set used for model parameter adjustment and early stopping strategy; and 15% of the test set used for evaluating the model and selecting the best model hyperparameters.
[0019] Preferably, in order to evaluate the reliability of the AI-MCSS model, an evaluation scale is used to verify and analyze its output results. By comparing and analyzing the prediction results of AI-MCSS with the retest results of the corresponding evaluation scale, the accuracy, sensitivity and specificity are evaluated and further optimized, which will be given later.
[0020] That is to say, when using the AI-MCSS model for prediction, the corresponding evaluation scale used for verification analysis is the retest result of the evaluation scale corresponding to the current moment of prediction, rather than the retest result at other moments.
[0021] 4. Model Correction Because blood sugar fluctuations at different times are caused by different medication taking habits (regular, irregular, or partially irregular), different treatment effects, mood, or other unknown factors, it is impossible to accurately determine blood sugar levels using only the above-mentioned comprehensive modal features. Therefore, the above-mentioned basic AI-MCSS model needs to be further improved to achieve the above-mentioned further optimization.
[0022] In view of this, further, by nesting the correction model with the AI-MCSS model and combining the repeated retest results, the specific steps for further optimizing the parameters of the AI-MCSS model include: Step 1: Pre-training the calibration model, including: T1: Recollect 3,000-10,000 retrospective case data, divide these data into groups, and construct multiple correction models with the same number of groups as the grouped data. The recollected 3,000-10,000 retrospective case data are completely different from the 3,000-10,000 retrospective case data used to train the basic AI-MCSS model, and are therefore recollected data. T2 sends the data of each group together into a set of fully connected layers according to steps S1-S2 to form corresponding multiple fully connected features, which are divided into a correction training set and a correction verification set in a ratio of 14:3. The corresponding correction model is pre-trained and verified for each group, wherein the multiple fully connected features are composed of multiple groups of sub-fully connected features, each group of sub-fully connected features corresponds to a time node in the time series, and each group of sub-fully connected features contains multiple fully connected features, and they are all divided into the correction training set and the correction verification set; Step 2: Improve the basic AI-MCSS model: T3 removes the Softmax layer of the trained basic AI-MCSS model, connects the fully connected output end to the input end of each pre-trained correction model to form a model nesting, and inputs the training set of the trained basic AI-MCSS model into the basic AI-MCSS model again. The fully connected features obtained through S1-S3 are input to the input end of the corresponding correction model. The classification function and loss function are constructed at the output end. The classification result obtained by the classification function is compared with the re-test result of the corresponding evaluation scale to obtain the accuracy. The loss function is then back-propagated from the fully connected layer of the basic AI-MCSS to optimize the network parameters of the basic AI-MCSS model connected at each input end, and the validation set is used for verification until the verification accuracy stabilizes and the loss function is minimized. T4 removes the connection between the fully connected output end and the input end of each correction model, and reconnects each end back to the Softmax layer to obtain the improved AI-MCSS model corresponding to multiple moments.
[0023] Optionally, the correction model is a long short-term memory model, in which each node unit corresponds to each time node in the time series, and the time interval between node units is 1-6 hours.
[0024] Optionally, the method for grouping data according to different blood sugar fluctuation characteristics in time series includes the following steps: Q1 obtains the characteristics of blood glucose changes over time in 3,000-10,000 retrospective case data to form a blood glucose feature vector; Q2 constructs a multidimensional space, divides the feature vectors into vector training and maps them into the multidimensional space, and uses cluster analysis methods to divide different vectors into multiple feature clusters; Q3 clusters each type of feature as a group.
[0025] Optionally, the blood glucose feature vector is in the form of ,in, and At successive moments and The blood sugar peak value, at this time the dimension of the multidimensional space is 4, It is the time corresponding to the first blood glucose peak that occurs in chronological order between 0:00 and 24:00 on a 24-hour basis.
[0026] The blood glucose feature vector is formed by the blood glucose values at two characteristic moments, which is used to characterize the characteristics of blood glucose changes over time in different retrospective case data.
[0027] Preferably, the blood glucose feature vector is in the form of , that is, with more groups of different time and blood glucose values, is half the dimension of the multidimensional space, For the moment blood sugar levels.
[0028] Optionally, the element moment in the blood glucose feature vector The subsequent times are selected from 6:00, 7:00, 8:00, ..., and 24:00, which are measured in a 24-hour system with intervals of one hour.
[0029] It should be understood that there is a significant difference between the above four-dimensional space and the multi-dimensional space scheme, that is, the former does not specify the specific time of the successive moments. As long as the first blood sugar peak occurs between 0:00 and 24:00, it is determined. , and the peak values at subsequent moments The difference between the latter is that it stipulates The subsequent time must be selected from 6 o'clock, 7 o'clock, 8 o'clock, etc. in the 24-hour system, and it is not necessarily the peak time. Because each person has different blood sugar fluctuation characteristics due to various aforementioned factors, as long as the first peak and its subsequent peak It is confirmed that the main fluctuation characteristics have been determined, and the subsequent ones are secondary fluctuation characteristics.
[0030] 5. Beneficial effects 1. Using facial, voice, clock drawing, and clinical multimodal data, combined with the retest results of two standardized cognitive assessment scales, MMSE and MoCA, we obtained the basic AI-MCSS model.
[0031] 2. Using the constructed multidimensional space, we perform mapping cluster analysis on the blood glucose feature vectors that reflect time-dependent blood glucose changes to form data groups. Based on the retest results, we perform nested optimization of the basic AI-MCSS model and the pre-trained correction model LSTM to optimize the basic AI-MCSS model parameters. Finally, we untie the nested model to obtain the optimized basic AI-MCSS model at different time nodes. This model also takes into account the impact of blood glucose changes on the prediction results, making it more accurate and sensitive. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a configuration diagram of a tool for early screening of diabetic cognitive impairment according to an embodiment of the present invention; Figure 2 is a flow chart of a method for constructing a basic AI-MCSS model according to an embodiment of the present invention; Figure 3a to Figure 3f This is an integrated diagram of the form and acquisition method of modeling data according to an embodiment of the present invention; Figure 4 This is a graph showing changes in blood sugar over time according to an embodiment of the present invention; Figure 5 1 is a schematic diagram of clustering of blood glucose feature vector mapping in four-dimensional space according to an embodiment of the present invention; Figure 6 1 is a flowchart of pre-training of correction models I-IV of four groups I to IV according to an embodiment of the present invention; Figure 7 4 is a flow chart of an improved basic AI-MCSS model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] Figure 1 A specific configuration diagram of the early screening tool for diabetic cognitive impairment of an embodiment of the present invention is given, including a test subject's computer equipped with an evaluation test system (software) and a clock drawing test intelligent evaluation system, a medical institution computer equipped with the facial behavior quantitative analysis system and the multimodal acoustic feature intelligent analysis system, a backend server, a high-definition camera communicating with the medical institution computer, and a high-definition recording device (such as a high-definition voice recorder or professional recording device).
[0034] The test subjects answered the MoCA and MMSE assessment scales sent via the Internet on their computers, and uploaded the data to the backend server through the clock drawing test intelligent assessment system; the facial behavior quantitative analysis system was used to analyze the facial feature values of the test subjects, and the multimodal acoustic feature intelligent analysis system was used to intelligently analyze and model the test subjects' voices and output voice feature data. Both uploaded the facial feature values and voice feature data to the backend server.
[0035] The backend server is used to construct an AI-MCSS model based on the received facial feature values, voice feature data, clock drawing data, and acquired clinical information data, and to nest the correction model with the AI-MCSS model, and to further optimize the parameters of the AI-MCSS model in combination with the repeated retesting results.
[0036] Figure 2A specific method for constructing a basic AI-MCSS model based on facial feature values, voice feature data, clock drawing data, and acquired clinical information data is given in an embodiment of the present invention, that is, a basic AI-MCSS model construction method.
[0037] Before the formal construction, data acquisition preparation is required as a preliminary step of the construction.
[0038] Specifically, the method for obtaining facial feature values is: use the OpenFace tool to extract 17 types of facial action units (AUs), three-axis head postures, namely pitch / yaw / roll, and gaze direction features, and then count their time series mean, peak value (mean and peak value corresponding to 17 types of facial action units), jitter amplitude (jitter amplitude of three-axis head posture), and gaze stability feature (gaze direction feature) parameters as the facial feature values of facial video data.
[0039] The method for obtaining speech feature data is: using 16kHz sampling, pre-emphasis and Hamming window processing, and extracting MFCC, fundamental frequency, formant, speaking rate, pause rate and eGeMAPS emotional features as key features of speech data.
[0040] The method for obtaining the clock drawing features is as follows: 1024×768 resolution input is converted into a grayscale image, and then binarized to highlight the content of the clock drawing; opening and closing operations (structural element 3×3) are used to remove noise and enhance image quality.
[0041] The clinical information data were obtained by normalizing continuous variables using Z-score and interpolating missing values.
[0042] Figure 3a to Figure 3f Given Figure 1 and Figure 2 The specific form and acquisition method of modeling data are integrated in the figure. This is to better sort out the acquisition of modeling data and clarify the specific form of preparatory work. Among them, Figure 3a For the MoCA cognitive scale, Figure 3b It represents the MMSE cognitive level. Figure 3c For patient information records, Figure 3d This is a diagram of the facial behavior quantitative analysis system. Figure 3e This is a schematic diagram of the multimodal acoustic feature analysis model. Figure 3f Schematic diagram of the intelligent evaluation system for clock drawing test.
[0043] Formal modeling steps are as follows Figure 2 As shown, S1 multimodal feature mapping steps, specifically including facial feature value feature mapping based on Bi-LSTM network, voice feature data feature mapping based on 1D-CNN+LSTM series model, clock drawing data feature mapping based on MobileNetV2, and acquired clinical information data encoding and feature mapping through multi-layer perceptron (MLP), four types of modal feature mapping; In the feature fusion step of the S2 modality, the feature map of step S1 is fed into the Transformer-style multi-head attention module to achieve weighted fusion of cross-modal information, and the semantic relevance between the modalities is modeled as the weight of the weighted fusion to be input into the multi-head attention module; S3 embeds the fused features into a set of fully connected layers + SoftMax to output the final three classifications: normal cognition, MCI, and suspected MCI.
[0044] The model also introduces at least one completion mechanism among masking mechanism, migration completion network, and multi-path output fusion.
[0045] Still like Figure 2 As shown in the figure, the specific training method of the basic AI-MCSS model is: internal validation is performed based on 5,000 retrospective case data, with 70% of the training set used for model parameter learning; 15% of the validation set used for model parameter adjustment and early stopping strategy; and 15% of the test set used for evaluating the model and selecting the best model hyperparameters.
[0046] To evaluate the reliability of the AI-MCSS model, the MoCA and MMSE assessment scales were used to validate and analyze its output results. The prediction results of AI-MCSS were compared with the retest results of the corresponding assessment scales to evaluate the accuracy, sensitivity, and specificity and further optimize the model.
[0047] Figure 2 From the fully connected layer, the The downward arrow of the symbol character is used to describe the construction of the correction model and the improvement of the basic AI-MCSS model.
[0048] Specifically, by nesting the correction model with the AI-MCSS model and combining the repeated retest results, the specific steps for further optimizing the parameters of the AI-MCSS model include: Step 1: Pre-training the calibration model, including: T1 re-collected 5,000 retrospective case data and grouped them.
[0049] like Figure 4 As shown, the specific grouping methods include: Q1 obtains the characteristics of blood glucose changes over time in 5000 retrospective case data. Figure 4 The changes of blood glucose over time are given in the table, from which the blood glucose peak values bs1 and bs2 corresponding to t1 and t2 are given, and the blood glucose feature vector is formed. ,in, and respectively in succession Figure 4 Middle time and The blood sugar peak value, at this time the dimension of the multidimensional space is 4, It is the time corresponding to the first blood glucose peak that occurs in chronological order between 0:00 and 24:00 on a 24-hour basis.
[0050] Q2 Figure 5 As shown, construct a four-dimensional space {bs1, t1, bs2, t2} (origin is O), and The feature vectors are divided into vector training and mapped into the four-dimensional space. The cluster analysis method is used to divide different vectors into multiple feature clusters. The figure shows the four types of intersection clusters I-IV as an example; Q3 clusters each type of feature as a group, and obtains four groups from group I to group IV. For more groups, the corresponding feature vectors are obtained in the same way. , and all the following steps are carried out similarly to obtain the corresponding multiple correction models.
[0051] Among them, the 5,000 retrospective case data re-collected in Q1 are completely different from the 5,000 retrospective case data on which the basic AI-MCSS model is trained, and are re-collected data.
[0052] T2 Figure 6 As shown, the data of each of the four groups are sent to a set of fully connected layers (such as Figure 2 As shown), forming corresponding multiple fully connected features (with The arrow output, Figure 6 ), which is divided into a correction training set and a correction validation set with a ratio of 14:3 (i.e., 70%:15%), and the corresponding correction model is pre-trained and validated for each group.
[0053] The four groups are divided into correction model I to correction model IV, and the correction models are all LSTM long short-term memory models.
[0054] Among them, multiple fully connected features are composed of multiple groups of sub-fully connected features, each group of sub-fully connected features corresponds to a time node in the time series and also corresponds to a node unit in LSTM. Each group of sub-fully connected features contains multiple fully connected features, and they are divided into the correction training set and the correction verification set, so that they can be respectively as follows Figure 6Different LSTM node units are input for pre-training and verification. The node outputs are sent to the rectified fully-connected layer (named differently to distinguish it from the fully-connected layer of the basic AI-MCSS model) and then to the classification function (which can be a softmax function, a sigmond function, etc.). The final output is normal cognition, MCI, and suspected MCI. The loss function corresponding to each unit node is calculated simultaneously: loss 1, loss 2, ..., loss n (in this example, n = 2). The accuracy is verified using the rectified validation set until the total loss function loss 1 + loss 2 + ... + loss n is minimized. Step 2: Improve the basic AI-MCSS model T3 Figure 7 As shown, the Softmax layer of the trained basic AI-MCSS model is removed (the S The symbol indicates a Softmax layer, and this symbol disappears after it is removed). The fully connected output end is connected to the input end of each pre-trained correction model to form a model nesting. The training set of the training basic AI-MCSS model is input into the basic AI-MCSS model again. The fully connected features obtained through S1-S3 are input into the input end of the corresponding correction model, and the classification function and loss function are constructed at the output end.
[0055] In order to distinguish it from the pre-trained correction model, that is, the pre-trained LSTM model, the loss function corresponding to each unit node here is expressed as loss 1 , loss 2 , . . . , loss , the total loss function loss is 1 +Loss 2 +... + loss .
[0056] The classification results obtained by the classification function are compared with the retest results of the corresponding evaluation scale to obtain the accuracy, and the loss function is used to calculate the accuracy. , loss 2 , . . . , loss Backpropagate from the fully connected layer of the basic AI-MCSS to optimize the network parameters of the basic AI-MCSS model for each input-out connection, and use the validation set for verification until the verification accuracy stabilizes and the loss function is minimized.
[0057] Figure 5 In the paper, we take group I as an example to give the pre-trained correction model I. Figure 6 Shown in. Figure 7 The improved basic AI-MCSS model also takes group I as an example, and the other groups II to IV are constructed in the same way as step T3. Figure 7Each classification result must be compared with the retest result made at the time node corresponding to the corresponding LSTM unit node.
[0058] T4 is still Figure 7 As shown in the figure, we removed the connections between the fully connected output and the input of each correction model and reconnected them to the Softmax layer to obtain the improved AI-MCSS model at multiple time points. Finally, we obtained an AUC greater than 0.92 before and after optimization, as shown in the following table:
[0059] The method of providing patients with two sets of prediction results before and after optimization for reference is specifically to perform the basic AI-MCSS model prediction before optimization for the current MoCA and MMSE assessment scales, and simultaneously input the improved AI-MCSS model corresponding to the corresponding LSTM node unit. If the two prediction results are the same, the diagnosis result is the corresponding prediction result. If the two are different, it indicates that further examination and confirmation are required.
Claims
1. A tool for early screening of cognitive impairment in diabetes, characterized by: It includes an assessment test system for testing and evaluating the test subjects using two standardized cognitive assessment scales, MMSE and MoCA. The two standardized cognitive assessment scales, MMSE and MoCA, each consist of multiple completely different fill-in question papers, which are used to repeatedly test the test subjects at different times. Facial behavior quantitative analysis system, used to analyze the facial feature values of the test subject, Multimodal acoustic feature intelligent analysis system, used to intelligently analyze and model the test subject's speech and output speech feature data. The clock drawing test intelligent evaluation system is used to give the test subject a clock drawing test, perform intelligent modeling, and obtain clock drawing data. On-site recording equipment is used to record the facial expressions and voices of the test subjects, and communicate with the facial behavior quantitative analysis system and the multimodal acoustic feature intelligent analysis system to upload the recorded data. The backend server is used to construct an AI-MCSS model based on the received facial feature values, voice feature data, clock drawing data, and acquired clinical information data, and to nest the correction model with the AI-MCSS model, and to further optimize the parameters of the AI-MCSS model in combination with the results of the repeated retests. The correction model is a pre-trained retest prediction model for blood glucose levels at different periods. Its input end inputs the fully connected features obtained by the fully connected layer of the AI-MCSS model through the fusion features of the modal features acquired based on the facial feature values, voice feature data, clock drawing data, and acquired clinical information data. The output end is the corresponding retest prediction result, and two sets of prediction results before and after optimization are provided to the patient for reference.
2. The tool according to claim 1, characterized in that The on-site recording equipment is a high-definition camera and a high-definition recording device.
3. The tool according to claim 1, characterized in that The assessment scale is sent to the testee via the Internet and is filled out on the testee's computer. The assessment test system and the clock drawing test intelligent assessment system are set up in the testee's computer, and the facial behavior quantitative analysis system and the multimodal acoustic feature intelligent analysis system are set up in the medical institution's computer.
4. The tool according to claim 3, characterized in that The facial feature values are obtained by using the OpenFace tool to extract 17 types of facial action units, head three-axis posture and gaze direction features, and then calculating the time series mean, peak value, jitter amplitude, and gaze stability characteristic parameters as the facial feature values of the facial video data; The method for acquiring speech feature data is as follows: using 16kHz sampling, pre-emphasis and Hamming window processing, extracting MFCC, fundamental frequency, formant, speech rate, pause rate and eGeMAPS emotion features; The method for obtaining the clock drawing features is as follows: 1024×768 resolution input, converted to grayscale image, and then binarized to highlight the clock drawing content; Use opening and closing operations to remove noise and enhance image quality; The clinical information data were obtained by normalizing continuous variables using Z-score and interpolating missing values.
5. The tool according to claim 4, characterized in that The specific method for constructing a basic AI-MCSS model based on facial feature values, voice feature data, clock drawing data, and acquired clinical information data includes the data acquisition step and the following subsequent steps: S1 multimodal feature mapping steps, specifically including facial feature value feature mapping based on Bi-LSTM network, voice feature data feature mapping based on 1D-CNN+LSTM series model, clock drawing data feature mapping based on MobileNetV2, and acquired clinical information data encoding and feature mapping through multi-layer perceptron (MLP), four types of modal feature mapping; In the feature fusion step of the S2 modality, the feature map of step S1 is fed into the Transformer-style multi-head attention module to achieve weighted fusion of cross-modal information and model the semantic relevance between the modalities as the weight of weighted fusion. S3 embeds the fused features into a set of fully connected layers + SoftMax to output the final three classifications: normal cognition, MCI, and suspected MCI.
6. The tool according to claim 5, characterized in that The basic AI-MCSS model introduces at least one of the following completion mechanisms: masking mechanism, migration completion network, and multi-path output fusion.
7. The tool according to claim 5 or 6, characterized in that The specific training method of the basic AI-MCSS model is: internal validation is based on 3,000-10,000 retrospective case data, with 70% of the training set used for model parameter learning; 15% of the validation set used for model parameter adjustment and early stopping strategy; and 15% of the test set used for evaluating the model and selecting the best model hyperparameters.
8. The tool according to claim 7, characterized in that To evaluate the reliability of the AI-MCSS model, an evaluation scale was used to verify and analyze its output results. By comparing and analyzing the prediction results of AI-MCSS with the retest results of the corresponding evaluation scale, the accuracy, sensitivity, and specificity were evaluated and further optimized, which will be given later.
9. The tool according to claim 8, characterized in that By nesting the correction model with the AI-MCSS model and combining the repeated retest results, the specific steps for further optimizing the parameters of the AI-MCSS model include: Step 1: Pre-training the calibration model, including: T1: Recollect 3,000-10,000 retrospective case data, divide these data into groups, and construct multiple correction models with the same number of groups as the grouped data. The recollected 3,000-10,000 retrospective case data are completely different from the 3,000-10,000 retrospective case data used to train the basic AI-MCSS model, and are therefore recollected data. T2 sends the data of each group together into a set of fully connected layers according to steps S1-S2 to form corresponding multiple fully connected features, which are divided into a correction training set and a correction verification set in a ratio of 14:
3. The corresponding correction model is pre-trained and verified for each group, wherein the multiple fully connected features are composed of multiple groups of sub-fully connected features, each group of sub-fully connected features corresponds to a time node in the time series, and each group of sub-fully connected features contains multiple fully connected features, and they are all divided into the correction training set and the correction verification set; Step 2: Improve the basic AI-MCSS model T3 removes the Softmax layer of the trained basic AI-MCSS model, connects the fully connected output end to the input end of each pre-trained correction model to form a model nesting, and inputs the training set of the trained basic AI-MCSS model into the basic AI-MCSS model again. The fully connected features obtained through S1-S3 are input to the input end of the corresponding correction model. The classification function and loss function are constructed at the output end. The classification result obtained by the classification function is compared with the re-test result of the corresponding evaluation scale to obtain the accuracy. The loss function is then back-propagated from the fully connected layer of the basic AI-MCSS to optimize the network parameters of the basic AI-MCSS model connected at each input end, and the validation set is used for verification until the verification accuracy stabilizes and the loss function is minimized. T4 removes the connection between the fully connected output end and the input end of each correction model, and reconnects each end back to the Softmax layer to obtain the improved AI-MCSS model corresponding to multiple moments.
10. The tool according to claim 9, characterized in that The correction model is a long short-term memory model, in which each node unit corresponds to each time node in the time series, and the time interval between node units is 1-6 hours.
11. The tool according to claim 9, characterized in that The method for grouping data based on different blood sugar fluctuation characteristics of time series includes the following steps: Q1 obtains the characteristics of blood glucose changes over time in 3,000-10,000 retrospective case data to form a blood glucose feature vector; Q2 constructs a multidimensional space, divides the feature vectors into vector training and maps them into the multidimensional space, and uses cluster analysis methods to divide different vectors into multiple feature clusters; Q3 clusters each type of feature as a group.
12. The tool according to claim 9, characterized in that The blood glucose feature vector is in the form of ,in, and At successive moments and The blood sugar peak value, at this time the dimension of the multidimensional space is 4, It is the time corresponding to the first blood glucose peak that occurs in chronological order between 0:00 and 24:00 on a 24-hour basis.
13. The tool according to claim 9, characterized in that The blood glucose feature vector is in the form of , that is, with more groups of different time and blood glucose values, is half the dimension of the multidimensional space, For the moment The blood glucose value, the element in the blood glucose feature vector is The subsequent times are selected from 6:00, 7:00, 8:00, ..., and 24:00, which are measured in a 24-hour system with intervals of one hour.
14. The tool according to any one of claims 9 to 13, characterized in that The method of providing patients with two sets of prediction results before and after optimization for reference is specifically to perform the basic AI-MCSS model prediction before optimization for the current MoCA and MMSE assessment scales, and simultaneously input the improved AI-MCSS model corresponding to the corresponding LSTM node unit. If the two prediction results are the same, the diagnosis result is the corresponding prediction result. If the two are different, it indicates that further examination and confirmation are required.