Cognitive impairment identification method and device, electronic equipment and storage medium

By using feature extraction from multimodal data and prediction via logistic regression models, this method addresses the issue of low accuracy in identifying cognitive impairments in existing technologies, achieving efficient and accurate screening for cognitive impairments, and is applicable to large-scale populations.

CN119745322BActive Publication Date: 2025-12-05SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411726430.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-12-05
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Existing intelligent prediction models lack generalization ability and cannot adapt to the real screening needs of large-scale populations, resulting in low accuracy in identifying cognitive impairment.

Method used

By acquiring multimodal data, including voice data and demographic data, machine learning methods are used for feature extraction and classification. A logistic regression model is then used to predict cognitive impairment. An automated feature extraction and classification process is employed, and demographic factors are integrated for prediction.

Benefits of technology

It improves the accuracy and efficiency of cognitive impairment identification, reduces computational complexity, is suitable for large-scale screening applications, has good generalization ability, and can adapt to the needs of different populations and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119745322B_ABST
    Figure CN119745322B_ABST
Patent Text Reader

Abstract

The application provides a cognitive impairment identification method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises: acquiring multi-modal data of a target; performing feature extraction on the multi-modal data to obtain demographic characteristics and a plurality of acoustic characteristics of the target; each acoustic characteristic corresponds to a category of voice data; a category score corresponding to each category is calculated based on each acoustic characteristic; the category score is used to indicate a confidence value of voice data belonging to the corresponding category; and a cognitive prediction model is used to predict cognitive impairment of the target according to the demographic characteristics of the target and the category scores, to obtain a prediction result. The application solves the problem of low accuracy of cognitive impairment identification in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, electronic device, and storage medium for recognizing cognitive impairment. Background Technology

[0002] Mild cognitive impairment (MCI) is a cognitive state that falls between normal aging and dementia. It is primarily characterized by a mild decline in memory, attention, or executive function, but does not yet severely affect daily living abilities. As an early stage of cognitive decline, MCI has significant clinical implications. Studies have shown that MCI patients have a significantly higher risk of progressing to more severe cognitive impairments such as Alzheimer's disease than cognitively normal individuals. Therefore, early identification and intervention for MCI are crucial for slowing disease progression.

[0003] Traditional methods for diagnosing MCI typically rely on neuropsychological tests or imaging examinations. However, these methods are time-consuming, complex, or expensive, making them unsuitable for large-scale screening and long-term monitoring. In recent years, with the development of artificial intelligence technology, intelligent predictive models have provided new solutions for the early screening and precise intervention of MCI.

[0004] However, current intelligent prediction models lack sufficient generalization ability and cannot adapt to the real screening needs of large-scale populations, resulting in poor classification performance.

[0005] As can be seen from the above, how to improve the accuracy of cognitive impairment identification still needs to be addressed. Summary of the Invention

[0006] This application provides a method, device, electronic device, and storage medium for recognizing cognitive impairment, which can solve the problem of low accuracy in recognizing cognitive impairment in related technologies. The technical solutions are as follows:

[0007] According to one aspect of this application, a method for identifying cognitive impairment includes: acquiring multimodal data of a target; the multimodal data reflecting speech data related to cognitive impairment and demographic data related to individual health of the target; extracting features from the multimodal data to obtain demographic features and multiple acoustic features of the target; each acoustic feature corresponding to a category of the speech data; calculating a category score corresponding to each category based on each acoustic feature; the category score indicating the confidence value that the speech data belongs to the corresponding category; and using a cognitive prediction model to predict cognitive impairment of the target based on the demographic features and the category scores of the target, obtaining a prediction result; the cognitive prediction model is trained based on an initial logistic regression model.

[0008] According to one aspect of this application, a cognitive impairment identification device includes: a data acquisition module for acquiring multimodal data of a target; the multimodal data reflecting speech data related to cognitive impairment and demographic data related to individual health of the target; a feature processing module for extracting features from the multimodal data to obtain demographic features and multiple acoustic features of the target; each acoustic feature corresponds to a category of the speech data; a score calculation module for calculating a category score corresponding to each category based on each acoustic feature; the category score is used to indicate the confidence value that the speech data belongs to the corresponding category; and a cognitive prediction module for using a cognitive prediction model to predict cognitive impairment of the target based on the demographic features and the category scores of the target, and obtaining a prediction result; the cognitive prediction model is trained based on an initial logistic regression model.

[0009] According to one aspect of this application, an electronic device includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, implements the cognitive impairment recognition method as described above.

[0010] According to one aspect of this application, a storage medium having a computer program stored thereon, which, when executed by one or more processors, implements the cognitive impairment identification method as described above.

[0011] According to one aspect of this application, a computer program product includes a computer program that, when executed by one or more processors, implements the cognitive impairment identification method as described above.

[0012] The beneficial effects of the technical solution provided in this application are:

[0013] In the above technical solution, the classification accuracy of MCI is effectively improved by combining the output of the audio learning model and demographic data. Meanwhile, the use of a cognitive recognition model based on logistic regression to process multimodal data significantly reduces computational complexity, making it suitable for large-scale screening applications.

[0014] Furthermore, the automated feature extraction and classification process reduces the need for manual intervention, improves screening efficiency, and possesses good generalization ability, enabling it to adapt to the needs of different populations and scenarios. Finally, the reclassification considers demographic factors, further improving the accuracy and practicality of MCI prediction and expanding the model's applicability. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0017] Figure 2 This is a flowchart illustrating a cognitive impairment identification method according to an exemplary embodiment;

[0018] Figure 3 yes Figure 2 A flowchart of step 330 in one embodiment corresponds to the following example;

[0019] Figure 4 yes Figure 3 A flowchart of step 333 in one embodiment corresponds to the following example;

[0020] Figure 5 yes Figure 2 The steps for training the audio learning model in the corresponding embodiment are shown in a flowchart of one embodiment;

[0021] Figure 6 This is a schematic diagram illustrating the specific implementation of a cognitive impairment recognition method in an application scenario;

[0022] Figure 7 yes Figure 6 A schematic diagram illustrating the specific implementation of the voice data processing flow involved in the corresponding embodiment;

[0023] Figure 8 This is a structural block diagram of a cognitive impairment recognition device according to an exemplary embodiment;

[0024] Figure 9 This is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation

[0025] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0026] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this disclosure means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0027] As mentioned earlier, intelligent prediction models provide new solutions for early screening and precise intervention of MCI.

[0028] Intelligent predictive models built upon data such as speech, behavior, and electroencephalography (EEG) offer new solutions for early screening and precise intervention of microcognitive impairment (MCI). Among these, speech, as a non-invasive and readily available biosignal, can effectively reflect changes in cognitive function by extracting the acoustic features of a patient's speech, and has become one of the important research directions for identifying MCI.

[0029] However, on the one hand, current technologies rely heavily on small sample data for model training, lacking sufficient generalization ability and failing to meet the real-world screening needs of large-scale populations, resulting in poor classification performance. On the other hand, feature extraction and cleaning processes are highly dependent on manual engineering, leading to inefficiency and failing to meet the automation and speed requirements of large-scale screening. Furthermore, the failure to fully consider health-related features during model training limits the accuracy and applicability of the models.

[0030] As can be seen from the above, the relevant technologies still suffer from low accuracy in recognizing cognitive impairments.

[0031] Therefore, the cognitive impairment identification method provided in this application can effectively improve the accuracy of cognitive impairment identification. Accordingly, the cognitive impairment identification method is applicable to cognitive impairment identification devices, which can be deployed on electronic devices. The electronic devices can be computer devices configured with the von Neumann architecture, such as desktop computers, laptops, servers, etc.

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0033] Please see Figure 1 , Figure 1 This is a hardware structure diagram of a server according to an exemplary embodiment.

[0034] It should be noted that this server is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Nor should this server be interpreted as requiring or depending on any specific feature. Figure 1 One or more components of the exemplary server 200 shown.

[0035] The hardware architecture of server 200 can vary significantly due to differences in configuration or performance, such as Figure 1 As shown, server 200 includes: power supply 210, interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.

[0036] Specifically, power supply 210 is used to provide operating voltage for the various hardware devices on server 200.

[0037] Interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices.

[0038] Of course, in other examples adapted in this application, interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc. Figure 1 As shown, this does not constitute a specific limitation.

[0039] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it include the operating system 251, application programs 253, and data 255, etc., and the storage method can be temporary storage or permanent storage.

[0040] The operating system 251 is used to manage and control the various hardware devices and application programs 253 on the server 200, so as to enable the central processing unit 270 to perform calculations and processing on the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0041] Application 253 is a computer program formed by computer-readable instructions based on operating system 251 to perform at least one specific task, and may include at least one module ( Figure 1 (Not shown), each module can contain corresponding computer-readable instructions. For example, the cognitive impairment recognition device can be considered as an application 253 deployed on server 200.

[0042] Data 255 can be photos, pictures, etc. stored on a disk, or multimodal data, etc., stored in memory 250.

[0043] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer programs stored in the memory 250, thereby performing calculations and processing on massive amounts of data 255 stored in the memory 250. For example, a cognitive impairment recognition method may be implemented by the central processing unit 270 reading an application program 253 stored in the memory 250.

[0044] Furthermore, this application can also be implemented through hardware circuits or a combination of hardware circuits and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, or combination thereof.

[0045] Please see Figure 2 This application provides a method for recognizing cognitive impairment, applicable to electronic devices, the hardware structure of which can be as follows: Figure 1 As shown.

[0046] In the following method embodiments, for ease of description, the execution subject of each step of the method is an electronic device, but this does not constitute a specific limitation.

[0047] like Figure 2 As shown, the method may include the following steps:

[0048] Step 310: Obtain the multimodal data of the target.

[0049] The target is the object of cognitive impairment identification, which can be a patient or a healthy person, without limitation.

[0050] Regarding multimodal data, it reflects speech data related to cognitive impairment and demographic data related to individual health. Speech data reflects the target's cognitive function status in areas such as speech expression, fluency of thought, and memory ability; demographic data can include the target's basic health information, such as age, gender, education level, lifestyle habits, and medical history, providing important background information for risk assessment of cognitive impairment. Certain demographic information is strongly correlated with the occurrence of cognitive impairment.

[0051] Step 330: Extract features from the multimodal data to obtain the demographic features and multiple acoustic features of the target.

[0052] Each acoustic feature corresponds to a category of speech data.

[0053] First, it should be noted that machine learning methods can be used to classify multimodal data, such as random forests and support vector machines, without limitation here.

[0054] In one possible implementation, such as Figure 3 As shown, step 330 includes the following steps:

[0055] Step 331: Perform structured processing on the demographic data in the multimodal data to obtain demographic features.

[0056] Specifically, structured processing includes cleaning, standardizing, and encoding the target demographic data (such as age, gender, education level, medical history, etc.); and transforming the demographic data into demographic features usable by the model through feature selection and transformation, such as numerical, categorical, or embedded feature representations, to form clear demographic features that can be input into the model for analysis.

[0057] In one possible implementation, demographic characteristics include at least one of years of education, sex, and age.

[0058] Step 333: Based on the voice data, classify the data to obtain at least one class.

[0059] Among them, acoustic features can describe the physical properties of sound, such as frequency, amplitude, and power spectrum, and can also describe the human ear's perception and understanding of sound, conforming to human auditory characteristics, such as speech rate, pause duration and distribution, and speech intelligibility.

[0060] Acoustic features can be used to uncover subtle sound attributes from a physical perspective, and can also capture potential cognitive and emotional cues in sound from a perceptual perspective. This dual descriptive capability of physical and perceptual aspects enables acoustic features to play an important role in the assessment and screening of cognitive impairments.

[0061] The categories include at least one of the following: speech fluency test speech, logical memory test speech, and Johns Hopkins language learning test speech.

[0062] Speech fluency tests assess a target's ability to fluently express words within a limited time, such as saying related words according to a specific topic or starting with a letter. Logical memory tests assess a speaker's memory ability, such as repeating heard sentences or stories with accuracy and completeness. The Hopkins Language Learning Test assesses language learning ability, such as distinguishing similar words or repeating certain phrases.

[0063] It is understandable that the above categories reflect the cognitive level of the target in different aspects, and can provide an important basis for subsequent cognitive recognition.

[0064] Another possible implementation is to automatically determine the category based on the context and task characteristics of the speech data using a classification model.

[0065] Step 335: Based on each category, perform feature extraction processing on the speech data to obtain the acoustic features corresponding to each category.

[0066] First, it should be noted that feature extraction processing can include using audio learning models to extract features from speech data, capturing the time-domain and frequency-domain characteristics of the speech data, such as Mel spectrograms.

[0067] Therefore, for each category, the audio learning model can extract relevant acoustic features. Specifically, for speech fluency test speech, acoustic features such as speech rate, pause distribution, and pitch smoothness are extracted; for logical memory test speech, acoustic features such as key content words and semantic consistency are extracted; and for Hopkins language learning test speech, acoustic features such as speech clarity and speech repetition quality are extracted.

[0068] Different acoustic characteristics can help analyze potential impairments in different cognitive functions. For example, the distribution of pauses in speech responses in speech fluency tests may indicate difficulties in language production or memory retrieval; the clarity of speech responses in the Hopkins Language Learning Tests can be correlated with a decline in the target's control of articulatory organs or language expression ability; the semantic consistency of speech responses in logical memory tests may indicate impaired language comprehension or logical reasoning if the target's expression is incoherent or illogical, etc., without further specific limitations.

[0069] In one possible implementation, step 3333 further includes the following steps: calculating the power spectrum of the speech data based on each category to obtain the power spectrum information corresponding to each category; extracting the Mel spectrogram of the power spectrum information to obtain the acoustic features corresponding to each category.

[0070] First, it should be noted that before calculating the power spectrum of the speech data based on each category, it is necessary to classify the speech data to ensure that the subsequent acoustic feature calculations correspond to the categories. Specifically, based on the above categories, speech data corresponding to different categories can be obtained.

[0071] Then, power spectrum calculation can be performed on the speech data corresponding to different categories. Specifically, it can include the following steps: pre-emphasis, frame segmentation and windowing are performed on the original and unprocessed speech data to obtain the processed speech data; the speech data of each frame is converted from the time domain to the frequency domain by fast Fourier transform to obtain the frequency features corresponding to the speech data; the absolute value of the frequency domain features is squared to obtain the power spectrum.

[0072] Furthermore, after obtaining the power spectrum, it can be processed by a Mel filter to generate spectral information that conforms to human hearing characteristics. The absolute value of the spectral information is then squared, and the processed spectral information is output as a Mel spectrogram. This allows each category to be labeled with a corresponding Mel spectrogram, thereby obtaining the acoustic features corresponding to different categories.

[0073] In addition, deep learning models, such as LSTM and GRU, can be used for acoustic feature extraction, without limitation.

[0074] Through the above process, multimodal data of speech data can be comprehensively utilized to extract key acoustic features reflecting an individual's cognitive state from the speech data, thereby improving the model's ability to identify and predict an individual's cognitive state.

[0075] Step 350: Calculate the category score corresponding to each category based on each of the acoustic features.

[0076] The category score is used to indicate the confidence value of the speech data belonging to the corresponding category. The confidence value refers to the degree of confidence that the classification model is in judging a certain category during the prediction process. The category score can be the logits score, which is not limited here.

[0077] One possible implementation is to process each acoustic feature using a multi-layer neural network of a deep learning model, i.e., an audio learning model. The network structure of this audio learning model is the RESNET18 network, and the final output is a category score.

[0078] Step 370: Use a cognitive prediction model to predict cognitive impairment of the target based on the target's demographic characteristics and scores in each category, and obtain the prediction results.

[0079] The cognitive prediction model is trained based on an initial logistic regression model. It's worth noting that logistic regression was chosen as the foundation for the cognitive prediction model because it effectively handles high-dimensional features and provides clear classification results. Furthermore, this logistic regression model has low computational complexity, making it suitable for processing multimodal data.

[0080] Because there are multiple category scores and demographic features, feature fusion is needed to make full use of the above information.

[0081] Regarding feature fusion, one possible implementation involves fusing the scores of each category with demographic features to obtain fused features; and then using a cognitive prediction model to predict cognitive impairment of the target based on the fused features.

[0082] Specifically, demographic features and scores from various categories can be unified into a single comprehensive representation, i.e., a fusion feature, through weighted averaging, splicing, or other feature aggregation methods.

[0083] For example, demographic features and category scores can be directly concatenated into a long vector. The resulting fused features can be used in cognitive prediction models, which are capable of handling high-dimensional data and automatically learning the complex relationships between features.

[0084] Figure 4 The document illustrates a specific implementation of a cognitive prediction model, such as... Figure 4 As shown, demographic characteristics include the target's gender, age, and years of education. The category features corresponding to the speech fluency test, logical memory test, and Johns Hopkins Language Learning Test are processed by the RESNET18 audio learning model to calculate their corresponding logits scores. The logits scores of the speech fluency test, logical memory test, and Johns Hopkins Language Learning Test are then fused with gender, age, and years of education to obtain fused features. Logistic regression is then used to predict cognitive impairment based on these fused features.

[0085] In one possible implementation, the cognitive impairment prediction process of the cognitive prediction model may include: predicting the cognitive state of the target using a supervised binary classification method based on the target's fusion features, wherein the cognitive state may include a normal state or mild cognitive impairment, without limitation.

[0086] Based on this, by using a cognitive prediction model to make cognitive predictions about the target, we can obtain corresponding prediction results, which can indicate whether the target has a cognitive impairment.

[0087] In one possible implementation, the training process of a cognitive prediction model may include the following steps: acquiring a large-scale sample dataset and dividing the sample dataset into a validation set, a training set, and a test set; constructing a cognitive prediction model based on an initial logistic regression model; training the constructed cognitive prediction model using the training set and adjusting the parameters of the cognitive prediction model according to the validation set; evaluating the predictive performance of the cognitive prediction model using the test set until the predictive performance indicates that the cognitive prediction model has completed training.

[0088] First, it should be noted that large-scale sample datasets include multimodal data from multiple samples, such as speech data, demographic data, and medical records. These data can correspond to different cognitive states, such as healthy individuals, patients with mild cognitive impairment (MCI), and patients with Alzheimer's disease (AD), thereby ensuring the diversity and scale of the dataset and improving the generalization ability of cognitive prediction models.

[0089] Regarding the division of the sample dataset, it can be divided into 60% training set, 20% validation set, and 20% test set, with the specific proportions determined based on the actual situation.

[0090] Furthermore, the parameters of the cognitive prediction model, such as the learning rate and batch size, can be adjusted using the validation set. Based on the results of the validation set, it may be necessary to adjust the parameters of the cognitive prediction model, add or remove features, or change the model structure to improve the generalization ability of the cognitive prediction model.

[0091] Therefore, by using a test set to evaluate the predictive performance of the cognitive prediction model, and calculating various metrics such as accuracy, cognitive rate, and F1 score on the test set, we can determine whether the cognitive prediction model has completed training. For example, if the cognitive prediction model has an accuracy of 95% on the test set, it is considered to have completed training.

[0092] Through the above process, by including multimodal data and covering different cognitive states, the diversity of large-scale datasets helps to improve the generalization ability of the model, enabling it to better adapt to different practical application scenarios. Choosing the logistic regression model as the base model can effectively handle high-dimensional features and provide clear classification results. At the same time, this model has low computational complexity and is suitable for processing multimodal data.

[0093] By combining category scores and demographic data, the classification accuracy of MCI is effectively improved in conjunction with the above embodiments. Furthermore, the use of a cognitive recognition model based on logistic regression to process multimodal data significantly reduces computational complexity, making it suitable for large-scale screening applications.

[0094] Furthermore, the automated feature extraction and classification process reduces the need for manual intervention, improves screening efficiency, and possesses good generalization ability, enabling it to adapt to the needs of different populations and scenarios. Finally, the reclassification considers demographic factors, further improving the accuracy and practicality of MCI prediction and expanding the model's applicability.

[0095] Please see Figure 5 In one exemplary embodiment, the training process of an audio deep learning model may include the following steps:

[0096] Step 410: Obtain sample data.

[0097] The sample data includes sample categories and their corresponding acoustic features;

[0098] First, it should be noted that a pre-trained residual network model is selected as the base network to construct the audio learning model, so as to use the audio learning model to calculate the category score corresponding to each category based on each of the acoustic features.

[0099] In one possible implementation, a Residual Network (ResNet) pre-trained on a large-scale dataset is chosen as the base network. ResNet mitigates the vanishing gradient problem through its residual block structure and is well-suited for efficiently processing Mel spectrograms (i.e., acoustic features).

[0100] In one possible implementation, an audio learning model can be built through transfer learning, which may include the following steps: processing the task based on the acoustic features of the samples, adjusting the input layer, adjusting the output layer, freezing some layers, fine-tuning the model, and evaluating and optimizing the model.

[0101] Step 430: Select the cross-entropy loss function as the optimization objective for the constructed audio learning model.

[0102] It should be noted that the cross-entropy loss function was chosen as the optimization objective because it is suitable for classification problems and can effectively measure the difference between the model's prediction and the actual label.

[0103] The optimization objective is to train the model to improve the accuracy of classification or prediction by minimizing the cross-entropy loss function.

[0104] Step 450: Based on the optimization objective, train the audio learning model using the acoustic features of the samples.

[0105] During training, the acoustic features of the samples are input into the audio learning model. The difference between the predicted output and the true label is calculated using the selected optimization objective (cross-entropy loss function), and the model parameters are adjusted through backpropagation.

[0106] In addition, an appropriate learning rate can be initialized to control the step size of each model parameter update, thereby balancing the model's convergence speed and stability.

[0107] Step 470: During the training process, an adaptive moment estimation optimization method is used to accelerate the gradient descent process and improve the convergence performance of the audio learning model until the convergence performance indicates that the audio learning model has been successfully trained.

[0108] Additionally, one possible implementation involves introducing a dynamic learning rate adjustment strategy. Specifically, this can be achieved by combining adaptive moment estimation optimization methods (such as the Adam optimizer) with the learning rate to dynamically adjust the learning rate based on gradient changes during training. This ensures that the loss value decreases rapidly in the early stages of training and can be finely adjusted in the later stages to improve the final performance of the model.

[0109] Under the above embodiments, the residual network is fine-tuned through transfer learning to adapt to the characteristics of speech data. The cross-entropy loss function is selected as the optimization objective, and an adaptive optimization method is used to accelerate the training process, ensuring that the model can converge quickly and obtain optimized results.

[0110] Figure 6 This is a schematic diagram illustrating the specific implementation of a cognitive impairment identification method in an application scenario. In this scenario, the multimodal data is collected from multiple central hospitals using a unified standard, with no significant differences in recording equipment and quality, and no differences in participant fit. The positive sample label is "MCI" with a positive sample size of 173, and the negative sample label is "NC" with a negative sample size of 239.

[0111] Now combined Figure 6 This section explains the identification of cognitive impairment in this application scenario:

[0112] In step 801, feature processing is performed on the speech data and the demographic data respectively.

[0113] Step 803 involves preprocessing the demographic data.

[0114] Preprocessing includes handling outliers, missing values, and duplicate values. The demographic data distribution is as follows.

[0115]

[0116] Step 805 involves structuring the demographic data to obtain demographic characteristics.

[0117] Demographic characteristics include sex, age, and years of education.

[0118] In step 807, three different categories of speech data are input, and the ResNet18 deep learning model is used to process each type of speech data to obtain the corresponding logits score.

[0119] Specifically, regarding the processing of voice data, Figure 7 A schematic diagram illustrating a specific implementation of a voice data processing flow is shown, such as... Figure 7 As shown, the steps for processing voice data may include:

[0120] (1) Pre-emphasis framing and windowing: The original audio signal is pre-emphasized to enhance high-frequency components. Then, the original audio signal is framed and windowed to reduce the impact of signal truncation on the spectrum.

[0121] (2) Fast Fourier Transform (FFFT): The audio signal of each frame is converted from the time domain to the frequency domain by fast Fourier transform to obtain the frequency signal.

[0122] (3) Squaring of absolute value of signal: The absolute value of the transformed frequency signal is squared to obtain the power spectrum, which is then used for feature extraction.

[0123] (4) Mel filter: The power spectrum is passed through the Mel filter to generate a spectrum based on human auditory characteristics.

[0124] (5) Squaring the absolute value of the signal: The absolute value of the spectrum after passing through the Mel filter is squared to calculate the energy.

[0125] (6) Output Mel spectrum: The processed signal is output as a Mel spectrum, which is the acoustic feature.

[0126] (7) Calculate the corresponding logits score using a deep learning model based on each acoustic feature.

[0127] In step 809, the logits score and demographic features are fused using a splicing method to form a fused feature.

[0128] In step 813, the fused features are classified using a logistic regression model to output a prediction result of mild cognitive impairment.

[0129] In this application scenario, end-to-end deep learning algorithms can be used to directly classify the MCI (Morbidity, Injury, and Cognitive Impairment) of the target population, eliminating the need for manual intervention in intermediate feature extraction and enabling an automated and rapid screening process. Furthermore, it can comprehensively consider demographic characteristics (age, gender, and years of education) to improve the accuracy of MCI prediction, thereby enhancing the model's practicality and reliability. In addition, this application scenario can be extended to other areas such as cognitive impairment screening and mental health assessment, further enhancing its clinical application value.

[0130] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0131] The following are embodiments of the apparatus described in this application, which can be used to execute the cognitive impairment identification method involved in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of the cognitive impairment identification method involved in this application.

[0132] Please see Figure 8 This application provides a cognitive impairment recognition device 900, including but not limited to: a data acquisition module 910, a feature processing module 930, a score calculation module 950, and a cognitive prediction module 950.

[0133] Among them, the data acquisition module 910 is used to acquire the target's multimodal data; the multimodal data reflects the target's speech data related to cognitive impairment and demographic data related to individual health;

[0134] The feature processing module 930 is used to extract features from multimodal data to obtain the demographic features and multiple acoustic features of the target; each acoustic feature corresponds to a category of speech data.

[0135] The score calculation module 950 is used to calculate the category score corresponding to each category based on each acoustic feature; the category score is used to indicate the confidence value of the speech data belonging to the corresponding category;

[0136] The cognitive prediction module 950 is used to predict cognitive impairment of the target based on the target's demographic characteristics and scores in each category using a cognitive prediction model, and obtain the prediction results; the cognitive prediction model is trained based on an initial logistic regression model.

[0137] It should be noted that the cognitive impairment recognition device provided in the above embodiments is only illustrated by the division of the above functional modules when performing cognitive impairment recognition. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the cognitive impairment recognition device will be divided into different functional modules to complete all or part of the functions described above.

[0138] Furthermore, the embodiments of the cognitive impairment recognition device and the cognitive impairment recognition method provided in the above embodiments belong to the same concept, and the specific way in which each module performs its operation has been described in detail in the method embodiments, and will not be repeated here.

[0139] Please see Figure 9 This application provides an electronic device 4000, which may include: a desktop computer, a laptop computer, a server, etc.

[0140] exist Figure 9 In this context, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0141] Data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. This communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.

[0142] Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0143] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0144] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing computer programs having instruction or data structure forms and accessible by electronic device 400, but not limited to these.

[0145] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002.

[0146] The computer program is executed by one or more processors 4001 to implement the cognitive impairment recognition method in the above embodiments.

[0147] Furthermore, this application provides a storage medium storing a computer program, which is executed by one or more processors to implement the cognitive impairment identification method described above.

[0148] This application provides a computer program product, including a computer program that is executed by one or more processors to implement the cognitive impairment identification method as described above.

[0149] Compared to related technologies, combining the output of an audio learning model with demographic data effectively improves the classification accuracy of MCI. Furthermore, employing a recognition model based on logistic regression to process multimodal data significantly reduces computational complexity, making it suitable for large-scale screening applications.

[0150] Furthermore, the automated feature extraction and classification process reduces the need for manual intervention, improves screening efficiency, and possesses good generalization ability, enabling it to adapt to the needs of different populations and scenarios. Finally, the reclassification considers demographic factors, further improving the accuracy and practicality of MCI prediction and expanding the model's applicability.

[0151] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for identifying cognitive impairment, characterized in that, The method includes: Acquire multimodal data of the target; the multimodal data reflects speech data related to cognitive impairment and demographic data related to individual health of the target; Feature extraction is performed on the multimodal data to obtain the demographic features and multiple acoustic features of the target; each acoustic feature corresponds to a category of the speech data. Each category score is calculated based on each of the acoustic features; the category score is used to indicate the confidence value of the speech data belonging to the corresponding category, and the confidence value refers to the degree of confidence that the classification model has in judging a certain category during the prediction process; Using a cognitive prediction model, cognitive impairment is predicted for the target based on the target's demographic characteristics and scores in each category, and the prediction result is obtained; the cognitive prediction model is trained based on an initial logistic regression model.

2. The method as described in claim 1, characterized in that, The feature extraction process on the multimodal data yields the demographic features and multiple acoustic features of the target, including: The demographic data in the multimodal data is structured to obtain demographic features; Based on the speech data, at least one category is determined; the category includes at least one of speech fluency test speech, logical memory test speech, and Johns Hopkins language learning test speech. Based on each of the categories, feature extraction processing is performed on the speech data to obtain the acoustic features corresponding to each category.

3. The method as described in claim 2, characterized in that, The step of performing feature extraction processing on the speech data based on each of the categories to obtain the acoustic features corresponding to each category includes: Power spectrum calculation is performed on the speech data based on each of the categories to obtain power spectrum information corresponding to each category; Mel spectrum diagrams of the power spectrum information are extracted to obtain the acoustic features corresponding to each category.

4. The method as described in claim 1, characterized in that, The method of using a cognitive prediction model to predict cognitive impairment of the target based on the target's demographic characteristics and scores in each category includes: The scores of each category are fused with the demographic features to obtain the fused features; The cognitive prediction model is used to predict cognitive impairment in the target based on the fusion features.

5. The method as described in claim 1, characterized in that, A pre-trained residual network model is selected as the base network to construct an audio learning model, and the audio learning model is used to calculate the category score corresponding to each category based on each of the acoustic features. The training process of the audio learning model includes: Acquire sample data, which includes sample categories and corresponding sample acoustic features; The cross-entropy loss function is selected as the optimization objective for the constructed audio learning model; Based on the optimization objective, the audio learning model is trained using the acoustic features of the samples; During training, an adaptive moment estimation optimization method is used to accelerate the gradient descent process and improve the convergence performance of the audio learning model until the convergence performance indicates that the audio learning model has been successfully trained.

6. The method as described in claim 1, characterized in that, The training process of the cognitive prediction model includes: Obtain a large-scale sample dataset and divide the sample dataset into a validation set, a training set, and a test set; Based on the initial logistic regression model, a cognitive prediction model is constructed; The constructed cognitive prediction model is trained using the training set, and the parameters of the cognitive prediction model are adjusted according to the validation set. The prediction performance of the cognitive prediction model is evaluated using the test set until the prediction performance indicates that the cognitive prediction model has completed training.

7. The method according to any one of claims 1-6, characterized in that, The demographic data includes at least one of years of education, gender, and age.

8. A cognitive impairment recognition device, characterized in that, include: The data acquisition module is used to acquire multimodal data of the target. The multimodal data reflects speech data related to cognitive impairment and demographic data related to individual health. The feature processing module is used to extract features from the multimodal data to obtain the demographic features and multiple acoustic features of the target. Each of the acoustic features corresponds to a category of the speech data; The score calculation module is used to calculate the category score corresponding to each category based on each of the acoustic features; the category score is used to indicate the confidence value of the speech data belonging to the corresponding category, and the confidence value refers to the degree of confidence of the classification model in judging a certain category during the prediction process; The cognitive prediction module is used to predict cognitive impairment of the target based on the demographic characteristics and category scores of the target using a cognitive prediction model, and to obtain the prediction result; the cognitive prediction model is trained based on an initial logistic regression model.

9. An electronic device comprising at least one processor and at least one memory, wherein, The memory stores a computer program, characterized in that, when the computer program is executed by the processor, it implements the cognitive impairment recognition method as described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by one or more processors, it implements the cognitive impairment recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cognitive disorder detection method, system and related device

    CN118412120A

  • System, method, and computer program for cognitive training

    US20210312942A1